Skip to content

Map FRS EMPSTATI 11 to OTHER_INACTIVE, not LONG_TERM_DISABLED - #526

Draft
MaxGhenis wants to merge 4 commits into
mainfrom
fix/frs-empstati-other-inactive
Draft

MaxGhenis wants to merge 4 commits into
mainfrom
fix/frs-empstati-other-inactive

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What was wrong

create_frs set employment_status with categorical(person.empstati, 1, range(12), EMPLOYMENTS).fillna("LONG_TERM_DISABLED"). EMPLOYMENTS has 11 entries, so zip(range(12), EMPLOYMENTS) stops at code 10 and EMPSTATI 11 had no mapping. The fillna then made every code-11 adult LONG_TERM_DISABLED.

In the UKDS FRS 2024-25 data dictionary (SN 9563, adult table, EMPSTATI "Adult - Employment Status - ILO definition"), 11 is "Other Inactive" and 9 is "Permanently sick/disabled". policyengine-uk's EmploymentStatus already has OTHER_INACTIVE.

In the raw FRS 2024-25 adult table, 916 adults (2.05m by FRS grossing weight) have code 11, beside 1,678 (3.24m) with code 9. All of them came out LONG_TERM_DISABLED.

The fix

  • FRS_EMPSTATI_EMPLOYMENT_STATUS maps codes 1-11 explicitly, one data-dictionary label per line, using EmploymentStatus member names (a rename in policyengine-uk breaks loudly).
  • derive_employment_status_from_frs(empstati, is_adult_record): child-table rows are CHILD. An adult whose code is missing or not in the table raises a ValueError naming the codes, instead of falling back to a guessed status.
    • CI build logs are public, so the message lists the codes (for example "0, 12, blank") without per-code counts, and reports a total under 10 adults as "Fewer than 10 FRS adults".
    • The codes are formatted from the numeric values rather than Series.astype(str), whose NaN handling changed in pandas 3.
    • Why raise rather than keep a default: the old default is how code 11 became "long-term disabled". LONG_TERM_DISABLED is a health status that feeds the ESA proxies, and OTHER_INACTIVE would equally be a guess for a code nobody has read.
    • This cannot break a current build. Every adult in the 2020-21, 2022-23, 2023-24 and 2024-25 releases on disk has a code from 1 to 11; no adult is blank.
    • Children: the child table has no EMPSTATI, and the person table fills the blank with 0. The old code reached CHILD through code 0; the new code uses the row's table. categorical's NaN default of 1 (FT_EMPLOYED) never applied, because the combined table is already filled.

Who reads employment_status

In uk-data (git grep at origin/main b45c373, and the diffs of all 30 open uk-data PRs):

In policyengine-uk (git grep of origin/main at 3c48247eb and at 84f5ad43a, and of the heads of all 129 open PRs):

  • No variable formula reads employment_status.
  • dynamics/labour_supply.py calculate_excluded_from_labour_supply_responses reads it only for self-employment and STUDENT, so it is unchanged.
  • No variable is defined for esa_health_condition_proxy, esa_support_group_proxy or legacy_jobseeker_proxy. They are saved in the dataset but nothing in the model reads them.
  • In open PRs, employment_status appears only as test input (#2007 sets STUDENT) or in labour-supply exclusion tests (#1915), so none of them is affected.

Impact (real build and microsimulation)

Build. I ran production builds (512 epochs, OA clones, torch seed 0, web targets from one shared HTTP cache) of main b45c373 and of this branch at d398400. Two builds of main are output-identical, so seeded builds reproduce.

  • Every column of the household and benefit-unit tables is identical between main and the branch, including weights, so calibration did not move.
  • In the person table, only employment_status, esa_health_condition_proxy and esa_support_group_proxy differ.
  • The later commits change only the text of an error that no current FRS release reaches.
Dataset Change Person records Survey households Weighted (m)
FRS LONG_TERM_DISABLED → OTHER_INACTIVE 916 876 2.05
FRS leave the ESA health-condition proxy 839 803 1.94
FRS leave the ESA support-group proxy 833 797 1.93
Enhanced FRS LONG_TERM_DISABLED → OTHER_INACTIVE 3,006 876 2.30
Enhanced FRS leave the ESA health-condition proxy 2,778 803 2.10
Enhanced FRS leave the ESA support-group proxy 2,754 797 2.09

In the enhanced FRS, the 3,006 records are the 916 FRS rows plus their SPI copies, capital-gains clones and CGT band donors. No record enters either proxy, and the JSA proxy is unchanged. In the FRS, LONG_TERM_DISABLED falls from 5.30m to 3.24m.

Microsimulation. For each enhanced FRS, I ran one policyengine-uk main (3c48247eb) Microsimulation that calculated every one of the 976 variables for 2025 and 2026, and fingerprinted each array.

  • Every array is bit-identical between main and the branch except employment_status itself.
  • Neither run had a calculation error.
  • The labour-supply exclusion mask is unchanged.
  • So the change moves no tax, benefit, income, poverty or budget figure.

In the 2025 microsimulation, LONG_TERM_DISABLED goes from 5.81m to 3.50m and OTHER_INACTIVE from 0 to 2.31m.

A rerun on policyengine-uk main 84f5ad43a (2.109.1) is queued behind the shared build slot; its result will be added here.

Tests

  • test_frs_employment_status.py:
    • Codes 0-11 checked exhaustively, as adult and as child rows.
    • The code table pinned to the data dictionary's labels.
    • Adult codes plus CHILD are a bijection onto EmploymentStatus.
    • Unknown adult codes (0, negative, 12+, non-integer, NaN) raise.
  • Hypothesis properties:
    • The mapping is row-wise and commutes with row order.
    • Any unknown adult code anywhere fails.
    • For any ages, State Pension ages, hours and reported flags, the ESA health proxy is exactly reported ∧ working age ∧ code ∈ {9, 10}.
    • The support group is a subset of the health proxy and only code 9.
    • Code 11 never sets either proxy.
  • test_legacy_benefit_proxies.py: create_frs runs end to end for every code 1-11 (status and both ESA proxies), for a child row next to a code-11 adult, and for code 12, which is rejected.
  • Built-dataset checks (CI's make data output) on the FRS and enhanced FRS:
    • Every status is a valid EmploymentStatus.
    • OTHER_INACTIVE is present.
    • Neither ESA proxy is set for anyone outside the two sick/disabled statuses.
  • Failure-message tests check that no per-code count appears, that a total under 10 is suppressed, and that a blank code next to a numeric one is listed as "0, 12, blank".
  • Mutation check: 7 of 7 mutations are killed. They include the old behaviour, 11 → LONG_TERM_DISABLED, unknown codes → a default, 9 and 10 swapped, OTHER_INACTIVE added to the ESA health statuses, and child rows treated as adults.

Adds hypothesis as a dev dependency. The pyproject.toml and uv.lock change is identical to #522's and #525's, so whichever lands second merges cleanly.

axiom: n/a: data mapping of an FRS survey code to an input variable; no policy rule changes.

microcosm's UK runtime (uk_runtime/frs_employment.py) keeps the old map on purpose, for parity. A separate session is porting this fix there, with a differential test against this PR's function.

Reviews: r1 approved d398400 (low findings, addressed in 5f9912d). r2 on 5f9912d is running.

Part of the batched uk-data release (d833); not to be merged on its own.

🤖 Generated with Claude Code

The FRS employment_status mapping zipped range(12) with 11 statuses, so
EMPSTATI 11 ("Other inactive" in the UKDS FRS 2024-25 data dictionary) was
unmapped and the fillna fallback made it LONG_TERM_DISABLED. That also put
other-inactive working-age adults into the ESA health-condition and
support-group proxies.

Map the adult codes through an explicit table keyed to the data dictionary,
make child-table rows CHILD, and fail the build on any adult code the table
does not know rather than guessing a status. Every adult in the 2020-21,
2022-23, 2023-24 and 2024-25 releases carries a code from 1 to 11.

Tests cover codes 0-11 exhaustively, Hypothesis properties of the mapping and
the ESA proxies, create_frs end to end for every code and a child row, and
the built datasets. Adds hypothesis as a dev dependency (same pyproject and uv.lock
change as #522 and #525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 2 commits October 2, 2026 13:20
CI build logs are public, so the error now names the unknown codes without
per-code counts and reports a total under 10 adults as "fewer than 10".
Raised by the microcosm port session; no change on current FRS releases,
where every adult has a code from 1 to 11.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under pandas 3, Series.astype(str) keeps NaN as a float, so sorting a set
holding a blank code and another unknown code raised TypeError and hid the
codes. Format from the float codes instead ("0, 12, blank"), which behaves
the same under pandas 2 and 3. Found by the microcosm port session's
differential under pandas 3.0.3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The row-order property recomputed the mapping for every index; compute it
once. The test comment now says what was checked for the older FRS releases
(every adult has a code from 1 to 11), not that their labels match. Drop
redundant parentheses in the create_frs test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant