Skip to content

Draw SPI incomes within earnings groups set by FRS employment status - #529

Draft
MaxGhenis wants to merge 8 commits into
mainfrom
spi-imputation-employment-status
Draft

MaxGhenis wants to merge 8 commits into
mainfrom
spi-imputation-employment-status

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Draft. This changes the published dataset, so merging it is a data release. uk-data PRs land together as one batched release when Max says go (d833). No dataset was uploaded or released from this branch. Fixes #504.

Problem

impute_income stacks a zero-weight copy of 10,000 FRS households (household_is_spi_synthetic), and calibration later gives them weight. Their six incomes (employment, self-employment, savings interest, dividends, private pension, property) are redrawn by a QRF trained on the SPI, whose only predictors are age, gender and region. Every other column, including employment_status and hours_worked, stays the FRS donor's. So a row's incomes do not depend on whether it is a child, an employee, self-employed, or out of work.

On main's build (b45c373, seed 0), the SPI rows and their capital-gains clones hold 8.63m of 31.14m households. Share of people with earnings, by employment status, at 2024-25 calibrated weights:

Status FRS rows: pay > 0 SPI rows: pay > 0 FRS rows: profit > 0 SPI rows: profit > 0
Child (FRS child table) 0% 99.7% 0% 1.0%
Employee 98.4% 84.0% 2.8% 7.3%
Self-employed 6.4% 81.3% 89.5% 8.6%
Out of work (unemployed, retired, student, carer, sick or disabled) under 10 households 46.3% 0% 7.4%

In money, SPI-row children carry £73.3bn of employment income (#504 found the same, from the file as released). SPI rows out of work carry £54.3bn of pay and £6.8bn of profit. SPI-row self-employed people carry £41.3bn of pay but only £3.4bn of profit.

Everything that reads status and income together inherits this:

  • Universal Credit hours and conditionality;
  • the minimum income floor (policyengine-uk#2081 reads any profit as gainful self-employment when Set UC gainful self-employment from the FRS main-job status #525's input is absent);
  • income tax and NI on children's pay;
  • the second-stage QRF, which draws benefit reports from these incomes.

What the SPI can support

From the SPI 2022-23 Public Use Tape documentation (UKDS SN 9422, Annex A):

  • No ILO employment status, hours or full/part-time split. Status cannot be a QRF predictor, because the SPI has nothing to train it on.
  • Income sources, which it does have:
    • PAY ("Pay from employment net of benefits and foreign earnings"), with EPB and TAXTERM;
    • PROFITS ("Gross profits assessable for all sources of self-employment income"). Its range starts at 0, so losses are not recorded;
    • SEINC_NUM, the "Indicator for self-employed cases (those submitting pages in their tax return for income from trades or partnerships)". This finds traders whatever their profit: 17,298 records file self-employment pages with zero PROFITS;
    • MAINSRCE, the main source of income: pay, occupational pension, sole trader, partnership, other, or claims case. It is the largest income, not the main job, so this PR does not use it.

Change

1. Draw SPI incomes within earnings groups (imputations/income.py)

The FRS and the SPI share one thing that matters here: whether a person has pay and whether they have a trade. That gives four earnings groups.

Group SPI records (2022-23, weighted) FRS rows
EMPLOYEE pay > 0, no trade (33.0m) employee main job (FT/PT_EMPLOYED), or recorded pay, and no trade
SELF_EMPLOYED SEINC_NUM = 1 or PROFITS > 0, no pay (3.9m; 93.5% with a profit) self-employed main job, or recorded profit, and no pay
EMPLOYEE_AND_SELF_EMPLOYED both (1.5m; 81% with a profit) both, e.g. a self-employed main job plus an employee second job
NO_EARNINGS neither (12.3m) everyone else: retired, unemployed, student, carer, sick or disabled, other inactive
NOT_IMPUTED n/a the FRS child table, and anyone under 16. They keep their own incomes, which are zero
  • One QRF per group, fitted on that group's SPI records only, still on age, gender and region. So:
    • an employee always draws pay;
    • a self-employed person always draws a trade, whose profit is zero where the SPI trader's is;
    • someone out of work draws pension, property and investment income from SPI people with no earnings;
    • a child takes no draw.
  • The training sample is still a weighted resample of the SPI (100k; 10k in TESTING). It is now drawn within each group, with at least 10% of the sample per group: 65.1k employee, 24.2k no earnings, 10k self-employed and 10k both.
  • The cached model (income_spi_2022_23.pkl, 2.1 GB as before) holds one QRF per group. A cache in the old single-model format is retrained.
  • Main's FRS-half dividend draw goes through the same groups, so FRS children keep their own dividends. Keep FRS-reported dividends and key them on person_id #498 removes that draw; the two merge cleanly in either order.

Why not the other options

  • Status as a QRF predictor. The SPI has no status to learn from. A proxy built from income sources is these groups. As a soft predictor, the forest would usually split on it but not always. Separate models make the agreement hold by construction, and the tests check it.
  • Re-deriving status from the imputed incomes. That throws away what the FRS knows and the SPI does not: who is a child, retired, a student, a carer, sick, or unemployed. It would also need hours, which the SPI does not have.

2. Calibrate to LFS employees and self-employed (targets/sources/ons_labour_market.py)

Making rows coherent exposed a conflict in the calibration. The local targets include HMRC's count of income-tax payers with employment income in each constituency and local authority (SPI table 3.15, 30.0m in total). That count is annual. It includes people who had pay for part of the year but whose FRS status at interview is out of work, and in the FRS those people have no pay.

On main, SPI-row children and out-of-work rows with pay supplied part of that count. Once they draw no pay, calibration met the count by moving weight from people out of work to employees.

Nothing tied employment status to an official total, so this adds national targets for the ONS LFS levels of employees (MGRN) and self-employed (MGRQ), annual averages for 2022-2025, counted from the FRS ILO main-job status. The LFS also counts working dependants aged 16-19, whom the FRS child table records without a job. On FRS 2024-25 grossing weights the FRS has 28.1m employees, against the LFS's 29.1m for 2024.

3. Report, not train on, the local HMRC employment-income counts (create_datasets.py)

The two national LFS targets could not outweigh about 1,000 local count targets: with both trained, employees stayed at 32.1m. VALIDATION_ONLY_LOCAL_TARGETS moves hmrc/employment_income/count to the calibrator's existing validation mode, so it is still logged but no longer trained. The local amounts of employment income, the self-employment counts, and the national counts by income band still train. This is a calibration-methodology change; it is a separate commit (fc46a56) so it can be reverted on its own.

Labour-market composition across builds

All builds are seeded production builds (512 epochs, PE_UK_DATA_OA_CLONES=1, seed 0) with the same cached web targets. People by FRS employment status at 2024-25 calibrated weights, FRS and SPI rows together:

Build Employees Self-employed Out of work (16+) Children
main (b45c373) 28.2m 4.8m 21.2m 15.1m
placebo: main with only the SPI draw's seed changed 28.3m 4.6m 21.3m 15.1m
change 1 only (708ab3e) 32.2m 4.0m 18.1m 14.9m
changes 1 and 2 (40fbfea) 32.1m 4.4m 17.9m 14.9m
changes 1, 2 and 3 (this head) pending pending pending pending
ONS LFS, 2024 (2025) 29.1m (29.6m) 4.3m (4.4m)
FRS 2024-25 grossing weights 28.1m 4.2m 21.3m 14.7m

The placebo is main rebuilt with a different random draw for the SPI incomes and nothing else changed. It is the noise floor for every comparison below.

Target fit for this head: pending the final build.

Impact

Pending: the final build of this head is running; microsimulation results follow it.

Interplay with open uk-data PRs

Invariants (Hypothesis property tests)

tests/test_spi_income_earnings_groups.py:

  1. FRS rows. Children (FRS child table, or under 16) are NOT_IMPUTED. Every other row is in one group. That group has pay if and only if the main job is as an employee or the FRS records pay, and has a trade if and only if the main job is self-employment or the FRS records a profit. Every policyengine-uk EmploymentStatus is covered.

  2. SPI records. Pay if and only if PAY + EPB + TAXTERM > 0. A trade if and only if SEINC_NUM = 1 or PROFITS > 0.

  3. Both mappings are monotone, and each gives the same answer elementwise as row by row.

  4. Sample allocation. Every group with weight gets at least 10% of the sample. Groups without weight get none. Groups above the floor get their weighted share.

  5. generate_spi_table resamples each group only from its own records, in those counts.

  6. Draws from a real fitted per-group QRF.

    • Pay is positive exactly in the groups with pay.
    • Profit is zero in the groups without a trade.
    • Every draw is non-negative.
    • NOT_IMPUTED rows get no draw, and the index is preserved.
    • The trade groups draw both zero and positive profits.
  7. apply_income_draws overwrites drawn rows and leaves NOT_IMPUTED rows, other columns and its input untouched.

  8. The cache. A cache in the old format, or missing a group, is retrained; a current one round-trips.

  9. On a built enhanced FRS. On SPI rows:

    • every employee has pay;
    • every child has no earnings;
    • over 60% of the self-employed have a profit;
    • under 2% of people out of work have earnings.

    This check fails on main's build, at its first assertion.

tests/test_lfs_employment_targets.py:

  • the ONS values, pinned independently;
  • year resolution;
  • discovery of the source module;
  • a property for the household count column;
  • on a built enhanced FRS, employees and the self-employed are within 8% of the LFS. This also catches the local counts going back into training.

Mutation check. I made 10 deliberate defects in income.py and the tests caught each one: status ignored for pay or for a trade, children drawn, SEINC_NUM ignored, child rows overwritten, NaN draws kept, one model for every group, resample ignoring groups, no group floor, and a cache accepting a missing group.

Full suite. The full test suite on the changes 1 and 2 build passes, except the LFS employee check that change 3 fixes.

Not in this PR

Checklist

  • Data. Aggregates only; no record-level values. Cells under 10 survey households are suppressed. Nothing was uploaded.
  • axiom: n/a: survey imputation and calibration targets, no policy rule.

🤖 Generated with Claude Code

MaxGhenis and others added 8 commits October 2, 2026 12:37
The SPI income model drew every SPI-synthetic row's six incomes from age,
gender and region alone, while the row kept its FRS donor's employment
status. Children drew pay, the unemployed and retired drew earnings, and
self-employed rows rarely drew a profit.

Fit one QRF per earnings group (employee, self-employed, both, neither) on
the SPI records in that group, using PAY and the SPI self-employment
indicator SEINC_NUM, and draw each FRS row from the group its employment
status and recorded earnings put it in. Children keep their own incomes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same lines as #514 and #524, so whichever lands second merges cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis properties for the FRS and SPI group mappings, the training
sample allocation, the per-group model's draws and the donor values kept
for children, plus a check on a built enhanced FRS. The cache tests write
the per-group format.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With SPI incomes drawn by earnings group, the SPI-synthetic rows no longer
give pay to children and people out of work. Calibration then met the HMRC
counts of income-tax payers with employment income, which are annual and
include part-year earners, by moving about 4m people's weight from out of
work to employees (32.2m against the LFS's 29.1m for 2024). Nothing tied
employment status to an official count.

Add national targets for the ONS LFS employee and self-employed levels
(MGRN, MGRQ annual averages, 2022-2025), counted from the FRS ILO main-job
status. Move the status groups to utils/employment_status.py so the income
imputation and the targets share them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A loaded runner tripped Hypothesis's too_slow health check while the
properties themselves held. Drop the deadline and that health check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HMRC's counts of income-tax payers with employment income by area are
annual, so they include people with pay for part of the year whose FRS
status at interview is out of work. Trained on, they held employees at
32m against the LFS target of 29.6m. The area amounts and the national
counts by income band still train; the counts stay in the calibration
logs as validation targets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Values pinned against the ONS release, year resolution, source discovery,
a Hypothesis property for the household column, and a check that a built
enhanced FRS lands near both counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPI income imputation gives every child in a synthetic household employment income (£67bn on under-16s)

1 participant