Skip to content

Rebuild a subsampled simulation from its inputs, not its calculated values - #573

Draft
MaxGhenis wants to merge 1 commit into
masterfrom
fix-subsample-inputs-only
Draft

MaxGhenis wants to merge 1 commit into
masterfrom
fix-subsample-inputs-only

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

subsample rebuilt the simulation from to_input_dataframe(include_computed_variables=True), that is, from every stored value, calculated ones included, all reloaded as inputs. A formula result calculated before subsampling therefore replaced its formula for good, and results after subsampling depended on what had been calculated before:

  • a formula returning 7 that ends in 2012, calculated at 2012, then subsampled: 2013 is 7 (carried over), against 0 for a simulation subsampled straight away (executed on master b78b0ba);
  • the frozen value also survives apply_reform (_invalidate_all_caches keeps set_input values) and later set_input calls on its inputs.

Found while reviewing #562 (soundness review, finding 7).

The rule

subsample now exports the values the simulation was given: every period recorded by set_input (dataset loading included) and still stored, on the branches it reads, for every variable, including variables with a formula. So formula-backed structural IDs that come from the dataset (the reason for 56c8b82) and dataset values overriding a formula are kept, and calculated values are not. to_input_dataframe is unchanged; its loop moved into _to_person_dataframe, which both use.

Invariants

  • Whatever is calculated before subsample, the subsampled simulation stores exactly what a simulation subsampled straight after loading stores, and every later result agrees. Property test: tests/core/test_subsample_inputs_only_property.py (Hypothesis, 150 examples: calculations of formula, group, weight and input variables at the dataset year and others, carry-over on and off, then later requests). It fails on master.
  • Sampling probabilities and the total weight are unchanged for a simulation subsampled straight after loading.

Downstream impact

Real runs on master b78b0ba and on all three of #571/#572/#573 merged together (scratch merge 37aadada), each output compared array by array (compare.py). The PE-US 2026 run was also repeated on each branch alone, and the 2035 run on #572 alone; all identical too.

Run Outputs compared Result
policyengine-us 6c5170fd, eCPS 2024, 3,000-household subsample, 2026 (income tax, itemizing and SALT branches, state taxes, CTC, EITC, SNAP, household and SPM net income, weights, marginal tax rates) 25 arrays bitwise identical
same, 2035 25 arrays bitwise identical
policyengine-uk 7b9fc379, enhanced FRS 2024-25, full sample, 2026 (income tax, NI, UC and legacy benefits, Pension Credit, HB, council tax, HBAI income, poverty, marginal tax rates) 36 arrays bitwise identical

No published pipeline calls subsample (policyengine.py and the APIs build one simulation per year); policyengine-us uses it in its microsimulation tests and docs.

The PE-US runs subsample straight after loading (--subsample=3000), so they also show that this branch samples the same households with the same weights as master there.

Tests

  • tests/core/test_subsample_inputs_only.py: 6 regressions (4 fail on master: carry-over past a formula's end, stored values, following a new input, following a reform; 2 pin kept behaviour: dataset values for a formula variable, formulas recomputed on the sample).
  • tests/core/test_subsample_inputs_only_property.py, behind pytest.importorskip("hypothesis").
  • Existing subsample tests (formula-backed IDs, fast cache, baseline branch) pass.
  • Full suite: 1150 passed, 4 skipped, 1 xfailed.

Composition

Reads the _user_input_keys record. #561 keeps that record in step with storage (deletes, clones, subsample reset); with it merged, values deleted from storage after being set also stop being exported. Until then the existing intersection with stored periods covers the common case.

axiom: n/a: core engine, no policy encoding

🤖 Generated with Claude Code

subsample exported every stored value, calculated ones included, and
loaded them all back as inputs. A formula result calculated before
subsampling then replaced its formula for good: it was carried over
past the formula's end and survived apply_reform and later set_input
calls, so results after subsampling depended on what had been
calculated before.

It now exports the values the simulation was given (loaded from the
dataset or passed to set_input), for variables with a formula too, so
formula-backed structural IDs from the dataset are still kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant