Conversation
The State Pension age parameters held one age per year and stayed at 66 from 2020, so the model never applied the Pensions Act 2014 s.26 rise to 67 for people born on or after 6 April 1960. From 2026-27 every 66-year-old counted as over State Pension age. Encode Pensions Act 1995 Sch 4 para 1 by date of birth, row for row: age_by_birth_date (the age, in months) and day_by_birth_date (the day, where the statute sets one), for women and for men born on or after 6 December 1953, plus rule (1) for men born earlier. A person attains State Pension age on the later of the two. months_since_last_birthday places each date of birth within the year of age, measured at 6 October, the middle of the fiscal year. A fractional age is the exact age; single households use the middle of the year of age; representative microdata spreads each single year of age and sex evenly by weight, so the weighted share over State Pension age is the statutory share (three quarters of 66-year-olds in 2026-27, a quarter in 2027-28, none after). state_pension_age is now the person's own State Pension age, and is_SP_age, the basic/new State Pension split and the Savings Credit age test (SPCA 2002 s.3(1)(a), including its age-65 limb) all follow it. No additional State Pension is paid below State Pension age. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A birth instant within a day is the day starting at or after it, so a person's legal age on 6 October is their age and each age is attained at the commencement of the anniversary (Family Law Reform Act 1969 s.9(1)). The time-of-day carry onto the anniversary is gone. - One helper gives exact age in months (float64), capping months since the last birthday a few minutes short of 12 so it never rounds onto the next birthday. Hypothesis found that edge. - Birthdays are spread over the year in any simulation built from data (Simulation.built_from_dataset), not whenever weights exceed a million, so a constituency or local authority filtered from the data is still spread, and filter_dataset carries each person's place into an extract. - Rewrite the old Savings Credit cases from dates of birth, keep one that tests the state_pension_age override, and test the attainment day exactly in the day-level differential test. - male/born_before is a YYYYMMDD number like the scales; labels no longer hardcode the date; the NI reference notes its numbering; the removed fragment and docs say how to reform the timetable now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t simulation uc_deduction_random_draw, uc_deduction_type_random_draw and attends_private_school decided between microdata imputation and household defaults by testing whether total weight was below 1e6. policyengine.py builds constituency and local-authority simulations by filtering rows from the national data (RowFilterStrategy), so those fell below the threshold and got household defaults: no UC deductions and no private school attendance. Each now reads Simulation.built_from_dataset (added in #1899). filter_dataset carries each person's attends_private_school into an extract, since one household alone would rank at the 100th income percentile. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- attends_private_school no longer raises when no household has weight (the old 1e6 gate returned early; MicroSeries cannot rank zero weight). - The weight-scale property scales by powers of two, so invariance is exact; the national fixture carries national-scale weight, so the region test compares a national run with a constituency-sized one. - Replace the vacuous attends_private_school YAML cases with situation cases that fail under the old gate. - Document what filter_dataset carries, and that the changelog's filtered regions rank private school attendance locally. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1902
Stacked on #1899. Do not merge before it. This branch builds on #1899's head (c46893e), which adds
Simulation.built_from_dataset. Until #1899 merges, the diff below includes its commits. This PR's own changes are the last two commits.Summary
Three variables chose between microdata imputation and household defaults by testing whether total weight was below 1e6:
uc_deduction_random_drawuc_deduction_type_random_drawattends_private_schoolperson.household(...)), so in effect peoplepolicyengine.py builds constituency and local-authority simulations by filtering rows from the national data.
src/policyengine/countries/uk/regions.pyusesRowFilterStrategyonconstituency_code_oa/la_code_oa.filter_dataset_by_household_idsinsrc/policyengine/utils/entity_utils.pykeeps rows without rescaling weights.run()insrc/policyengine/tax_benefit_models/uk/model.pywraps them inUKSingleYearDatasetand callsMicrosimulation. So every constituency and local authority fell below the threshold and got household defaults: no UC deductions, and no private school attendance even under a private school VAT reform.Each variable now gates on
getattr(<entity>.simulation, "built_from_dataset", False), which is true for any simulation built from data, however little weight it carries.attends_private_schoolalso loses a deadhasattr(person.simulation, "dataset")check (Simulation.datasetis a class attribute, so it always held).filter_datasetnow carries each person'sattends_private_schoolinto the household it extracts, as #1899 does formonths_since_last_birthday. A household alone ranks at the 100th income percentile (rate 0.47 × 0.85), so without this about 40% of children in an extract would be assigned to private school. UC draws need no carrying: they hashbenunit_id, which the extract keeps.attends_private_schoolalso no longer raises when no household has weight. The old gate returned before ranking, andMicroSeriescannot rank zero total weight. Households without weight stay at percentile 0, as before.The
attends_private_schoolYAML cases were vacuous under the new gate (situations always return False), so they are replaced with situation cases: a household with 1e9 of weight and the top income attends no private school unless set, and a set value is kept.clone(),get_branch()andsubsample()copy or keep the instance__dict__, so the flag survives intobaselineand branch simulations.Invariants (stated and tested)
policyengine_uk/tests/test_data_built_imputations.py:uc_has_deduction,uc_deduction_combination,uc_deductionsandattends_private_schoolunchanged (Hypothesis). Powers of two scale exactly in floating point, so the property holds exactly and can't flake on percentile boundaries.splitmix64_uniform(benunit_id)draws, some deductions and some private school attendance.filter_datasetextract reproduces the full simulation's UC deductions and private school attendance, and the test asserts both sets are non-empty.All six fail on #1899's head, and the first new YAML case fails there too. The extract test also fails with the new gates but without the
filter_datasetcarry.Intended exception: private school attendance is not row-filter invariant. It ranks incomes within the simulated population, so a constituency ranks against itself (see caveats).
Constituency and local-authority runs, before and after
Real runs, following policyengine.py's path: filter the national tables by
constituency_code_oa/la_code_oa/region, buildUKSingleYearDataset+Microsimulation, and calculate 2026.632 constituencies (sums over constituency runs)
363 local authorities
months_since_last_birthdayspreads birthdays within the simulated population. UC paid differs by more than 0.1% in 168 constituencies. I did not trace every case.Caveats and follow-ups
shareholdingandcorporate_land_value. In a filtered constituency this puts the whole national total on the constituency. For E14001063 (87 records),corporate_tax_incidenceis £34,863m in the constituency run vs £28m for the same households nationally, andbusiness_ratesis £31,733m vs £25.6m. Household net income is −£35,634m vs £2,328m. This affects every constituency run in policyengine.py (follow-up task). The net-income figures above are differences, in which it cancels.filter_datasetaffect only data-built simulations. Household calculators (situations) are unchanged: they got the defaults before and still do.Tests run
test_data_built_imputations.py: 6 passed.test_uc_deductions.py+test_state_pension_age.py: 37 passed.contrib/labour/attends_private_school.yaml+private_school_vat.yaml: 5 passed.ruff format/ruff checkclean.axiom: n/a: microsimulation imputation
🤖 Generated with Claude Code