Skip to content

Make set_input on a branch drop values calculated from the input it replaces - #560

Draft
MaxGhenis wants to merge 22 commits into
masterfrom
branch-set-input-invalidation
Draft

MaxGhenis wants to merge 22 commits into
masterfrom
branch-set-input-invalidation

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A branch starts with every value its parent holds (get_branch → clone → InMemoryStorage.clone), and Holder.get_array serves a branch any key it lacks from its ancestors' keys and default. On master, set_input on a branch stores the new value and drops nothing. So a value the parent calculated before the branch existed answers for the branch, whatever the branch's inputs say, and branch results depend on what was calculated first. For example, the parent calculates income_tax, a branch sets salary, and the branch's income_tax is still the parent's.

This PR makes set_input on a branch drop the values that may have been calculated from the value it replaces. It decides which ones from the order in which values were stored. It also adds Simulation.drop_computed_arrays(), so country packages can stop editing private storage attributes (policyengine-us tools/override_branch.py, and the MTR delete_arrays loops in policyengine-us and policyengine-uk).

Draft: this changes core semantics, and the merge decision is Max's.

How it works

  • Sequence numbers. Every stored array gets a number from one process-wide counter (policyengine_core/data_storage/store_history.py), kept next to the array and copied with it. Storages also flag which values are inputs, meaning values stored through set_input.
  • Store history per simulation. Each simulation records, for each variable and period, the number of the first store its values may have been calculated from. It also records, per variable, the first value uprated or carried over from another period. Three things feed it:
    • its own stores, including values a holder calculates but does not keep, macro-cache reads and spiral defaults;
    • a copy of its parent's history, taken when the branch is created;
    • the history of any simulation its formulas calculate in, taken in each time calculate there returns or raises, fast-cache hits included.
    • the number from which values may have been calculated from anything: the first macro-cache read (its sources were never calculated here) and the values restored from a dump. An input for any variable drops those. This is how a formula that sets an input in a branch it created and calculates there hands back what it needs. A branch calculated from a thread that has no formula context hands its history to every ancestor with a calculation running.
  • The rule. set_input(variable, period, value) on a branch drops each non-input value numbered at or above the earliest recorded store of variable for a period that shares a day with period, or its earliest uprated or carried-over value. It then forgets the records numbered from there on and records again the inputs it kept. If a formula is still running in the branch, it keeps the records, because that formula may hold, in local variables, values it read before the drop. A calculation that was running then (one whose formula sets an input it read, say) is kept neither in storage nor in the macro cache, and neither is anything calculated from it in another simulation: each calculation records, for every other simulation it got a value from (directly or through its callees), that simulation's count of such changes when the value's calculation began, and keeps its result only if none has changed. A call into another simulation that settled before returning leaves its caller clean; unrelated simulations and other threads keep caching. The outermost calculation in the simulation whose input changed (calculate, or a direct calculate_add, whose terms run within it) runs again from the new inputs, inner ones included, until a run changes no input, at most ten times, and keeps that result, so later uprating and carry-over find its period. An input stored for the very period being calculated after the calculation began, under any branch name the simulation reads, is the result.
    • No record means no drop. When a formula creates the branch while still calculating the variable it overrides, nothing it holds was calculated from that variable, so nothing is dropped and the cost is one dict lookup. The exception is a second branch of the same comparison: there, only what the first one handed back, and what was calculated after it, is dropped.
  • Unchanged. set_input on a simulation that is not a branch drops nothing, as before.
  • New public API. Simulation.drop_computed_arrays() drops every non-input value and returns the count. Use it on a branch whose tax-benefit system or parameters change.

Why it is sound

A formula's result is stored after everything it read was stored or recorded, so a calculated value carries a larger number than each value it was calculated from, directly or transitively. Every value a simulation holds was calculated there, inherited from its parent, handed back by another simulation's calculate, or restored from a dump. In the first three cases, what it was calculated from is in the simulation's history, unless it came from a macro-cache read; restored values and values calculated from a macro-cache read are covered by the record that values from a number on may depend on anything. So any value that depends on the overridden variable at an overlapping period was stored after the earliest record of it. Uprating and carry-over read which periods hold values at all, so the first uprated or carried-over value counts too.

A drop at since removes every non-input value numbered since or later. When no formula is running in the simulation, the records numbered since or later therefore describe nothing it still holds, except the inputs it keeps, which are recorded again. When a formula is running, its local variables may still hold such values, so the records stay, and a result whose calculation began before the drop is returned but not cached: it may have been calculated from the old value. A calculation that raises still hands back its history, because whether it raised can depend on what it read. For a value summed or divided from other periods, the record that bounds it may be of those periods; it still overlaps.

The rule can drop more than necessary, but not less, outside these documented limits (docs/usage/branches.md, set_input docstring):

  • Policy changes are not tracked. A branch with a different tax-benefit system or parameters keeps values calculated under its parent's policy. Call drop_computed_arrays() after changing it.
  • In-place writes break the ordering. A formula that writes into an array it read (x += y) changes a value after its number was assigned. policyengine-us#9739 removes the known cases.
  • Inspecting storage is not tracked. Formulas that call get_known_periods or get_array, or read another simulation's storage directly, instead of calculating.
  • Unrelated simulations in other threads. A formula that calculates in a simulation other than its own branches, from a thread it starts without copying its context, does not take in that simulation's history. Its own branches are covered.
  • A branch a formula keeps between calls is a snapshot. Inputs set on its parent afterwards do not reach it. This is the multi-year issue that policyengine-us#9738 handles.
  • Values already read stay read. A formula that read a value from another simulation and then calculates there a formula that changes that simulation's input keeps what it read; its result is returned but not kept. A formula that catches an error raised after such a change returns its own fallback, unkept.
  • Existing child branches keep their values. An input set on a branch does not reach branches already created from it.

Other changes in this PR

All of these came from independent reviews of earlier versions of this branch: nine rounds of soundness review and a code review, run on Subfleet. Each has a test that fails without the fix.

  • Macro cache. A branch whose input dropped values stops reading macro-cache files, which are keyed by branch and period but not by inputs. drop_computed_arrays does too. Before this, the branch re-read its own stale file.
  • requires_computation_after. A prerequisite calculated successfully before a drop removed its values still counts, so there is no new error in branches; a failed request still does not.
  • Custom set_input handlers. Values they calculate are stored as calculated values, not inputs (Holder._set(is_input=...); put_in_cache passes False). Master also recorded them in _user_input_keys.
  • Disk storage writes a new file per store (<key>.<process token>.<number>.npy; a forked child gets its own token). restore takes each key's most recently written file, prefers this process's own on a timestamp tie, and still reads <key>.npy and <key>.<number>.npy names. Recalculating a value no longer changes what child branches, same-named branches in other lineages, or views restored in another process read from the shared directory; the same-name collision also happens on master. Files stay until the directory is removed. Between two other processes' files written in one clock tick, restore cannot tell which came last (documented).
  • Dumps record inputs. dump_simulation writes inputs.txt per variable. restore_simulation restores those values as inputs and every other value as calculated, under one number later than the inputs, and records that values from that number on may depend on anything: a dump does not say what each value was calculated from, nor what was read without being kept. An input for any variable on a branch of the restored simulation drops all of them. Dumps without inputs.txt (master's) restore every value as an input, as before.
  • Dumping a branch dumps the values the branch reads (its own, else its nearest ancestor's or the default). Master read the default branch's, so a branch's dump could not be restored.
  • Custom set_input handlers that calculate between their own stores: if the handler calculated anything (calculate, calculate_add or _calculate), the branch drops again, by the same rule, once it returns or raises.
  • Macro-cache reads also record that what follows may depend on anything, since the cached value's sources were never calculated here, and a branch stops reading macro-cache files as soon as an input is set on it (a file for its name may come from another simulation's branch of the same name).
  • derivative uses drop_computed_arrays(). It now keeps inputs set after the simulation was built; master dropped them, which gave −7 for a true derivative of 2.
  • apply_reform. _invalidate_all_caches keeps inputs by the per-array flag instead of replaying _user_input_keys, and PreservedUserInput is removed (no importers found).
  • Unnumbered values. A value written into storage without a number makes an input for that variable drop everything calculated.
  • Pickles. Storages pickled before numbering still work (__setstate__). Unpickling a storage or history advances the process's counter past the numbers it carries, because counters restart in every process.
  • Merging is incremental. Each history journals its changes since it was created or pruned, so repeated calls into a simulation whose history keeps growing do not re-read every record. A merge notes the source's position before reading, so a record added meanwhile (another thread) is read next time. The read positions are keyed weakly, so a temporary branch's entry goes with it. Histories pickled before journals existed load with empty journals.
  • Subsample. subsample starts a new history before rebuilding.

Invariants and tests

  • Fresh-simulation equivalence (property). Every calculation in a branch equals the same calculation in a new simulation given the branch's inputs first: its own, plus those it inherited when created. This holds whatever was calculated before, in any order.
    • Test file: tests/core/test_branch_input_invalidation_property.py.
    • Programs: Hypothesis generates random programs of calculations, branches (including reused and forgotten names), branch inputs, drop_computed_arrays, and dumps restored as new simulations (including calculate, dump, branch the restored one, override, calculate again).
    • Synthetic system: uprating, a monthly input divided from a year, lagged years, a year summed from months, and a formula that calculates in a branch it deletes.
    • Storage modes: memory, disk, not-kept and blacklisted.
    • Reference: a new simulation per checked value. The inputs a branch should hold are modelled independently, not read back.
    • Runs: CI runs 200 derandomized programs. POLICYENGINE_BRANCH_PROPERTY_EXAMPLES=N explores N random ones; a soak of 3,000 random programs passed.
  • Inputs are never dropped: dataset, situation, ancestor-branch, month-in-year and disk inputs, through set_input, drop_computed_arrays and apply_reform.
  • Precision. A sibling's or another arm's store does not cause drops. In a synthetic itemization-style model, an MTR branch and a reform's baseline arm each calculate shared values once; the previous version of this PR recalculated them three times.
  • Examples: tests/core/test_branch_input_invalidation.py has one per path and per review finding.

Auto carry-over is left out of the random programs. Core's carry-over depends on calculation order in any simulation, branched or not: calculating a later period first stops an earlier input carrying into the years between. That is reported separately and covered here by a controlled-order example.

Mutation check (mutation_check_v3.py, 1,000 random programs per mutant plus all examples, at 4b70aac1): every one of 20 mutants is caught. The fixes from rounds 3–5 are covered by examples that fail on the commit before them. mutation_check_v4.py adds 20 mutants for them (40 in all; mutation_check_v4_116_1000.out, at 116b0181): 39 are caught. The one that survives, numbering restored calculated values before the restored inputs, is equivalent since round 4: the restored simulation records that values from the restored number on may depend on anything, so an input drops every non-input value it holds whatever their order.

Mutant Caught by the random property Caught by the suite
No invalidation (master behaviour) yes yes
No tracking of uprated/carried-over values yes yes
Exact period match instead of overlap yes yes
No history import from simulations a formula calculates in yes yes
No import on fast-cache hits no yes
> instead of >= yes yes
Values not kept go unrecorded yes yes
Blacklisted values go unrecorded yes yes
Last store instead of first sometimes yes
Inputs dropped too yes yes
Fast cache kept after a drop yes yes
Prune without recording kept inputs again yes yes
Prune keeps read positions (no re-merge) no yes
Drop inside _set per stored period yes yes
Macro-cache reads stay on no yes
Merge forgets uprated/carried-over records yes yes
Prune while a formula runs no yes
No ancestor fallback without a formula context no yes
Disk files without the process token no yes
Unpickling does not advance the counter no yes

Full core suite at 94f21309 (local): 1,226 passed, 4 skipped, 1 xfailed. CI on 116b0181: all 18 checks pass, Windows included. Country-template YAML: 39 passed.

Effect on policyengine-us

Enhanced CPS 2024 file, 2026, 3,000-household subsample, real runs on this commit and on master (ms/v2_*). Core master is 7950c01; policyengine-us main is b4e80cf9; #9741 is policyengine-us cc2fdab6.

policyengine-us Core Same outputs in every calculation order? Versus the drop-everything run Median calc time, CPU (3 runs) Branch inputs that dropped anything / arrays dropped (policyengine.py order)
main master no: asking income_tax first changes 8 of 47 output arrays 12 of 47 arrays differ 9.50 s n/a
main this PR yes, bitwise, 3 orders bitwise identical, 47 of 47 9.52 s 6 of 14 / 350
#9741 master yes n/a 9.65 s n/a
#9741 this PR yes, and bitwise identical to master n/a 10.01 s 4 of 8 / 4
  • Main. On policyengine-us main, master's results depend on calculation order. With this PR they are bitwise identical across orders, and identical in all 47 output arrays to a run where every branch drops every inherited array at creation. That brute-force way of making every override effective takes 278 s here. The difference from master is the no_salt CTC-limit override, which the parent's cache was shadowing. On the subsample, this PR on main moves:

    • income tax by −$0.48m;
    • refundable CTC by +$0.48m;
    • income tax before credits by +$167m;
    • itemizers by −260k.

    (Subsample weights scale to national totals.)

  • #9741. policyengine-us#9741 limits the CTC by actual liability and handles branch overrides in policyengine-us. Against it, this PR is bitwise neutral in every order, so releasing core after #9741 changes no numbers. Released before it, core makes the no_salt override effective everywhere. #9741's full-sample "every branch override made effective" row (a brute-force run on core Carry over only values stored at the variable's definition period #557's branch) shows what that does against main in policyengine.py order: income tax −$13.0m, refundable CTC +$13.1m, itemizers −321k.

  • Drops. The 350 arrays on main are 346 in the no_salt branches plus one each in not_itemizing (after the itemizing arm handed back its result), Delaware, Virginia and Idaho. The previous version of this PR dropped 32,822, because a family-wide history counted every earlier store anywhere in the family. On #9741 it drops 4 arrays in every order.

  • Overhead.

    • Per store: about 315 ns per stored array, about 24,000 stores per run.
    • Per call: 39 ns per fast-path calculate (307 vs 268 ns, microbenchmark), about 125,000 calls per run; plus one context-variable set and reset per formula run, about 21,000 per run.
    • Total: that comes to roughly 10–15 ms of a 9.5 s run.
    • Medians: the CPU medians in the table differ by less than this host's run-to-run spread (±0.6 s across repeats of the same configuration). Peak RSS is 3.5 GB for both.

Other packages

  • policyengine-uk branch tests (labour supply response, CPS marriage reform, employer-cost MTR) pass under both master and this PR.

  • policyengine.py sets inputs only on root simulations, so nothing changes there.

  • policyengine-us's own tests for the branch call sites (tests/core, the CTC itemizing-cycle test, and the YAML suites for the CTC, deductions, Delaware, Virginia, Idaho and Missouri TANF), run against core master and this PR (f826d57a):

    policyengine-us Core pytest YAML
    main master 723 passed 1,056 passed
    main this PR 723 passed 1,056 passed
    #9741 master 741 passed 1,056 passed
    #9741 this PR 741 passed 1,056 passed

Composition with other core PRs

Status

  • Reviews. Five rounds of soundness review and one code review on Subfleet found the issues fixed in this branch; each now has a test. Round 3 (on 4b70aac1) found eight issues: an incremental merge that could skip a record added while it read, an obsolete result cached across its own input change, a caught child failure dropping its history, restored dumps losing store order, unreadable f826d57a dump names, timestamp ties, forked processes sharing a token, and old history pickles. 32d5bc43 fixes all eight; each reproduction now matches a fresh simulation. Round 4 (on bff21a79) confirmed those and found one regression of round 3's own (a result not kept across an input change left its period unknown to carry-over: 0 instead of 3) and five gaps that also exist on master: restored dumps without records of uprated, carried-over or unkept reads; a custom input handler calculating between its stores; a macro-cache read's sources; a formula's own input overwritten by its result; and dumping a branch. 2b31dcb7 fixes all six; every reproduction now matches a fresh simulation. Round 5 (on 2b31dcb7) confirmed those, found no regression, and found five more gaps that also exist on master: an override before a branch's first macro-cache read; an input set during a calculation under the default key, or with the formula returning None; a direct calculate_add mixing old and new terms; a custom handler that raises after calculating; and dumping a disk-backed branch whose name contains _. It also showed that a single rerun still let a second input change hide a period from carry-over. 9f8d513c fixes the first four and the rerun (now until no input changes, at most ten times); the _ case is documented (Read only periods the current branch can see when uprating or carrying over #552's key parsing fixes it, verified on the composition). Round 6 (on 9f8d513c) confirmed those and found what that commit broke or left: direct calculate_divide/calculate_add could return an input stored before the calculation (99 instead of 10, a regression of 9f8d513c), the rerun budget multiplied through nested calculations, a handler's private _calculate slipped past the post-handler drop, and a result calculated across its own input change was still written to the macro cache (also on master). ed843df9 fixes all four (an input wins only if stored after the calculation began; only the outermost calculation reruns; only kept results reach the macro cache; _calculate counts). Round 7 (on ed843df9) confirmed those and found that ed843df9's outermost-only rerun let a parent keep a branch's refused result when the branch's formula called back into the parent (2 instead of 6), that a direct calculate_add still gave each term its own reruns, and that a private _set with an earlier number could lose an input. 116b0181 fixes all three: a process-wide count keeps any result calculated across an in-flight drop out of every cache, while reruns stay with the simulation whose input changed; a sum counts as in flight; inputs are always numbered when stored. Round 8 (on 116b0181) confirmed those and found that the process-wide count went too far: an input change in an unrelated simulation left a period uncached here, so a later carry-over returned 0 instead of 6. 359c97fb scopes it: a result is not kept if a calculation it called returned one that was not kept (passed up the call chain, across simulations), or if it is running in this context in the family of the simulation whose input changed. Round 8 also showed a parent formula that reads a temporary branch's value and then calls a branch formula that changes it: the parent keeps what it read (7 where input-first gives 9; master gives 3). That is now a documented limit, and its result is not cached. Round 9 (on 359c97fb) found that rule still leaky and over-wide: a calculation that changed its input and then returned early (a carry-over default) or raised did not taint its caller; a formula that read another family's simulation, a root clone or its own branch from a plain thread kept its old result; and the family-wide refusal lost a settled carry-over period (0 instead of 6, where master gives 6). 94f21309 replaces both mechanisms with one rule: each calculation records what it read from other simulations and when those values' calculations began, hands that to its caller however it ends, and keeps its result only if none of those simulations' inputs has changed since. Every witness from rounds 4–9 now gives the fresh value, or the documented "already read" value without keeping it. A tenth round on that delta is running.

  • Confirmed on 10b1d01d (3,000-household subsample, ms/final_*): policyengine-us main with this PR is bitwise identical to the f826d57a results above, across the three calculation orders, and to the drop-everything run; #9741 with this PR is bitwise identical to master in all three orders; master stays order-dependent (8 of 47 arrays). Drop counts are unchanged (14 branch inputs, 6 dropping, 350 arrays on main; 8, 4, 4 on #9741), so the post-handler drop never fires in policyengine-us. CPU medians on a loaded host (other sessions' jobs running): main 15.7 s master vs 17.5 s this PR, #9741 15.7 s vs 16.9 s, with single runs of each ranging over about 5 s; the per-store and per-call costs measured above remain the better estimate of overhead.

  • Confirmed on 9f8d513c: the subsample outputs are bitwise identical to 10b1d01d's (main in two orders, #9741 in two), with the same drop counts.

  • Full sample (Enhanced CPS 2024 file, 2026, all households; current master b78b0ba9 vs 9f8d513c, ms/full9f8_*):

    policyengine-us Result CPU, master vs this PR Peak RSS
    #9741 bitwise identical to master 117.1 s vs 118.3 s 6.9 GB vs 6.9 GB
    main same outputs in policyengine.py and income-tax-first orders; versus master, 14 of 47 arrays differ: income tax −$13.08m, refundable CTC +$13.08m, income tax before credits +$203m, itemizers −321k 113.4 s vs 107.4 s 7.0 GB vs 7.1 GB

    The main row matches #9741's full-sample brute-force "every branch override made effective" row (−$13.0m, +$13.1m, −321k). Single runs on a shared host; the CPU differences are within run-to-run spread.

  • Confirmed on ed843df9: subsample outputs bitwise identical to 10b1d01d's (main in two orders, #9741 in two; same drop counts), and full-sample outputs bitwise identical to 9f8d513c's on main and #9741 (#9741 still identical to master).

  • Confirmed on 116b0181: subsample and full-sample outputs bitwise identical to the earlier runs (main and #9741; #9741 still identical to master), drop counts unchanged.

  • Confirmed on 359c97fb: subsample outputs bitwise identical to the earlier runs (main 2 orders, #9741 2 orders), drop counts unchanged.

  • Running. The tenth review round, and the policyengine-us confirmation (subsample and full sample) and mutation check of 94f21309.

🤖 Generated with Claude Code

MaxGhenis and others added 10 commits October 1, 2026 14:59
A branch starts with every value its parent holds, and set_input on a
branch used to keep everything calculated from the value it replaced, so a
value the parent calculated before branching answered for the branch
whatever its inputs said.

Every stored array now carries a sequence number from one process-wide
counter (store order), and storages record which values are inputs; both
are copied with clones. A simulation family shares a StoreHistory of the
first store of each variable and period (including values a holder does
not keep) and of the first uprated or carried-over value of each variable.
set_input on a branch drops the branch's non-input values stored at or
after the earliest of those for an overlapping period; when there is none,
nothing is dropped. Simulation.drop_computed_arrays() drops every non-input
value, for branches whose policy changes; _invalidate_all_caches now uses
the same input flags.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Neither is stored, but values calculated from them are, so set_input on a
branch must see them in the store history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ypothesis

The country-package smoke job installs no dev dependencies, so the
module-level hypothesis import failed its collection. The synthetic system
the examples and the property share moves to
tests/fixtures/branch_input_invalidation.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each simulation now keeps its own store history: a branch copies its
parent's, and a formula that calculates in another simulation takes in that
simulation's history when calculate returns (fast-cache hits included). A
drop prunes the records it made obsolete and records the kept inputs again.
An unrelated store in a sibling branch, or in the other arm of a
reform/baseline pair, no longer lowers a branch's threshold, so MTR-style and
baseline branches calculate shared values once.

Also from the soundness and code reviews:
- a branch whose input dropped values stops reading macro-cache files (they
  are keyed by branch and period, not inputs); drop_computed_arrays too;
- requires_computation_after counts a prerequisite requested before a drop;
- values a custom set_input handler calculates are not inputs;
- disk storage writes a new file per store, so a recalculation no longer
  changes what child branches or same-named branches read;
- a value written into storage without a number counts as a dependency;
- storages pickled before numbering still store (__setstate__);
- derivative keeps inputs set after the simulation was built;
- property test: an independent model of divided inputs, reused and
  forgotten branch names, a blacklist mode, derandomized CI examples with an
  environment soak; examples for each finding;
- docs and docstrings say what the code does, with the remaining limits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From the second soundness review (in progress):
- A drop while a formula is still running in the simulation (its own
  set_input, or drop_computed_arrays) no longer prunes the history: the
  formula may hold values it read before the drop and store a result from
  them. Calculations in flight are counted per simulation; a clone starts
  at zero.
- Merging is incremental: each history keeps a journal of its changes since
  it was created or pruned, so repeated calls into a simulation whose history
  keeps growing no longer re-read every record (quadratic before).
- The merge map is keyed weakly, so temporary branches' entries go with
  them, and is left out of pickles (weak references do not pickle).
- Unpickling a storage or history advances the process's sequence counter
  past the numbers it carries; disk restore keeps each key's most recently
  written file (numbers restart in every process).
- Tests: drops inside a running formula, an uprated value handed back by a
  temporary branch (also in the property's scenarios), cross-process
  unpickling, restore order. Docs: the in-flight rule and the thread limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… success

From the second soundness review:
- A branch calculated from a thread with no formula context (a worker a
  formula starts without copying its context) hands its store history to
  every ancestor with a calculation running, so the waiting formula's result
  keeps the dependency. Unrelated simulations are not touched.
- Disk files carry a per-process token as well as the sequence number:
  numbers restart in each process, so a later process could otherwise
  overwrite a file a restored view still maps.
- requires_computation_after counts a prerequisite only once its
  calculation succeeded.
- Docs and comments: the soundness argument states records for overlapping
  periods (summed and divided values), disk files stay until the directory is
  removed, the thread limit covers only unrelated simulations, and the
  history is copied into branches, not shared.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dation

Conflicts kept both sides: InMemoryStorage keeps #556's shared-array clone
and also copies the sequence numbers and input flags; drop_computed also
discards dropped keys from _shared; get_branch's docstring keeps both
paragraphs.

Two #556 tests encoded the shadowing this branch fixes:
- test_set_input_on_branch_leaves_parent now expects the branch's
  income_tax to follow its new salary (the parent stays unchanged);
- test_branch_copies_only_what_it_reads accepts that the input drops,
  rather than copies, values calculated after salary was first stored.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sk names

From the round-3 soundness review:
- A calculation still running when an input set on its simulation dropped
  values returns its result but does not cache it (an input epoch), so the
  next calculation uses the new input.
- A calculation that raises still hands its store history to the calling
  formula: whether it fails can depend on what it read.
- A merge records how far it read the source before reading, so a record
  added meanwhile (another thread) is read next time.
- dump_simulation records each variable's input periods (inputs.txt);
  restore_simulation restores inputs as inputs, then every other value under
  one later number, so a branch input drops what it may have fed. Older dumps
  restore every value as an input, as before.
- Disk restore reads the previous `<key>.<number>.npy` names, prefers this
  process's file on a timestamp tie, and a forked child gets its own token.
- StoreHistory pickles from before journals load with empty journals.
- Tests for each; docs updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The pickle test merged a history that died at once, so the weak map was
empty when pickled. Keep the source alive so the test covers the case the
__getstate__ exists for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A process pool forked from a multi-threaded process (the test process can
have threads by then) can deadlock; the child now only writes its token to
a pipe and exits. Still fails without the at-fork hook.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ro reads

From the round-4 soundness review (on bff21a7):
- A result refused because an input changed while it was calculated left
  its period unknown, so later carry-over found nothing (0 instead of 3).
  calculate now runs such a calculation once more from the new inputs and
  keeps that result; a second change still returns without keeping.
- An input a formula sets for the very period it calculates is the result
  and is no longer overwritten by the formula's return value.
- A dump does not say what each value was calculated from, nor what was read
  without being kept, and a macro-cache value's sources were never
  calculated here: the store history now records the number from which
  values may depend on anything (restored values, the first macro-cache
  read), so an input for any variable on a branch drops them.
- A custom set_input handler that calculates between its own stores: the
  branch drops again, by the same rule, once the handler returns.
- dump_simulation on a branch dumps the values the branch reads (its own and
  inherited) instead of the default branch's, which could not be restored.
- calculate ends the calculation (in-flight count, tracer) even if handing
  its history back fails.
- The property test also dumps and restores simulations, including a chunk
  that calculates, dumps, branches the restored simulation and overrides an
  input; it catches a history without the new record.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 2, 2026
Follow-up to the review of the first commit:

- Holder.delete_arrays no longer looks through the whole record. It
  compares the periods memory stores for the branch before and after the
  deletion and discards exactly those entries, so its cost does not grow
  with the record. Country marginal-rate code deletes every variable on a
  branch: with 9,000 entries and 3,024 variables the loop took 0.73 s with
  the scan and takes 0.018 s now (0.004 s without any pruning). A holder
  with disk storage, which cannot list its periods for every branch name,
  still looks through the record for the variable's entries in the deleted
  periods and drops those whose value neither storage holds.
- Holder._set records the period as storage keys the value: eternity for an
  eternal variable whatever period it was set for, and a Period for a
  handler that passes a string. Each entry names one stored value, so a
  string-period input is exported and deleting an eternal input drops its
  entry.
- put_in_cache stores with is_input=False (the same keyword as #560), so
  values a custom set_input handler calculates are formula results, which
  apply_reform recalculates, not inputs.
- subsample starts the record again before it rebuilds the simulation, so it
  records only what the rebuild stores.

Tests: a guard that a memory-only deletion never iterates the record; the
eternal, string-period, calculating-handler, disk and subsample cases; the
property's model now includes a handler that calculates and stores months
under string periods.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 7 commits October 2, 2026 02:40
… earlier files

It faked a later process by restarting the counter only, so with the same
token its first store could reuse the writer's first file name; and it aged
only the writer's last file, so on a coarse file clock (Windows) the other
four could tie with the later write. Passes with every write given the same
timestamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ted)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They found the cause: test_disk_restore_reads_the_latest_file_of_each_key
failed on Windows' coarse file clock (fixed in acc5f58), and its rerun by
pytest-rerunfailures dropped the module's setup-state entry without running
its finalizers, so the module-scoped tax_benefit_system leaked into later
modules; test_branch_shared_arrays traced it, and test_parameters and
test_reforms then got a traced parameter tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… handlers

From the round-5 soundness review (on 2b31dcb); all were also on master:
- A branch stops reading macro-cache files as soon as an input is set on
  it, not only when the input drops something: a file for its name may come
  from another simulation's branch of the same name, read before any record
  exists.
- An input set during a calculation for its own period wins under any
  branch name the simulation reads (Holder.set_input stores under
  "default"), and also when the formula returns None and core would fall
  back to the default for want of an earlier period.
- A direct calculate_add sums again when a term changed an input while
  summing (an earlier term could be obsolete).
- A custom set_input handler that calculated and then raised: the branch
  still drops again (the inputs it stored stay). The drop after a handler
  now runs only if the handler calculated anything.
- calculate runs a calculation again until a run changes no input, at most
  ten times (was once), so a formula changing a second input on its rerun
  no longer leaves the period unknown to carry-over.
- Docs: the above, and that disk-backed branches whose names contain "_"
  cannot be dumped (OnDiskStorage key parsing, as on master; #552 fixes it).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ne rerun budget

From the round-6 soundness review (on 9f8d513):
- Direct calculate_divide/calculate_add could return an input stored
  before the calculation (they do not look in storage first), after any
  unrelated input change: 99 instead of 10, 100 instead of 12. An input
  now wins only if its sequence number is after the calculation began.
- Only the outermost calculation in a simulation (calculate or a direct
  calculate_add) runs again after an input change; inner ones run with it.
  The ten-rerun budget no longer multiplies through nesting (11/121/1,331
  leaf runs at depths 1/2/3, now 11 each).
- A result calculated across an input change is no longer written to the
  macro cache (only kept results are), so a branch of the same name
  elsewhere cannot read it (also on master).
- The post-handler drop counts _calculate, not only calculate, as
  calculating (a handler's private carry-over calculation was missed).
- Docs: the above, and the rerun cap's consequence for carry-over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imulations

From the round-7 soundness review (on ed843df):
- A parent formula calling into a branch whose formula calls back into the
  parent and then changes the branch's input: the parent kept the branch's
  refused inner result (2 instead of 6; also through calculate_add and
  calculate_divide). Drops made while calculations run now also bump a
  process-wide count, and no result calculated across one, in any
  simulation, is kept. Which calculation runs again is still decided per
  simulation: the outermost one in the simulation whose input changed, so a
  temporary branch settles locally and nesting does not multiply reruns.
- A direct calculate_add counts as in flight while it sums, so its terms do
  not each get their own reruns (121 runs, now 11).
- Inputs are numbered when stored, even if the caller passes a number
  (a private _set with an earlier number lost its input to "input wins").
- Tests for each; docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… family

From the round-8 soundness review (on 116b018): refusing to cache any
result calculated across an in-flight drop anywhere in the process let an
unrelated simulation's input change leave a period uncached, so a later
carry-over returned 0 instead of 6 (master, 9f8d513 and ed843df give 6).

A result calculated across an input change is now kept out of caches only
where it can matter:
- in the simulation whose input changed (its own epoch, as before);
- in any calculation that received a refused result, directly or through
  other calculations, in any simulation (a per-frame flag in
  _calculation_frames, passed to the caller when a frame returns);
- in calculations of the same family still running in this context, which
  may hold values read before the drop (their cache epoch is bumped).
Unrelated simulations and other threads keep caching. Reruns stay with the
outermost calculation in the simulation whose input changed; calculate_add
and calculate_divide run in frames too.

Docs: the scope, and that values a formula already read from a branch
before a formula there changed the branch's input stay read (its result is
returned but not kept). Tests isolate each part (taint across families,
the family bump, the unrelated family); new mutants M38/M39 are caught.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From the round-9 soundness review (on 359c97f):
- A calculation that changed its simulation's input and then left before
  any cache decision (a carry-over default returned early, an exception)
  did not tell its caller, which kept an obsolete result (0 instead of 6).
- A formula that read another simulation (another family, a root clone, its
  own branch from a plain thread) and then calculated there something that
  changed that simulation's input kept its old result (2 on the second
  request instead of 6).
- The family bump refused a settled, correct result, so a later carry-over
  found no period (0 instead of 6; master gives 6).

One rule replaces the taint flag and the family bump: each calculation frame
records, for every other simulation it got a value from (directly or
through its callees), that simulation's _input_epoch when the value's
calculation began; it keeps its result only if neither its own simulation's
epoch nor any of those has changed. A callee hands its reads to its caller
however it ends (value, early default, error), its own simulation counted
at its start; a rerun that settles restarts the count, so a settled call
leaves its caller clean. Fast-cache hits from another simulation are reads
too.

Docs updated ("values already read stay read" now covers any simulation and
caught errors). Tests for each witness; mutants M39 (no reads) and M41 (a
read counted from return, not start) are caught.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant