Skip to content

Integrate main into the graph-native US line (native restart 2026-09-23) - #1011

Draft
MaxGhenis wants to merge 633 commits into
mainfrom
native-integration-20260923
Draft

MaxGhenis wants to merge 633 commits into
mainfrom
native-integration-20260923

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Integrates the graph-native US development line onto main 6442030a2, preserving the September 21 descendant branches and the retention-seal implementation. Includes the comparison-only poverty guidance without changing release certification or publishing data.

Supersedes the review stack #893 (native US integration), #935 (verification reuse), #945 (scale transport), #949 (row ceilings), and #950 (content seals). All five recorded branch tips are ancestors of this branch. Part of #956.

The integration record documents all 26 prepared main conflict resolutions and the one additional conflict from newer main. It also records the retained test fixtures, facade export unions, shared fiscal compilation, adapter source ownership, and Medicaid substitution behavior.

Validation (the lane's runners, 123 selected files): 120 pass on their latest attempt. These include the survey-population (73), survey-puf55 (47) and other-disability-host (24/24 in two isolated runs) retries after the implementation-inventory repair. Two files still fail and will be fixed in follow-up commits on this branch:

  • test_us_graph_us_survey_enrichment.py: 9 setup errors, DEMOGRAPHIC_SOURCE_CONTRACT. A demographic source contract moved with main's ASEC person-column changes.
  • test_us_graph_immigration_enrichment.py: 1 failure. The refusal is now DETAIL_ARTIFACT_IDENTITY where the test expects ARTIFACT_PRODUCER_KEY.

The merge commit e8dfb41 records three parents: the native line, main 6442030 and poverty-comparison-only 07b69be. The lane applied the last two by content while its sandbox could not write git. The tree was checked against both of those sides.

Enrichment and disability continuation now distinguish retained-parent-to-cache comparisons from reconstructed-values-to-current-manifest comparisons. All 24 nullable ContentStore context regressions pass without relaxing mutation checks.

The integration record explains reviewed source qualification and fixture boundaries, exact field/hash derivations, and the static scanner correction. No actual-data native run, release certification, or publication was performed.

This PR must remain a draft.

🤖 Generated with Claude Code

MaxGhenis and others added 30 commits September 17, 2026 18:02
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rd it

survey_population_preparation._spill_roster gained a Path.read_bytes on
native-scale-transport at b6081ef, which added a resource_accesses entry and
left graph_implementation_inventory.json's declared
resource_accesses_sha256 stale. implementation_manifest therefore REFUSED
authenticated_survey_population_v1 -- the stage whose kernels' own
implementation_hash calls it -- at every commit from b6081ef to the branch
tip a64f7b7, so no 19-node graph run was possible there. The transport
report's section 4 'No committed pin moves' was recomputed before that commit
and is stale at its own tip.

The repository had no test over the whole inventory: the one caller built
three of the ten stages. test_us_implementation_inventory_contracts.py builds
all ten and checks all 124 declared contracts; reverting the re-pin turns both
arms red, which was verified before committing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
graph._bounded_json opens with

  _require(type(limit) is int and 0 < limit <= 64 * 1024**2, "TRANSPORT_LIMIT")

so the shared encoder refuses any cap above 64 MiB before it encodes a byte. A
larger MAX_PAYLOAD_BYTES in survey_origin_budget would not loosen the bound; it
would refuse the module. That turns "this takes the transport argument" from a
preference into a structural fact, and the test now drives it.

Also names, in section 5a, the test the rule applies to a bound that looks like
a row ceiling -- does the binding site stream, or materialise a per-row payload
-- and the three bounds it decides: current_survey_geography streams (moved),
current_survey_household_roles serialises its whole per-person table to JSON
under a 64 MiB artifact cap, and current_child_property_income_source does one
json.dumps under a 64 MiB projection cap. Both of those are the transport's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
asec_demographic_source._MAX_PERSONS was 600,000 over two rosters: a fixed ASEC
three-cohort 432,523 that no fraction grows, and the retained ACS persons, which
a full-source build grows to 3,422,888. One constant over two rosters takes the
larger, so the rule produces 14,000,000.

Three things had to hold and each was read. It is not a fixed-width encoding:
rows is a JSON header integer, the only struct.pack is "<I" over the header
length that _HEADER_MAX bounds separately, and the body budget is derived and
checked against actual bytes rather than capped. It is not an assertion about
the ACS file's size and could not be -- at 600,000 it sat 5.7x BELOW the genuine
3,422,888-row file, and a bound that would refuse the real file is not a claim
about it; ACS_SOURCE_ROW_SHAPE and ACS_CAPTURE_CHANGED make that claim exactly,
and DEMOGRAPHIC_COHORT_ROWS makes the ASEC one, so neither weakens. And no
per-row payload sits behind :1016, which takes two numpy arrays.

Boundary test drives ACS_ROWS at acs_household_reference_states. No pin moves:
the module is not inventoried and carries no external digest pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The base financial runner now names its declared consumers before the run --
financial.ATTACH_NODE, the property graph's attach node, and the tax gate when
the rebase is enabled -- detaches those three with the executor's own snapshot
function so object identity is exactly what it is today, and seals every other
observation on arrival and drops it. Both same_replayed_population call sites
become seal comparisons with the same codes.

_node_population_stamp takes either arm: a retained Population is stamped as
before, a sealed node re-derives seal_identity from its retained record, so
FINANCIAL_NODE_POPULATION_CHANGED stays non-vacuous on both. The completion
host keeps today's behaviour untouched: it runs its own observer, retains
everything, and passes no seal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both are byte transports, so neither is this lane's to move. Both are measured
and pinned in a test, because a reader who meets them in a build has lost hours.

ACS coverage authentication charges every selected row 6 * len(raw) + 1024
against MAX_BODY_BYTES before the reader allocates, refusing
SELECTED_BODY_BUDGET. Measured over 200,000 real records of the pilot's
captured public ACS PUMS archive: a person record averages 695.57 bytes, so the
charge is 5,197 bytes and 64 MiB admits 12,911 selected persons -- 0.38% of
source, 265x under at full source. It is the tightest ceiling censused, and it
refuses 265x before acs_person_coverage_columns.MAX_SELECTED_ROWS, which this
lane lifted.

And the preparation-receipt ceiling the transport lane reported as moved from
96,860 households to 6,206,000 is still enforced one module downstream:
_roster_payload returns one joined payload under MAX_ROSTER_BYTES = 4 GiB, and
graph_survey_population._checked_preparation checks those same bytes against
PREPARATION_MAX_BYTES, still 64 MiB, refusing PREPARATION_BYTES. From the
transport lane's own committed receipt: a 1/10 roster measured 109,804,304
bytes and was recorded accepted by the producer, 1.64x this cap; a full-source
roster measured 1,099,892,722 bytes, of which this cap admits 96,839
households. That is the number the transport lane lifted, still standing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
census.json: 146 bounds across five families, each with its enforcement site,
refusal code and exception type, what it protects, its counts at 1/10 and full
source, and whether it binds -- plus the 41 adversarial verdicts, none of which
overturned a headline binding call. Fourteen bind at full source, six at 1/10.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Test summaries, PR URL and the questions for Max follow once the final clean
run lands; those sections are added, not rewritten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ineteen

The census agents read a moving tree: five constants were lifted while it ran
and the agents correctly reported the post-lift values, recording both numbers
per row. At base nineteen bind; seven moved here; twelve remain for the
transport argument. Also corrects 41 verdicts to 42.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A probe fleet over the comparison surface found eleven discriminations the
first battery missed, all of the same shape: the pair is byte-identical and the
comparison refuses anyway, because it is asserting something about one operand.
A seal cannot carry those as content, so it must assert them when it is built,
and the agreement driver proves it does -- a MassChangeRecord or MassRecord
subclass, a Population subclass, a pyarrow or sparse masked carrier, a uint8
mask, an ndarray-subclass backing, a numpy str_ cell, metadata key order,
object NaN sign, and non-finite metadata, which the store codec spells in hex
rather than refusing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e code does

I wrote in the design note that the roles projection costs "roughly 200 bytes
of JSON per person" and admits "a few hundred thousand". That was reasoned, not
computed, and this lane's own standard forbids it. Measured through the module's
own encoder: 320.08 bytes per person row, so the 64 MiB artifact cap admits
209,661 persons -- 5.88% of source, against a row bound that admits 58.8%. The
byte cap refuses ten times earlier, which is exactly why MAX_PERSONS is not this
rule's to move.

Also two precision fixes. The financial graph DOES call into
survey_origin_budget -- for _config_payload and the _live() producer seal -- so
"no node executes it" becomes "no node executes freeze_survey_origin_budget,
where GROUP_COUNT_BOUND is checked, and neither of those calls reaches
_initial". And MAX_SELECTED_ROWS at 14,000,000 now sits above MAX_ROWS at
6,000,000 in the same module, which is not an inconsistency -- one bounds
caller-supplied keys, the other the archive's own rows -- but does mean the
effective ceiling there is 6,000,000 by containment. Both said in the note
rather than left for a reviewer to find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
399 passed over the seven touched test files at a fixed commit with the tree
untouched for the run; ci_test_groups --verify ok; the stdlib matrix contract's
15 tests OK; spec-engine coverage 42156/42156 and 41/41; ruff clean and 21 files
already formatted; all three pin generators idempotent.

Records honestly that an earlier full-file run reported 8 failed while this
session was editing the tree underneath it, and that it measured nothing -- the
same tests pass at a fixed commit. And records the base-worktree run that proves
the inherited pin break is the base's: 7 failed there with the exact
Unclassified-contract ValueError, all 7 pass here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR #949, draft against native-scale-transport, MERGEABLE. CI does not run on it
by design: test.yml triggers on pull_request: branches: [main], and gh pr checks
reports none, so section 8's local gates are the only ones this branch has.

Section 11 says what the report does not claim -- including that census.json's
per-row prose is the agents', checked but not rewritten, so a row may carry a
superseded classification beside its verified binding call; and that the roles
projection figure is on a faithfully shaped invented table, not a real artifact.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 1/1000 launcher runs the transport lane's committed replay harness
unedited, byte-identical, at NEED_GB 40. The 1/15 harness is that lane's
committed tenth harness with exactly three changes -- the fraction, the RSS
ceiling the retention brief names, and the label -- verified by diffing the
bodies.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dangerous one first: an object axis sealed through pd.Series(index.array)
lost its null-sentinel identity, because pandas 3 re-infers that Series to str
and maps both None and pd.NA to nan -- so the seal ACCEPTED a pair
Index.identical refuses. The axis now folds np.asarray(index.array).

Four code divergences, each pinned: array_equivalent is byte-tolerant for
complex and bool as well as float, so the value fold collapses all three and
the byte difference keeps reaching NATIVE_BITS; Index.identical compares
dtypes with == only, so the dtype CLASS check belongs to the series code and
not to AXIS; _array_bytes_equal does not raise, so a structured dtype with an
object field must refuse under its caller's code; and the masked canonical
flag is np.all(data[mask] == 0), which an empty string fails and truthiness
does not. Also: _comparables are compared element-wise with ==, as identical
does, rather than as a tuple that short-circuits on identity.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by test rather than by reading: graph_survey_completion_host:814-816
uses the base run's whole per-node population roster as its own expected
populations and hands them to _states, which reads each frame's tables. The
recursive base call now sets a private retention flag, so that path retains
exactly what it retains today, and the design note says where the saving does
not apply.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A declared consumer is sealed after the snapshot detaches it -- which is the
object the comparison ran against before this change -- and every other node is
sealed live, with no copy taken at all. One seal per node instead of two, and
the property that makes the two equivalent is now a test rather than a runtime
double-seal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…agnostic

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The root journal is cumulative: a lane appends a section above the rule and
leaves everything below it. This lane's first commit replaced the whole file,
deleting 1,623 lines of four earlier lanes' history. Restored verbatim from
a64f7b7, with this lane's section above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
108 -> 112 tests. The battery now records each comparison's verdict on both
paths when MICROCOSM_BATTERY_RECEIPT is set, and the committed receipt shows
the tally: every comparison agrees. Two codes are unreachable and the receipt
says why -- FRAME_TYPE because Population validates its frame, STRING_POLICY
because StringDtype equality already compares storage and na_value. The strata
name is only reachable after construction, because Frame normalises it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 1/15 harness takes the same per-file F401 ignore the native-scale lane's
three harnesses take, for the same reason: it is committed byte-identical to
what ran and its imports feed the stray-module assertion. The two probe scripts
were reformatted and RE-RUN, so the committed copy is the one that produced the
output the report quotes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Counted off the module rather than off the table. The battery reaches twenty of
them plus eleven accepted pairs; the two it cannot reach are named with the
reason each is unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Twenty-one files import a module this lane touched. One pytest process per
file, bounded parallelism: the native-scale lane's section 9a is the write-up
of what happens when they share one process, and this avoids it by
construction rather than by explaining it afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…iagnostic

Round 1 named a second catch-all. Round 2 -- patching the other catch-alls --
refused in 3 CPU-s because acs_native_coverage_binding pins the sha256 of four
ACS module files, which is itself worth recording: those modules cannot be
instrumented. Round 3 observes the raise with sys.monitoring instead, changing
no byte of the measured tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MaxGhenis and others added 14 commits September 21, 2026 15:49
…egration line

Brings the AGI-tail pure selection contract, its review fixes, the
structural expansion contract and compatible matching (b2d7a57,
43c39a9, 5113093, 1b35534).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the native integration line

Brings the engine-free survey handoff (d053f39), its test constructor
fix (c1ae4f2) and operation-local child reconstruction reuse with
recheck on close (050c58a). Same tip as
native-child-verification-targeted-counter-20260921.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e integration line

Brings the explicit survey financial path before atomic geography
(48ce525) and the household-domain qualification with retained source
custody (5067f21, patch-identical to 292c268).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…native integration line

292c268 is patch-identical to 5067f21, already merged, so this merge
records ancestry only and changes no file content.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ntegration line

Brings the development singleton Social Security beneficiary convention
(fa72c60): a descriptive pure core, with no live source, engine, graph
or clone-binding authority.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the native integration line

Brings the optional other-disability completion graph (a5fc3e9,
patch-identical to 0e073d1), the opt-in host continuation (1623488),
filter-frame materialization before population reconstruction
(0184355) and the inherited manifest population-context fix
(aa2eaab).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…native integration line

0e073d1 is patch-identical to a5fc3e9, already merged through
aa2eaab, so this merge records ancestry only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on line

Conflicts (2 files), resolved so both retention answers survive, separated
by profile:

- pyproject.toml: union of the per-file-ignores blocks (row-ceilings and
  byte-transports from this line, the retention-seal harnesses from #950).
- graph_atomic_survey_financial.py (12 hunks):
  - _population_retention="compact" (ee70493) is unchanged: witness every
    node, retain the compact roster, keep the executor's detachment.
  - _population_retention="all" (the default) adopts #950: seal every node on
    arrival, retain only the declared consumers (every node when
    _retain_every_node_population), detach them itself and pass
    _population_observer_detach=False; replay comparisons use the seals and
    node_populations carries the seal record for each dropped node.
  - Declared consumers come from _compact_retained_roster (financial attach,
    development attach when present, property attach, tax gate): a superset
    of #950's three that adds the development attach node this line makes
    final_node.
  - _retain_every_node_population is accepted by _run_survey_financial and
    both public wrappers; the recursive completion base call sets it exactly
    when the profile is "all"; the flag check also refuses the compact
    profile.
  - docs/us-native-retention-seal.md gains an integration note saying so.

Targeted tests on the resolved tree: test_graph_executor,
test_us_survey_population_replay, test_us_implementation_inventory_contracts,
test_us_graph_pre_geography_survey_financial and
test_us_graph_survey_completion pass. test_us_graph_atomic_survey_financial
has one failure, test_final_owner_return_cannot_mutate_materialized_geography
("DID NOT RAISE"): its profiler waits for a caller that is
run_atomic_survey_financial, which became a thin wrapper in 48ce525, so the
mutation never fires. That predates this merge and is fixed in a follow-up
commit. Completion-host and compact-retention files were still running.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lane's code (through 9af56aa, the tree the 1/15 run used) is already an
ancestor of this line. This brings the two remaining experiments-only
commits: the launch scripts' unclosed-quote fix (703c9d5) and the 1/15
launch record with its pre-registered prediction (25d66f3). No conflicts;
no source or test file changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion line

Ancestry only: the merged tree equals this line's tree (0025f01). The
commit ("Return verified graph-store byte buffers") was already cherry-picked
into native-release-integration-20260919 by subject; this records the
original so the branch is contained rather than duplicated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
48ce525 made run_atomic_survey_financial a thin wrapper over
_run_survey_financial, which is where verify_materialized_current_survey_
predictors returns and where ``result`` lives. The test's profiler still
waited for the wrapper as the caller, so the mutation never fired and the
test failed with "DID NOT RAISE". Wait for the shared runner instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ative integration line

Integrates main through 6442030 onto the graph-native US line (epic #956):
the 26 conflict resolutions against d9a248d, then the one further conflict
from newer main (test_uk_calibration_run.py: both fixtures kept), then the
clean merge of poverty-comparison-only-20260921. docs/native-integration-20260923.md
records every resolution and the identity repairs. The integrate lane
applied the second and third merges by content while its sandbox could not
write git, so this commit records all three parents explicitly. Its tree
was checked against both later sides: every file they changed matches
their version unless the native line also changed it. main's retirement
of uk/cgt_source_stages.json (821c134) wins over the poverty branch's
older copy.

Validation (lane runners, 123 selected files): 120 pass on their latest
attempt, including the survey-population (73), survey-puf55 (47) and
other-disability-host (24/24 in two isolated runs) retries. Still failing,
and to be fixed in follow-up commits:
- test_us_graph_us_survey_enrichment.py: 9 setup errors,
  DEMOGRAPHIC_SOURCE_CONTRACT (a demographic source contract moved with
  main's ASEC changes);
- test_us_graph_immigration_enrichment.py: 1 failure, the refusal is now
  DETAIL_ARTIFACT_IDENTITY where the test expects ARTIFACT_PRODUCER_KEY.
The other-disability-host errors seen once under concurrent runs did not
reproduce in isolation.

No actual-data native run, release certification or publication.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 15 commits September 26, 2026 07:37
…on line

test_us_graph_us_survey_enrichment.py: the module fixture's hours helper
rewrites every ASEC person member but re-pinned only the coverage and
restoration owners, so the demographic owner refused the final bytes
(DEMOGRAPHIC_SOURCE_CONTRACT wrapping STUDENT_SOURCE_SIZE). This predates
main: it reproduces at c15661e and at ac15e9f, the native commit that
added the helper. The fixture now re-pins the demographic owner too, and a
new fixture test checks the final member pins against the bytes. Two
assertions enumerated every declared amount group; they now read the
executed default groups from the run receipt and require the opt-in routes
to stay absent and missing.

test_us_graph_immigration_enrichment.py: replacing only producer_key always
refused in the shared typed-edge identity check (DETAIL_ARTIFACT_IDENTITY)
before any producer comparison. The control now asserts each refusal
separately: producer-key-only mutation refuses DETAIL_ARTIFACT_IDENTITY, and
a self-consistent foreign identity refuses
CURRENT_SURVEY_AMOUNTS_ARTIFACT_PRODUCER_KEY.

No production module, source pin, implementation inventory or F0 identity
changed. docs/native-integration-20260923.md records both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings in the 19 main merges since 6442030: the Route A remediation stack
(#1016, #1018, #1017, #1025, #1024, #1028), #1029, #1005, #1008, #1004, #992,
#1031, #1015, #1033, #994, #954, #966, #1006 and #1010.

Eight conflicted paths, each recorded in docs/native-integration-20260923.md:
- test.yml keeps the native sharded matrix (main's --durations=25 is already
  in every native pytest call);
- the us_runtime facade stays lazy and gains main's three fiscal-target
  exclusion exports (the parent union, 920 names; the facade-union test is
  re-pinned to that union's digest);
- reform_validation.py keeps main's _released_engine_state and the native
  explicit-constructor default_simulate_factory;
- build_us_fiscal_refresh_release.py carries the native explicit consumer
  seams (formula metadata, dataset and microsimulation constructors, SPM
  selection) onto main's household-batched post-export scorer and batched
  base materialization; omitted seams keep main's exact calls, and _main
  supplies none. Three native test files that addressed the removed
  frame-based factory now address the scorer with the same assertions;
- identity pins observed on the merged tree: five EXPECTED_HASHES entries,
  the regenerated F0 coverage report (42,239/42,239 fields, 41/41 checks),
  the US spec identity 2dfa51b8... and the loader golden f2047cb9...

tools/generate_us_bundle_from_constants.py --check, tools/ci_test_groups.py
--verify, the CI matrix contract and ruff pass. The fiscal consumer,
formula-metadata, shared target/solve and calibration-attachment files pass;
the rest of the fiscal battery and the known pre-existing failures are
follow-up work on this branch.

No actual-data native run, release certification or publication.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tools/spec_seed_identity_diagnostics.py pinned LOCK_SHA256 95e8297d..., the
digest of the native line's uv.lock at e47adc6. The lock has since moved:
main's policyengine-uk 2.100.0 relock (26c449d) plus the native
source-io extra give the current uv.lock 345e014a..., which
us_runtime/worker_identity.py already binds. The earlier integration
re-pinned worker_identity but not this diagnostic, so
test_checked_in_lock_passes_actual_source_stamps refused with LOCK on every
lane that runs it (engine-shared, fast rest 6/6, wheels 2/8) and the
spec-seed-diagnostics job failed. The pin is the observed digest of the
committed uv.lock; the refusal test for a wrong pin still passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two native test observers ended with sys.monitoring.clear_tool_id, which
exists only from Python 3.14; the workspace floor is 3.13. On the 3.13 lanes
the AttributeError fired before free_tool_id, so the tool id stayed claimed:
test_us_child_property_verification_operation.py failed directly and every
later test_us_current_survey_household_domains.py observer refused with
RETURN_OBSERVER_TOOL_OCCUPIED (fast 3.13 rest 6/6, wheels 3.13 3/8).

Both now release the way test_us_graph_survey_puf55.py already does: zero
each local event mask, unregister the observer's one callback, check that no
global or local events remain, then free the id. Every ownership, thread and
profile assertion is unchanged. A standalone script replaying both
sequences on CPython 3.13.15 (no clear_tool_id) and 3.14.4 fires the
callbacks only while installed and reclaims both ids afterwards. The genuine
fixtures behind these tests need the real disk probe's ~17 GiB free, which
this machine lacks, so the 3.13 CI lanes verify the tests themselves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The native line moved the hours gate body into
us_hours_worked_gate_from_summary, which required the summary to count all
three output columns. Main's #941 (microcosm#765) lets the ACS/pool surface,
which drops weeks_worked, call us_hours_worked_signal_gate with
required_columns=US_HOURS_WORKED_POOL_OUTPUT_COLUMNS, and its summary counts
only the columns present. Combined, every pool-scoped call refused with
"requires the complete original summary": two test_us_hours_worked.py gate
tests and five test_us_acs_local_release_tool.py finalize tests.

us_hours_worked_gate_from_summary now takes the same required_columns, with
the complete three-column default, and us_hours_worked_signal_gate forwards
its scope. A summary is complete when the required columns are counted and
every counted column is a declared output. The constant-column check still
covers every counted column in declared order, and the bands and value
checks are unchanged. The native pool callers gate the full three-column
surface before dropping weeks_worked, so they keep the strict default.

Invariant, property-tested with hypothesis: for any summary, the scope only
decides which counted columns must be present. Pass/fail and failures are
identical under every scope the summary satisfies. The default refuses any
summary without weeks_worked, and each constant counted column fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Native commit 071f563 bound the immigration stage's undocumented-worker
spill to ASEC labor-force status: A_LFSR joined
US_IMMIGRATION_REQUIRED_SOURCE_COLUMNS, and the immigration_status stage now
refuses a person table without it. Its source-stage notes say A_LFSR is
"restored from the same pinned official ASEC archives by exact PERIDNUM
join". On the legacy base stage nothing did that, because none of the three
census_cps H5 inputs carries A_LFSR. test_us_plan's provider-closure test
caught this on every lane. On real inputs, main's base-stage build would
have refused at immigration_status once this line merged.

A_LFSR now joins main's #720 Census person-column restoration
(asec_census_person_columns), which appends reviewed columns per vintage
from the pinned person member by exact PERIDNUM before pooling. It is
declared separately as ASEC_CENSUS_PERSON_COLUMNS_BEYOND_720, so the #720
review still equals the offline fix's column list exactly. Its domain
{0,1,2,3,4,7} was observed on 2026-09-26 in all three pinned members
(archive and member SHA-256 verified; complete; code 0 is the only code
under age 15). The cited readers resolve and name the column. A pool without a
Census person source still refuses at immigration_status; there is no
fallback. test_us_plan's closure now counts the restoration register as a
declared provider.

Two related register fixes. The plan's immigration DonorSpec still cited
Pew's 2024 short read, while source_stages.json cites the 2025 report the
controls come from (test_every_donor_stage_has_matching_source_spec). And
NOW_CAID was recorded as having "no build reader" while the native health
projection names it. That projection captures its own pinned ASEC member and
never reads the pooled H5, so NOW_CAID now carries that reviewed reason. A new
test fails if any other module names it.

Tests: test_us_asec_census_person_columns (33), test_us_plan (36),
test_us_puf_support_base_builder and test_us_immigration pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Native commit 071f563 ported #779's paired immigration transfer, which
requires exactly one row-level citizenship source (ASEC PRCITSHP or ACS CIT)
with its birth-country and entry-year partners on every donor and recipient
row. It did not port #779's ACS loader change (3c3bdbf) that carries
CIT/POBP/YOEP, so any multispine pool with ACS rows refused
(test_every_pool_transfer_family_accepts_its_produced_physical_dtype on
every lane: missing_rows=2, the two ACS-channel clones).

acs_pums now reads CIT/POBP/YOEP as optional person columns, like #765's
hours columns. Every ACS PUMS person file carries them. A spine loaded
without them stays loadable, and a paired immigration transfer onto it still
refuses; nothing is defaulted. The reviewed acs_pums.py fingerprint in
acs_native_coverage_binding is re-pinned with that review noted. The pool
fixture's ACS rows carry the evidence, as #779's did, and the transfer
gains its three evidence predictors (32 -> 35). A new loader test checks
carriage when supplied and absence otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The native input_coverage_profile froze the release coverage manifest at
05d254a (163 required inputs). Main since requires 167: the SPM
independence role (#959) and the reported receipts receives_snap,
receives_tanf and receives_wic (#978). The profile's anti-rot test failed on
every lane. The historical roster is re-extracted in manifest order with the
manifest digest re-pinned. Because the identifiers encode their counts, they
move to us_release_167_v1, us_national_cd_165_v1 and
us_native_national_cd_163_v1. The derivation rules are unchanged: national/CD
omits block/tract, and native also omits the prior-year family. The
survey-enrichment coverage count moves 161 -> 165 with the national/CD
profile.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The wheels lane installs no hypothesis. The module-level import added in
a57e987 failed collection there, which interrupted the whole wheels 2/8
shard. The property test now imports hypothesis behind
pytest.importorskip, as test_us_fiscal_targets.py does; it still runs in
every lane that has hypothesis.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test_gate_evidence_files_are_the_files_their_gates_write (Route A PR-3)
counted three solve calls in _main. On this line the dense and L0 solves
live in the shared _calibrate_fiscal_support, which _main calls once beside
the exact-k ladder. The test now counts those two calls, pins that the
helper holds exactly calibrate and calibrate_l0_refit, and keeps the
invariant that _calibration_runtime() precedes every solve.

test_pinned_feed_national_state_surface_restores_the_fences (it runs only
where the pinned feed exists, not in CI) pinned main's 32,842 / 5,694. On the
merged tree the pinned feed compiles 32,858 / 5,710. The 16 extra rows are
exactly the native SOI Table 1.1 size-of-AGI classes (#958: eight national
classes from $100k, AGI and return count), observed by their
requires_agi_size_distribution_rebase flag. The test now asserts main's
counts plus those 16 and the new surface version. Every Route A fence
assertion is unchanged, and all of them hold: 32 CHIP rows and none for an
M-CHIP state, no other-income row, no tips return count, and the tips
amount kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sses

On Python 3.14, PEP 649 compiles class annotations into an __annotate__
function closed over the class namespace through a __classdict__ cell.
survey_population_preparation._function_seal recorded every closure cell by
value. The first harmless read of a producer class's annotations caches them
in that namespace, so the live seal no longer equalled the import-time seal,
and every later preparation refused with PRODUCER_CHANGED. Observed with a
seal diff: test_us_native_calibration_attachment.py read
microcosm.frame.weights.MassChangeRecord's annotations, and the only moved
entry was ('microcosm.frame.weights', 'MassChangeRecord',
'__annotate_func__'). Reproduced locally by running that file before
test_us_native_child_support.py; it fails on the 3.14 CI lanes only
(us-not, fast rest 1/6, wheels 1/8), because 3.13 has no such cell.

A __classdict__ cell is now bound by the namespace's identity and type. The
namespace's functions and classes are already sealed individually, and
every other closure cell is still sealed by value. A new test reads the
annotations and requires the seal to stay equal, then swaps the namespace
behind the cell and requires the seal to change. It fails under the old
rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The native line's run_uk_calibration pins the executing checkout's git
commit (git_code_pin(_REPOSITORY)) and refuses when it cannot. The wheel
lane runs from /tmp/wheels-venv, which is not a checkout. Native tests
request the invented_code_pin fixture, but main's newer seam tests drove the
real seam without it, and failed there with "Could not resolve the local git
code pin": three test_uk_rowwise_national_role tests (wheels 4/8), and
test_run_uk_calibration_accepts_a_caller_minted_attempt_id in wheels 2/8,
which that run's collection error had hidden.

The rowwise module re-exports the calibration tests' fixture, and those four
tests request it. Production stays strict. With GIT_DIR pointed at a
nonexistent path (git unavailable, as in the wheel venv) all four pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…alled

The native line's CI adds a sharded wheels group that runs the whole suite
against installed wheels; main's wheel job never selected this file. The
donor-receipt producer identity binds this checkout's source paths and git
commit by design, so under /tmp/wheels-venv it refuses at cps_carried.py
before any injected foreign module. Four tests expected either success or a
specific foreign-module message, and failed in wheels 6/8.

They now skip only when the producer's own cps_carried module is not this
checkout's file. The local-shards refusal test still runs everywhere and
covers that refusal. In a checkout all seven producer tests pass, and a
simulated installed-module path takes the skip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One table row per failure class fixed after the 2bce661 merge: cause and
fix, as observed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fix)

Independent review of 55ce189 found that sealing the __classdict__ cell
by identity alone dropped protection the old seal had on Python 3.14.
Deleting a MassChangeRecord dataclass field in place left the seal
unchanged, while dataclasses.asdict then drops the field. It also found
that the stated cause was wrong. A snapshot diff of the live namespace
across test_us_native_calibration_attachment.py shows the only moved entry
is __slotnames__, which copyreg._slotnames caches when an instance is
copied (the attachment bridge deep-copies the mass log). Annotations played
no part.

The cell is now sealed by identity and by its contents minus __slotnames__
alone, so every other class attribute, dataclass fields included, stays
sealed. The regression test requires a copy to pass, an in-place field edit
to refuse, and a replaced namespace to refuse. It fails under both the
original by-value rule and the identity-only rule. The failing CI order
(the attachment tests, then native child support) passes.

Also from the review: docs/us-input-coverage-diagnostic.md now gives the
current 163/165/167-name profiles, and the integration record qualifies the
default-call-shape claim (some internal helper calls pass explicit None
seams, with no behavioral effect).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant