From 3e733a8d144ee17a7071723d76ec6503cd9654f9 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:30:29 -0400 Subject: [PATCH 01/22] Open the lane journal for the row-count ceilings Records the brief, the four bounds the transport lane's census found binding between 1/10 and full source, and the one-argument decision. Co-Authored-By: Claude Opus 5 --- PROGRESS-native-row-ceilings.md | 45 +++++++++++++++++++++++++++++++++ 1 file changed, 45 insertions(+) create mode 100644 PROGRESS-native-row-ceilings.md diff --git a/PROGRESS-native-row-ceilings.md b/PROGRESS-native-row-ceilings.md new file mode 100644 index 000000000..476a10c2e --- /dev/null +++ b/PROGRESS-native-row-ceilings.md @@ -0,0 +1,45 @@ +# Lane journal: lift the native build's row-count ceilings (`native-row-ceilings`) + +Branch `native-row-ceilings`, from `origin/native-scale-transport` at `a64f7b733`. +PR base is `native-scale-transport`, draft, and it stays draft. + +This file is a session-handoff journal. Per the repo's CLAUDE.md, its +"State"/"Next" sections are accurate when written and historical afterward — +check git and GitHub for current truth. + +## The brief + +The transport lane's census (`docs/us-native-scale-transport.md` §5) found that +after the 64 MiB transport ceilings were lifted, a full-source native US build +(1,587,376 source households; ~3,471,000 stacked and ~6,943,000 cloned persons) +still meets at least four row-count bounds: + +- `acs_person_coverage_columns.MAX_SELECTED_ROWS` — 1,000,000 +- `asec_current_money.MAX_PERSONS` — 1,000,000 +- `acs_pums.MAX_EXACT_HOUSEHOLDS` / `MAX_EXACT_PERSON_ROWS` — 1,000,000 +- `survey_observed_age.MAX_ROWS` — 2,000,000 + +and that `survey_origin_budget.MAX_GROUPS` (1,000,000) was not established. +Max decided (2026-09-17) they are lifted as **one change with one argument**, +not one build at a time. + +## State + +Census in progress. Nothing implemented yet. + +## Done + +- Worktree verified at `a64f7b733`; `uv sync --all-packages --locked --extra us` exit 0. +- Read the transport lane's §2c/§2e (the `MAX_ROSTER_BYTES` argument shape: + an explicit resource ceiling at a fixed multiple above the full-source count) + and its §4 (how it re-derived pins through `graph_implementation._dependency_contract`). + +## Next + +1. Census every `MAX_*` bound in `us_runtime/` reachable from the 19-node + financial graph and the 45-node pilot graph, from code at this head. +2. Establish `survey_origin_budget.MAX_GROUPS`. +3. Write `docs/us-native-row-ceilings.md` — the one argument. +4. Implement; a test per moved bound. +5. Re-derive every moved pin through its generator. +6. Draft PR against `native-scale-transport`. From 1d5cfafb111ef89863a7746697e39f8372ee8c80 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:32:26 -0400 Subject: [PATCH 02/22] Record the 1/1000 roster ratios the census scales from The recovered pilot artifact's own rosters: 1,584 selected households carry 3,464 stacked persons (2.1869/hh) and six entity rosters. Every count in the census is this artifact scaled to 1,587,376 supplied households, so the derivation is stated rather than assumed. Co-Authored-By: Claude Opus 5 --- .../native-row-ceilings/roster-census.json | 32 +++++++++++++++++++ 1 file changed, 32 insertions(+) create mode 100644 experiments/native-row-ceilings/roster-census.json diff --git a/experiments/native-row-ceilings/roster-census.json b/experiments/native-row-ceilings/roster-census.json new file mode 100644 index 000000000..7e52eac6c --- /dev/null +++ b/experiments/native-row-ceilings/roster-census.json @@ -0,0 +1,32 @@ +{ + "artifact": "/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json", + "artifact_sha256": "34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0", + "scope": "Census input only. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "fraction": [ + 1, + 1000 + ], + "supplied_households": 1587376, + "selected_households": 1584, + "excluded_households": 65, + "cells": 5, + "origin_household_rows": 1584, + "origin_person_rows": 3464, + "origin_entities": { + "family": 1594, + "household": 1584, + "marital_unit": 2774, + "person": 3464, + "spm_unit": 1585, + "tax_unit": 2128 + }, + "per_household": { + "family": 1.0063131313131313, + "household": 1.0, + "marital_unit": 1.7512626262626263, + "person": 2.186868686868687, + "spm_unit": 1.0006313131313131, + "tax_unit": 1.3434343434343434 + } +} From 1bccb2e0485789f51a4535a8d66832fabc3a1c1f Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:36:17 -0400 Subject: [PATCH 03/22] Measure the full-source counts from the artifact's catalogues, not by scaling The recovered artifact carries the whole ACS and ASEC catalogues, which do not scale with the selection fraction. Their selectable households reconcile to selection.supplied_households exactly (1,348,408 + 84,422 + 98,784 + 55,762 = 1,587,376), so a full-source selection is the catalogue and the counts are measured rather than extrapolated: ACS 1,531,614 households / 3,422,888 persons ASEC 55,762 households / 142,125 persons stacked 1,587,376 / 3,565,013; combined clone 3,174,752 / 7,130,026 The transport lane's ~3,471,000 stacked and ~6,943,000 cloned were the 1/1000 per-household ratio extrapolated; the catalogue figures supersede them. Co-Authored-By: Claude Opus 5 --- .../native-row-ceilings/roster-census.json | 96 ++++++++++---- .../native-row-ceilings/roster_census.py | 117 ++++++++++++++++++ 2 files changed, 188 insertions(+), 25 deletions(-) create mode 100644 experiments/native-row-ceilings/roster_census.py diff --git a/experiments/native-row-ceilings/roster-census.json b/experiments/native-row-ceilings/roster-census.json index 7e52eac6c..0c039d6a8 100644 --- a/experiments/native-row-ceilings/roster-census.json +++ b/experiments/native-row-ceilings/roster-census.json @@ -1,32 +1,78 @@ { "artifact": "/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json", "artifact_sha256": "34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0", - "scope": "Census input only. Not a build, not a certification, not release eligible.", + "scope": "Census input for the row-count ceiling lane. Reads one recovered development artifact and writes one JSON summary. Not a build, not a certification, not release eligible.", "release_eligible": false, - "fraction": [ - 1, - 1000 - ], - "supplied_households": 1587376, - "selected_households": 1584, - "excluded_households": 65, - "cells": 5, - "origin_household_rows": 1584, - "origin_person_rows": 3464, - "origin_entities": { - "family": 1594, - "household": 1584, - "marital_unit": 2774, - "person": 3464, - "spm_unit": 1585, - "tax_unit": 2128 + "catalogue_reconciles_to_supplied_households": true, + "catalogues": { + "acs": { + "canonical_record_bytes": 374174650, + "canonical_record_sha256": "d200e4e16eb8c6baacc29244921c1f7b25f46aebd9129a4aa06049509c25e3f2", + "households": 1631969, + "institutional_gq": 84422, + "literal_batches": 35, + "noninstitutional_gq": 98784, + "occupied_hu": 1348408, + "people": 3422888, + "vacancies": 100355 + }, + "asec": { + "households": 55762, + "persons": 142125, + "unrepresented_households": 33170 + } }, - "per_household": { - "family": 1.0063131313131313, - "household": 1.0, - "marital_unit": 1.7512626262626263, - "person": 2.186868686868687, - "spm_unit": 1.0006313131313131, - "tax_unit": 1.3434343434343434 + "full_source": { + "derivation": "the catalogues themselves; a full-source selection supplies every selectable household in both, which the reconciliation above proves", + "acs_households": 1531614, + "acs_persons": 3422888, + "asec_households": 55762, + "asec_persons": 142125, + "stacked_households": 1587376, + "stacked_persons": 3565013, + "combined_clone_households": 3174752, + "combined_clone_persons": 7130026 + }, + "sample_at_one_thousandth": { + "fraction": [ + 1, + 1000 + ], + "selected_households": 1584, + "origin_household_rows": 1584, + "origin_person_rows": 3464, + "households_by_source": { + "acs": 1529, + "asec": 55 + }, + "persons_by_source": { + "acs": 3324, + "asec": 140 + }, + "origin_entities": { + "family": 1594, + "household": 1584, + "marital_unit": 2774, + "person": 3464, + "spm_unit": 1585, + "tax_unit": 2128 + }, + "entities_per_selected_household": { + "family": 1.0063131313131313, + "household": 1.0, + "marital_unit": 1.7512626262626263, + "person": 2.186868686868687, + "spm_unit": 1.0006313131313131, + "tax_unit": 1.3434343434343434 + } + }, + "one_tenth": { + "derivation": "the full-source counts above, times 1/10", + "stacked_households": 158737, + "acs_households": 153161, + "acs_persons": 342288, + "asec_persons": 14212, + "stacked_persons": 356501, + "combined_clone_persons": 713002 } } diff --git a/experiments/native-row-ceilings/roster_census.py b/experiments/native-row-ceilings/roster_census.py new file mode 100644 index 000000000..da4b030cb --- /dev/null +++ b/experiments/native-row-ceilings/roster_census.py @@ -0,0 +1,117 @@ +"""Derive the full-source row counts every ceiling in this lane is measured against. + +Reads the recovered 1/1000 pilot preparation artifact and reports two things: + +1. Its **catalogues**, which are counts of the whole upstream ACS and ASEC files + and therefore do not scale with the selection fraction. These give the + full-source selected counts exactly, not by extrapolation: a full-source + build selects every selectable household in both catalogues. +2. Its **origins** rosters at 1/1000, which give the per-household ratios and + let the catalogue figures be cross-checked against an observed sample. + +The reconciliation that makes (1) authoritative is asserted here rather than +asserted in prose: the ACS catalogue's occupied, institutional-GQ and +noninstitutional-GQ households plus the ASEC catalogue's households equal +``selection.supplied_households`` exactly, so "the whole catalogue" and "what a +full-source selection supplies" are the same set. + +No graph runs; nothing outside the output path is written. Not a build, not a +certification, not release eligible. + + python roster_census.py +""" + +from __future__ import annotations + +import collections +import hashlib +import json +import pathlib +import sys + + +def main() -> int: + artifact = pathlib.Path(sys.argv[1]) + out = pathlib.Path(sys.argv[2]) + raw = artifact.read_bytes() + document = json.loads(raw) + selection = document["selection"] + origins = document["origins"] + acs = document["catalogues"]["acs"]["counts"] + asec = document["catalogues"]["asec"]["counts"] + + acs_selectable = ( + acs["occupied_hu"] + acs["institutional_gq"] + acs["noninstitutional_gq"] + ) + supplied = selection["supplied_households"] + reconciles = acs_selectable + asec["households"] == supplied + if not reconciles: + raise SystemExit( + "catalogue households do not reconcile to supplied_households; " + "the full-source counts below would be extrapolations, not measurements" + ) + + columns = origins["persons"]["columns"] + source_at = columns.index("source") + sample_persons = collections.Counter( + row[source_at] for row in origins["persons"]["rows"] + ) + sample_households = collections.Counter( + row["source"] for row in origins["households"] + ) + selected = len(selection["selected"]) + + record = { + "artifact": str(artifact), + "artifact_sha256": hashlib.sha256(raw).hexdigest(), + "scope": ( + "Census input for the row-count ceiling lane. Reads one recovered " + "development artifact and writes one JSON summary. Not a build, not " + "a certification, not release eligible." + ), + "release_eligible": False, + "catalogue_reconciles_to_supplied_households": reconciles, + "catalogues": {"acs": acs, "asec": asec}, + "full_source": { + "derivation": ( + "the catalogues themselves; a full-source selection supplies every " + "selectable household in both, which the reconciliation above proves" + ), + "acs_households": acs_selectable, + "acs_persons": acs["people"], + "asec_households": asec["households"], + "asec_persons": asec["persons"], + "stacked_households": supplied, + "stacked_persons": acs["people"] + asec["persons"], + "combined_clone_households": supplied * 2, + "combined_clone_persons": (acs["people"] + asec["persons"]) * 2, + }, + "sample_at_one_thousandth": { + "fraction": selection["fraction"], + "selected_households": selected, + "origin_household_rows": len(origins["households"]), + "origin_person_rows": len(origins["persons"]["rows"]), + "households_by_source": dict(sample_households), + "persons_by_source": dict(sample_persons), + "origin_entities": {k: len(v) for k, v in origins["entities"].items()}, + "entities_per_selected_household": { + k: len(v) / selected for k, v in origins["entities"].items() + }, + }, + "one_tenth": { + "derivation": "the full-source counts above, times 1/10", + "stacked_households": supplied // 10, + "acs_households": acs_selectable // 10, + "acs_persons": acs["people"] // 10, + "asec_persons": asec["persons"] // 10, + "stacked_persons": (acs["people"] + asec["persons"]) // 10, + "combined_clone_persons": (acs["people"] + asec["persons"]) * 2 // 10, + }, + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(json.dumps(record["full_source"], indent=1)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) From e660ec6ccc1675390630c0b19c5c0b2e7bc20878 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:41:01 -0400 Subject: [PATCH 04/22] Measure the origin-budget payload, which binds harder than anything the census inherited survey_origin_budget streams one ~731-byte origin record per allocation group into one bytearray under a 64 MiB cap. Measured through the module's own _reference and _json at full-source household-id widths: 764 bytes per group, so the cap admits 87,838 households -- 5.53% of source, below 1/10 and below the 96,860-household preparation-receipt ceiling the transport lane lifted. A full-source payload is 1.13 GiB, 18.07x the cap. Also records where the module runs: no node in the 19-node financial graph (graph.json of the recovered run lists all nineteen) or in the completion host executes it; survey_age_calibration and graph_survey_budget do, over the same full-source selection. Co-Authored-By: Claude Opus 5 --- .../origin-budget-size.json | 48 ++++++++ .../native-row-ceilings/origin_budget_size.py | 111 ++++++++++++++++++ 2 files changed, 159 insertions(+) create mode 100644 experiments/native-row-ceilings/origin-budget-size.json create mode 100644 experiments/native-row-ceilings/origin_budget_size.py diff --git a/experiments/native-row-ceilings/origin-budget-size.json b/experiments/native-row-ceilings/origin-budget-size.json new file mode 100644 index 000000000..62c96f750 --- /dev/null +++ b/experiments/native-row-ceilings/origin-budget-size.json @@ -0,0 +1,48 @@ +{ + "scope": "Encoder-measured size law for the origin-budget payload. Builds one faithful record through the module's own _reference and _json. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "max_payload_bytes": 67108864, + "max_groups": 1000000, + "full_source_households": 1587376, + "measurements": [ + { + "household_id_magnitude": 1000, + "record_bytes": 719, + "header_ids_and_groups_bytes": 20, + "per_group_bytes": 740, + "households_at_64MiB": 90687, + "full_source_payload_bytes": 1174658240 + }, + { + "household_id_magnitude": 1000000, + "record_bytes": 731, + "header_ids_and_groups_bytes": 32, + "per_group_bytes": 764, + "households_at_64MiB": 87838, + "full_source_payload_bytes": 1212755264 + }, + { + "household_id_magnitude": 1587376, + "record_bytes": 731, + "header_ids_and_groups_bytes": 32, + "per_group_bytes": 764, + "households_at_64MiB": 87838, + "full_source_payload_bytes": 1212755264 + }, + { + "household_id_magnitude": 3174752, + "record_bytes": 731, + "header_ids_and_groups_bytes": 32, + "per_group_bytes": 764, + "households_at_64MiB": 87838, + "full_source_payload_bytes": 1212755264 + } + ], + "binding": { + "per_group_bytes_at_full_source_ids": 764, + "households_the_64MiB_cap_admits": 87838, + "fraction_of_source_that_admits": 0.05533534587898519, + "full_source_payload_bytes": 1212755264, + "full_source_payload_over_cap": 18.07146167755127 + } +} diff --git a/experiments/native-row-ceilings/origin_budget_size.py b/experiments/native-row-ceilings/origin_budget_size.py new file mode 100644 index 000000000..888e1938c --- /dev/null +++ b/experiments/native-row-ceilings/origin_budget_size.py @@ -0,0 +1,111 @@ +"""Measure the origin-budget payload's per-group byte cost, using the module's own encoder. + +``survey_origin_budget._budget_payload`` streams one origin record per allocation +group into a single bytearray guarded by ``MAX_PAYLOAD_BYTES``. The record is a +fixed-shape JSON object whose only variable-width field is the raw native key, so +its encoded size is measurable from one faithfully constructed record rather than +estimated. This builds that record through the module's own ``_reference`` and +``_json`` and reports the household count at which the 64 MiB cap is met. + +The per-group cost also includes the two ``household_ids`` and two +``group_indices`` entries the header carries for each group's two clone roles. + +No graph runs, no source is read, nothing outside the output path is written. +Not a build, not a certification, not release eligible. + + python origin_budget_size.py +""" + +from __future__ import annotations + +import json +import pathlib +import sys +from fractions import Fraction + +sys.path[:0] = [ + str(path) + for path in sorted( + (pathlib.Path(__file__).resolve().parents[2] / "packages").glob("*/src") + ) +] + +from microcosm.build.us_runtime import survey_origin_budget as owner # noqa: E402 + +FULL_SOURCE_HOUSEHOLDS = 1_587_376 # measured; see roster-census.json + + +def _record(native_id: str, household_id: int, clone_ids: tuple[int, int]) -> dict: + design = Fraction(137, 1) # a representative ACS WGTP + probability = Fraction(1584, 1587376) + share = Fraction(1, 2) + actual = float(design * share / probability) + reference = owner._reference(design, probability, share, actual) + return { + "source": "acs", + "source_year": 2024, + "survey_year": 2024, + "raw_native_id": native_id, + "selected_receiving_household_id": household_id, + "combined_household_id": household_id, + "statistical_unit": "occupied_housing_unit", + "original_design_float64_bytes": "0" * 16, + **reference, + "members": [[clone_ids[0], 0], [clone_ids[1], 1]], + "incoming_clone_float64_bytes": ["0" * 16, "0" * 16], + } + + +def main() -> int: + out = pathlib.Path(sys.argv[1]) + rows = [] + # Measure at household-id magnitudes a full-source build actually reaches, so + # the integer widths in the encoded record are the real ones. + for magnitude in (1_000, 1_000_000, FULL_SOURCE_HOUSEHOLDS, 3_174_752): + record = _record("2024HU%07d" % (magnitude % 10_000_000), magnitude, (magnitude, magnitude * 2)) + encoded = len(owner._json(record)) + # ",": one separator per record after the first. + # header ids/groups: two household ids and two group indices per group. + ids = len(str(magnitude * 2)) * 2 + 2 + groups = len(str(magnitude)) * 2 + 2 + per_group = encoded + 1 + ids + groups + rows.append( + { + "household_id_magnitude": magnitude, + "record_bytes": encoded, + "header_ids_and_groups_bytes": ids + groups, + "per_group_bytes": per_group, + "households_at_64MiB": owner.MAX_PAYLOAD_BYTES // per_group, + "full_source_payload_bytes": per_group * FULL_SOURCE_HOUSEHOLDS, + } + ) + worst = max(rows, key=lambda r: r["per_group_bytes"]) + record = { + "scope": ( + "Encoder-measured size law for the origin-budget payload. Builds one " + "faithful record through the module's own _reference and _json. Not a " + "build, not a certification, not release eligible." + ), + "release_eligible": False, + "max_payload_bytes": owner.MAX_PAYLOAD_BYTES, + "max_groups": owner.MAX_GROUPS, + "full_source_households": FULL_SOURCE_HOUSEHOLDS, + "measurements": rows, + "binding": { + "per_group_bytes_at_full_source_ids": worst["per_group_bytes"], + "households_the_64MiB_cap_admits": worst["households_at_64MiB"], + "fraction_of_source_that_admits": worst["households_at_64MiB"] + / FULL_SOURCE_HOUSEHOLDS, + "full_source_payload_bytes": worst["full_source_payload_bytes"], + "full_source_payload_over_cap": worst["full_source_payload_bytes"] + / owner.MAX_PAYLOAD_BYTES, + }, + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(json.dumps(record["measurements"], indent=1)) + print(json.dumps(record["binding"], indent=1)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) From 14defbfc0850216074b211684cadde46705b93b9 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:44:06 -0400 Subject: [PATCH 05/22] Lift the four binding row-count ceilings under one rule Four times the measured full-source count of exactly what each bounds, rounded up to the next whole million: acs_pums.MAX_EXACT_HOUSEHOLDS 1,000,000 -> 7,000,000 (1,531,614) acs_pums.MAX_EXACT_PERSON_ROWS 1,000,000 -> 14,000,000 (3,422,888) acs_person_coverage_columns.MAX_SELECTED_ROWS 1,000,000 -> 14,000,000 (3,422,888) survey_observed_age.MAX_ROWS 2,000,000 -> 14,000,000 (3,422,888) survey_origin_budget.MAX_GROUPS 1,000,000 -> 7,000,000 (1,587,376) Every refusal keeps its code and its expression; only the number moves. acs_person_coverage_columns.MAX_ROWS stays at 6,000,000 because it asserts the source file's own size rather than a roster this build chooses, and survey_origin_budget.MAX_PAYLOAD_BYTES stays at 64 MiB because it is a byte transport and takes the segmented-transport argument, not this one. Co-Authored-By: Claude Opus 5 --- .../build/us_runtime/acs_person_coverage_columns.py | 8 +++++++- .../src/microcosm/build/us_runtime/acs_pums.py | 10 ++++++++-- .../microcosm/build/us_runtime/survey_observed_age.py | 7 ++++++- .../microcosm/build/us_runtime/survey_origin_budget.py | 8 +++++++- 4 files changed, 28 insertions(+), 5 deletions(-) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_person_coverage_columns.py b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_person_coverage_columns.py index 53b692dad..5f1149743 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_person_coverage_columns.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_person_coverage_columns.py @@ -30,8 +30,14 @@ DICTIONARY_SHA256 = "929c2752995b0af1c16d5c64de8cdc43b4aa7d388ee2d45b4b4df90fecce1dff" KEYS = ("SERIALNO", "SPORDER") READ_COLUMNS = (*KEYS, "AGEP", "MIL", "ESR") +# MAX_ROWS bounds the source file's own person records and stays where it +# is: the 2024 ACS person file holds 3,422,888, and the bound is the +# structural assertion that a genuine file is near that size. Only the +# requested-roster ceiling moves, to four times the 3,422,888 persons a +# full-source selection requests, rounded up to the next whole million. +# See docs/us-native-row-ceilings.md. MAX_ROWS = 6_000_000 -MAX_SELECTED_ROWS = 1_000_000 +MAX_SELECTED_ROWS = 14_000_000 MAX_CSV_RECORD_CHARS = 100_000 _NON_CSV_CONTROLS = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]") diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_pums.py b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_pums.py index f4431a860..e1c44ee3e 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_pums.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_pums.py @@ -46,8 +46,14 @@ ACS_2024_1YR_SPINE = "acs_2024_1yr" ACS_2024_1YR_VINTAGE = 2024 DEFAULT_CHUNKSIZE = 100_000 -MAX_EXACT_HOUSEHOLDS = 1_000_000 -MAX_EXACT_PERSON_ROWS = 1_000_000 +# Explicit resource ceilings, four times the measured full-source count of +# exactly what each bounds, rounded up to the next whole million. A +# full-source ACS selection is the 1,531,614 selectable households of the +# 2024 catalogue, carrying 3,422,888 person rows. Neither is a statement +# about the source file, which MAX_ROWS in acs_person_coverage_columns +# still makes at 6,000,000. See docs/us-native-row-ceilings.md. +MAX_EXACT_HOUSEHOLDS = 7_000_000 +MAX_EXACT_PERSON_ROWS = 14_000_000 _HOUSEHOLD_REQUIRED = ( "SERIALNO", diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/survey_observed_age.py b/packages/microcosm-build/src/microcosm/build/us_runtime/survey_observed_age.py index fde1388e8..804f2f6ab 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/survey_observed_age.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/survey_observed_age.py @@ -11,7 +11,12 @@ RULE = "microcosm.us.observed-age-normalization.v1" AGE_CONVENTION = "observed_interview_age_completed_years" -MAX_ROWS = 2_000_000 +# One channel's person rows at a time: _normalized_source_copy normalizes the +# ACS and ASEC native frames separately, so the ceiling is met by the larger +# channel. A full-source ACS channel carries 3,422,888 persons, and this is +# four times that, rounded up to the next whole million. +# See docs/us-native-row-ceilings.md. +MAX_ROWS = 14_000_000 MAX_EXACT_FLOAT64_INTEGER = 2**53 diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py b/packages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py index bdf2029fe..8c374314d 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py @@ -52,7 +52,13 @@ BUDGET_TYPE = ArtifactType("microcosm.us.sampling_origin_budget", 1) SUCCESSOR_TYPE = ArtifactType("microcosm.us.sampling_origin_weight_only_successor", 1) MAX_PAYLOAD_BYTES = 64 * 1024**2 -MAX_GROUPS = 1_000_000 +# One group per allocation instruction, and allocation_instructions requires +# one instruction per selected household, so a full-source budget has the +# 1,587,376 households the catalogues supply. Four times that, rounded up to +# the next whole million. MAX_PAYLOAD_BYTES above is a byte transport, not a +# row count, and is deliberately left alone: it takes the segmented transport +# argument, not this one. See docs/us-native-row-ceilings.md. +MAX_GROUPS = 7_000_000 MAX_SCALAR_CHARS = 4096 PRESCRIPTION = ( "development-provisional-8b-4a-v1", From 55ca820c7c3e1485c09fd640e33bae8dd23c1714 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:46:28 -0400 Subject: [PATCH 06/22] Re-pin the inventory contract the base branch moved and did not regenerate INHERITED, NOT THIS LANE'S. On native-scale-transport, b6081efcb ("Hash a spill segment that was already there") added path.read_bytes() inside _spill_roster. read_bytes is in graph_implementation._RESOURCE_CALLS, so it moves survey_population_preparation.py's resource_accesses_sha256 -- and the pin was not regenerated. implementation_manifest() therefore raises "Unclassified US dependency/resource contract" for every stage containing that module, including authenticated_survey_population_v1, which is the stage the nineteen-node path runs. The break is at origin/native-scale-transport and at a64f7b733, proven by recomputing the contract from each commit's own blobs; this branch does not touch that file. The transport lane's report was right that Path.write_bytes is invisible to resource_accesses_sha256 by construction. This is the converse: a later commit on the same branch added a read, which is visible. Regenerated with experiments/native-row-ceilings/regenerate_inventory_contract.py, which writes exactly what graph_implementation._dependency_contract returns. Never hand-edited. Serialization verified byte-identical before the edit, so one value moves and nothing else. resource_accesses_sha256 ccfed1c1acff... -> 0071f934801d... Co-Authored-By: Claude Opus 5 --- .../regenerate_inventory_contract.py | 67 +++++++++ experiments/native-row-ceilings/repin.py | 139 ++++++++++++++++++ .../graph_implementation_inventory.json | 2 +- 3 files changed, 207 insertions(+), 1 deletion(-) create mode 100644 experiments/native-row-ceilings/regenerate_inventory_contract.py create mode 100644 experiments/native-row-ceilings/repin.py diff --git a/experiments/native-row-ceilings/regenerate_inventory_contract.py b/experiments/native-row-ceilings/regenerate_inventory_contract.py new file mode 100644 index 000000000..71232952b --- /dev/null +++ b/experiments/native-row-ceilings/regenerate_inventory_contract.py @@ -0,0 +1,67 @@ +"""Regenerate an inventory contract through its own generator. Never hand-edit a pin. + +``graph_implementation.implementation_manifest`` recomputes every inventoried +module's dependency contract and refuses the stage if one differs from +``graph_implementation_inventory.json``. This rewrites the differing entries +with exactly what ``_dependency_contract`` returns, preserving the file's key +order and its trailing newline, and prints what moved. + + python regenerate_inventory_contract.py [--write] +""" + +from __future__ import annotations + +import json +import pathlib +import sys + +ROOT = pathlib.Path(__file__).resolve().parents[2] +sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] + +from microcosm.build.us_runtime import graph_implementation as implementation # noqa: E402 + +INVENTORY = ( + ROOT + / "packages/microcosm-build/src/microcosm/build/us_runtime" + / "graph_implementation_inventory.json" +) + + +def main() -> int: + write = "--write" in sys.argv[1:] + raw = INVENTORY.read_text() + inventory = json.loads(raw) + roots = implementation._package_roots() + moved = [] + for name, expected in inventory["contracts"].items(): + package, relative = name.split("/", 1) + payload = (roots[package] / relative).read_bytes() + actual = implementation._dependency_contract( + payload, name, implementation._covered_imports(name, inventory) + ) + if actual != expected: + moved.append((name, expected, actual)) + inventory["contracts"][name] = actual + for name, expected, actual in moved: + print(f"MOVED {name}") + for key in ("imports", "unbound_uses_sha256", "resource_accesses_sha256"): + if expected[key] != actual[key]: + print(f" {key}") + print(f" old: {expected[key]}") + print(f" new: {actual[key]}") + if not moved: + print("no contract moved") + return 0 + if not write: + print("(dry run; pass --write to apply)") + return 0 + # The committed file is json.dumps(..., indent=2, sort_keys=True) plus a + # trailing newline; this was verified byte-identical before any edit, so the + # rewrite moves the changed values and nothing else. + INVENTORY.write_text(json.dumps(inventory, indent=2, sort_keys=True) + "\n") + print(f"rewrote {INVENTORY.name}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/experiments/native-row-ceilings/repin.py b/experiments/native-row-ceilings/repin.py new file mode 100644 index 000000000..eec20704a --- /dev/null +++ b/experiments/native-row-ceilings/repin.py @@ -0,0 +1,139 @@ +"""Re-derive every pin these edits can move, each through its own generator. + +Three pin families touch the four edited modules: + +1. ``graph_implementation_inventory.json``'s per-module ``contracts`` entry -- + ``imports``, ``unbound_uses_sha256`` and ``resource_accesses_sha256``, + re-derived through ``graph_implementation._dependency_contract`` against the + module's own ``_covered_imports``, exactly as ``implementation_manifest`` + checks it. A constant's value is not an import, an unbound use or a resource + access, so these are expected to hold; the point is to prove it rather than + assume it. +2. ``acs_native_coverage_binding._ACCEPTED``, which pins four ACS modules by + whole-file sha256. ``acs_pums.py`` is one of them, so that pin moves. +3. Each declared stage's ``implementation_manifest``, whose ``modules`` map is + the same whole-file sha256 per inventoried module. These are computed, not + committed, but every node key and store address is downstream of them, so + the manifests are rebuilt here to prove they still build. + +Nothing is written outside the output path. Not a build, not a certification, +not release eligible. + + python repin.py +""" + +from __future__ import annotations + +import hashlib +import json +import pathlib +import sys + +ROOT = pathlib.Path(__file__).resolve().parents[2] +sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] + +from microcosm.build.us_runtime import ( # noqa: E402 + acs_native_coverage_binding as binding, +) +from microcosm.build.us_runtime import graph_implementation as implementation # noqa: E402 + +US = ROOT / "packages/microcosm-build/src/microcosm/build/us_runtime" +EDITED = ( + "acs_pums.py", + "acs_person_coverage_columns.py", + "survey_observed_age.py", + "survey_origin_budget.py", +) + + +def main() -> int: + out = pathlib.Path(sys.argv[1]) + inventory = json.loads((US / "graph_implementation_inventory.json").read_bytes()) + + # 1. Inventory contracts, re-derived through the generator for every entry. + moved, checked = [], 0 + for name, expected in inventory["contracts"].items(): + package, relative = name.split("/", 1) + roots = implementation._package_roots() + payload = (roots[package] / relative).read_bytes() + actual = implementation._dependency_contract( + payload, name, implementation._covered_imports(name, inventory) + ) + checked += 1 + if actual != expected: + moved.append({"contract": name, "old": expected, "new": actual}) + + # 2. The ACS whole-file pin that names one of the edited modules. + accepted = [] + for name, expected in binding._ACCEPTED.items(): + actual = hashlib.sha256((US / name).read_bytes()).hexdigest() + accepted.append( + { + "module": name, + "old": expected, + "new": actual, + "moved": actual != expected, + "edited_by_this_branch": name in EDITED, + } + ) + + # 3. Every declared stage manifest, rebuilt through its generator. + stages = {} + for stage in sorted(implementation.STAGE_DEPENDENCIES): + manifest = implementation.implementation_manifest(stage) + stages[stage] = { + "inventory_sha256": manifest["inventory_sha256"], + "module_count": len(manifest["modules"]), + "edited_modules_in_stage": { + name: digest + for name, digest in manifest["modules"].items() + if pathlib.PurePosixPath(name).name in EDITED + }, + } + + record = { + "scope": ( + "Pin re-derivation for the row-ceiling lane. Reads the working tree and " + "rebuilds each pin through its own generator. Not a build, not a " + "certification, not release eligible." + ), + "release_eligible": False, + "edited_modules": { + name: hashlib.sha256((US / name).read_bytes()).hexdigest() + for name in EDITED + }, + "inventory_contracts": { + "generator": ( + "graph_implementation._dependency_contract(payload, name, " + "graph_implementation._covered_imports(name, inventory))" + ), + "checked": checked, + "moved": moved, + }, + "acs_native_coverage_binding_accepted": { + "generator": "hashlib.sha256(.read_bytes()).hexdigest()", + "entries": accepted, + }, + "stage_manifests": { + "generator": "graph_implementation.implementation_manifest(stage)", + "built": len(stages), + "stages": stages, + }, + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(f"inventory contracts checked: {checked}; moved: {len(moved)}") + for row in moved: + print(" MOVED", row["contract"]) + print("acs_native_coverage_binding._ACCEPTED:") + for row in accepted: + flag = "MOVED" if row["moved"] else "unchanged" + print(f" {row['module']}: {flag}") + if row["moved"]: + print(f" old {row['old']}") + print(f" new {row['new']}") + print(f"stage manifests built: {len(stages)}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/graph_implementation_inventory.json b/packages/microcosm-build/src/microcosm/build/us_runtime/graph_implementation_inventory.json index cf1c55f03..1959dc535 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/graph_implementation_inventory.json +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/graph_implementation_inventory.json @@ -946,7 +946,7 @@ "numpy", "pandas" ], - "resource_accesses_sha256": "ccfed1c1acff5a1c50b538424dc129930f69c374d9841e233b3fb593a440c3aa", + "resource_accesses_sha256": "0071f934801d15c14e4be66af12ff73daee5f65c89f61d655a23e462afdcd061", "unbound_uses_sha256": "d114117ca19e9b866982147e5497d601ed91ca603d0568730d4051bdd4005910" }, "microcosm.frame/__init__.py": { From 2ebd246f1bb1da4ded464522d7c7a0fd3c471e64 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:48:25 -0400 Subject: [PATCH 07/22] Re-pin the ACS accepted digest my acs_pums edit moves, and prove the rest hold acs_native_coverage_binding._ACCEPTED pins four ACS modules by whole-file sha256 and refuses UNREVIEWED_PREPARATION on a mismatch. acs_pums.py is one of them, so the ceiling edit moves it: 6ecf79f0dfb0c0bc... -> e79a2a4ecbc81e52... Regenerated with experiments/native-row-ceilings/regenerate_accepted_pin.py, which computes exactly the sha256 the check itself computes. Never hand-edited. repin.py now reports: 124 inventory contracts checked, 0 moved; the other three _ACCEPTED entries unchanged; all ten stage manifests built. Both generators are idempotent -- rerun, they report nothing moved. The pin tools set sys.path to this worktree and assert beneath the import block that each module resolved inside it, since a pin re-derived from another checkout would be the wrong value; pyproject carries the E402 ignore that ordering needs, beside the transport lane's own harness ignores. Co-Authored-By: Claude Opus 5 --- .../origin-budget-size.json | 2 +- .../native-row-ceilings/origin_budget_size.py | 8 +- .../regenerate_accepted_pin.py | 51 ++++++++ .../regenerate_inventory_contract.py | 13 +- experiments/native-row-ceilings/repin.json | 118 ++++++++++++++++++ experiments/native-row-ceilings/repin.py | 33 ++--- .../us_runtime/acs_native_coverage_binding.py | 2 +- pyproject.toml | 7 ++ 8 files changed, 211 insertions(+), 23 deletions(-) create mode 100644 experiments/native-row-ceilings/regenerate_accepted_pin.py create mode 100644 experiments/native-row-ceilings/repin.json diff --git a/experiments/native-row-ceilings/origin-budget-size.json b/experiments/native-row-ceilings/origin-budget-size.json index 62c96f750..1be652906 100644 --- a/experiments/native-row-ceilings/origin-budget-size.json +++ b/experiments/native-row-ceilings/origin-budget-size.json @@ -2,7 +2,7 @@ "scope": "Encoder-measured size law for the origin-budget payload. Builds one faithful record through the module's own _reference and _json. Not a build, not a certification, not release eligible.", "release_eligible": false, "max_payload_bytes": 67108864, - "max_groups": 1000000, + "max_groups": 7000000, "full_source_households": 1587376, "measurements": [ { diff --git a/experiments/native-row-ceilings/origin_budget_size.py b/experiments/native-row-ceilings/origin_budget_size.py index 888e1938c..84fe5feca 100644 --- a/experiments/native-row-ceilings/origin_budget_size.py +++ b/experiments/native-row-ceilings/origin_budget_size.py @@ -30,7 +30,7 @@ ) ] -from microcosm.build.us_runtime import survey_origin_budget as owner # noqa: E402 +from microcosm.build.us_runtime import survey_origin_budget as owner FULL_SOURCE_HOUSEHOLDS = 1_587_376 # measured; see roster-census.json @@ -62,7 +62,11 @@ def main() -> int: # Measure at household-id magnitudes a full-source build actually reaches, so # the integer widths in the encoded record are the real ones. for magnitude in (1_000, 1_000_000, FULL_SOURCE_HOUSEHOLDS, 3_174_752): - record = _record("2024HU%07d" % (magnitude % 10_000_000), magnitude, (magnitude, magnitude * 2)) + record = _record( + f"2024HU{magnitude % 10_000_000:07d}", + magnitude, + (magnitude, magnitude * 2), + ) encoded = len(owner._json(record)) # ",": one separator per record after the first. # header ids/groups: two household ids and two group indices per group. diff --git a/experiments/native-row-ceilings/regenerate_accepted_pin.py b/experiments/native-row-ceilings/regenerate_accepted_pin.py new file mode 100644 index 000000000..9001e0d33 --- /dev/null +++ b/experiments/native-row-ceilings/regenerate_accepted_pin.py @@ -0,0 +1,51 @@ +"""Regenerate acs_native_coverage_binding._ACCEPTED. Never hand-edit a pin. + +``_ACCEPTED`` pins four ACS modules by whole-file sha256 and refuses with +``UNREVIEWED_PREPARATION`` when one differs. Its generator is exactly +``coverage._sha(Path(__file__).with_name(name).read_bytes())``, which is +sha256 of the file's bytes -- the same call the check itself makes. This +rewrites each entry with that value and prints what moved. + + python regenerate_accepted_pin.py [--write] +""" + +from __future__ import annotations + +import hashlib +import pathlib +import re +import sys + +ROOT = pathlib.Path(__file__).resolve().parents[2] +US = ROOT / "packages/microcosm-build/src/microcosm/build/us_runtime" +OWNER = US / "acs_native_coverage_binding.py" + + +def main() -> int: + write = "--write" in sys.argv[1:] + text = OWNER.read_text() + block = re.search(r"_ACCEPTED = \{\n(.*?)\n\}\n", text, re.S) + if block is None: + raise SystemExit("could not locate the _ACCEPTED block") + body = block.group(1) + moved, replaced = [], body + for name, old in re.findall(r'^ "([^"]+)": "([0-9a-f]{64})",$', body, re.M): + new = hashlib.sha256((US / name).read_bytes()).hexdigest() + if new != old: + moved.append((name, old, new)) + replaced = replaced.replace(f'"{name}": "{old}"', f'"{name}": "{new}"') + for name, old, new in moved: + print(f"MOVED {name}\n old {old}\n new {new}") + if not moved: + print("no pin moved") + return 0 + if not write: + print("(dry run; pass --write to apply)") + return 0 + OWNER.write_text(text.replace(body, replaced)) + print(f"rewrote {OWNER.name}") + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/experiments/native-row-ceilings/regenerate_inventory_contract.py b/experiments/native-row-ceilings/regenerate_inventory_contract.py index 71232952b..27a68a388 100644 --- a/experiments/native-row-ceilings/regenerate_inventory_contract.py +++ b/experiments/native-row-ceilings/regenerate_inventory_contract.py @@ -18,7 +18,12 @@ ROOT = pathlib.Path(__file__).resolve().parents[2] sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] -from microcosm.build.us_runtime import graph_implementation as implementation # noqa: E402 +from microcosm.build.us_runtime import graph_implementation + +# The sys.path prelude above must win: regenerating a pin from some other +# checkout would write the wrong value into this one. +if not pathlib.Path(graph_implementation.__file__).resolve().is_relative_to(ROOT): + raise SystemExit(f"graph_implementation resolved outside {ROOT}") INVENTORY = ( ROOT @@ -31,13 +36,13 @@ def main() -> int: write = "--write" in sys.argv[1:] raw = INVENTORY.read_text() inventory = json.loads(raw) - roots = implementation._package_roots() + roots = graph_implementation._package_roots() moved = [] for name, expected in inventory["contracts"].items(): package, relative = name.split("/", 1) payload = (roots[package] / relative).read_bytes() - actual = implementation._dependency_contract( - payload, name, implementation._covered_imports(name, inventory) + actual = graph_implementation._dependency_contract( + payload, name, graph_implementation._covered_imports(name, inventory) ) if actual != expected: moved.append((name, expected, actual)) diff --git a/experiments/native-row-ceilings/repin.json b/experiments/native-row-ceilings/repin.json new file mode 100644 index 000000000..d10589c4c --- /dev/null +++ b/experiments/native-row-ceilings/repin.json @@ -0,0 +1,118 @@ +{ + "scope": "Pin re-derivation for the row-ceiling lane. Reads the working tree and rebuilds each pin through its own generator. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "edited_modules": { + "acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c", + "acs_person_coverage_columns.py": "ddcc0f009af6d0956d7ec8f116fa544da65b0ee31c111ee3991e72171d6e06c3", + "survey_observed_age.py": "675951305f1ddf5db86bd09045f0683f03103f2bfa2ced3ed502472f19272681", + "survey_origin_budget.py": "8da559ee32b469d96911b12158de9d95db4eb07992f7e0cc5396651fa2c968f7" + }, + "inventory_contracts": { + "generator": "graph_graph_implementation._dependency_contract(payload, name, graph_graph_implementation._covered_imports(name, inventory))", + "checked": 124, + "moved": [] + }, + "acs_native_coverage_binding_accepted": { + "generator": "hashlib.sha256(.read_bytes()).hexdigest()", + "entries": [ + { + "module": "acs_pums.py", + "old": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c", + "new": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c", + "moved": false, + "edited_by_this_branch": true + }, + { + "module": "acs_inputs.py", + "old": "aa4a8aeaba63dfef2f3e04fb89de59766deb088ed7f4d290aeba0425739916da", + "new": "aa4a8aeaba63dfef2f3e04fb89de59766deb088ed7f4d290aeba0425739916da", + "moved": false, + "edited_by_this_branch": false + }, + { + "module": "acs_housing_universe_source.py", + "old": "beb46a4a05a13580a868be423809a77946441dcafc3a0f57160561157e93e9a3", + "new": "beb46a4a05a13580a868be423809a77946441dcafc3a0f57160561157e93e9a3", + "moved": false, + "edited_by_this_branch": false + }, + { + "module": "acs_person_coverage_authentication.py", + "old": "9ec68721a4cf480ef412c51ab354db9000eb7a7989e35e6e09b1574d88e8e49f", + "new": "9ec68721a4cf480ef412c51ab354db9000eb7a7989e35e6e09b1574d88e8e49f", + "moved": false, + "edited_by_this_branch": false + } + ] + }, + "stage_manifests": { + "generator": "graph_graph_implementation.implementation_manifest(stage)", + "built": 10, + "stages": { + "acs_codec_2024": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 53, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c" + } + }, + "acs_housing_universe_2024": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 58, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c" + } + }, + "asec_codec_v4": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 42, + "edited_modules_in_stage": {} + }, + "asec_prepared_v3": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 83, + "edited_modules_in_stage": {} + }, + "assembly_harmonize": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 43, + "edited_modules_in_stage": {} + }, + "assembly_prepare": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 65, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c" + } + }, + "authenticated_survey_population_v1": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 92, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_person_coverage_columns.py": "ddcc0f009af6d0956d7ec8f116fa544da65b0ee31c111ee3991e72171d6e06c3", + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c", + "microcosm.build/us_runtime/survey_observed_age.py": "675951305f1ddf5db86bd09045f0683f03103f2bfa2ced3ed502472f19272681" + } + }, + "composed_asec_binding_v1": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 99, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c" + } + }, + "composed_population_v1": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 96, + "edited_modules_in_stage": { + "microcosm.build/us_runtime/acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c" + } + }, + "geography": { + "inventory_sha256": "2f98c788a247eb722de0fb4b533c0a183f5aecf16a10e964d8351244510ef58d", + "module_count": 41, + "edited_modules_in_stage": {} + } + } + } +} diff --git a/experiments/native-row-ceilings/repin.py b/experiments/native-row-ceilings/repin.py index eec20704a..561f943d9 100644 --- a/experiments/native-row-ceilings/repin.py +++ b/experiments/native-row-ceilings/repin.py @@ -4,12 +4,12 @@ 1. ``graph_implementation_inventory.json``'s per-module ``contracts`` entry -- ``imports``, ``unbound_uses_sha256`` and ``resource_accesses_sha256``, - re-derived through ``graph_implementation._dependency_contract`` against the + re-derived through ``graph_graph_implementation._dependency_contract`` against the module's own ``_covered_imports``, exactly as ``implementation_manifest`` checks it. A constant's value is not an import, an unbound use or a resource access, so these are expected to hold; the point is to prove it rather than assume it. -2. ``acs_native_coverage_binding._ACCEPTED``, which pins four ACS modules by +2. ``acs_native_coverage_acs_native_coverage_binding._ACCEPTED``, which pins four ACS modules by whole-file sha256. ``acs_pums.py`` is one of them, so that pin moves. 3. Each declared stage's ``implementation_manifest``, whose ``modules`` map is the same whole-file sha256 per inventoried module. These are computed, not @@ -32,10 +32,13 @@ ROOT = pathlib.Path(__file__).resolve().parents[2] sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] -from microcosm.build.us_runtime import ( # noqa: E402 - acs_native_coverage_binding as binding, -) -from microcosm.build.us_runtime import graph_implementation as implementation # noqa: E402 +from microcosm.build.us_runtime import acs_native_coverage_binding, graph_implementation + +# The sys.path prelude above must win: a pin re-derived from some other checkout +# would be worthless. Prove the modules resolved inside this tree. +for _module in (acs_native_coverage_binding, graph_implementation): + if not pathlib.Path(_module.__file__).resolve().is_relative_to(ROOT): + raise SystemExit(f"{_module.__name__} resolved outside {ROOT}") US = ROOT / "packages/microcosm-build/src/microcosm/build/us_runtime" EDITED = ( @@ -54,10 +57,10 @@ def main() -> int: moved, checked = [], 0 for name, expected in inventory["contracts"].items(): package, relative = name.split("/", 1) - roots = implementation._package_roots() + roots = graph_implementation._package_roots() payload = (roots[package] / relative).read_bytes() - actual = implementation._dependency_contract( - payload, name, implementation._covered_imports(name, inventory) + actual = graph_implementation._dependency_contract( + payload, name, graph_implementation._covered_imports(name, inventory) ) checked += 1 if actual != expected: @@ -65,7 +68,7 @@ def main() -> int: # 2. The ACS whole-file pin that names one of the edited modules. accepted = [] - for name, expected in binding._ACCEPTED.items(): + for name, expected in acs_native_coverage_binding._ACCEPTED.items(): actual = hashlib.sha256((US / name).read_bytes()).hexdigest() accepted.append( { @@ -79,8 +82,8 @@ def main() -> int: # 3. Every declared stage manifest, rebuilt through its generator. stages = {} - for stage in sorted(implementation.STAGE_DEPENDENCIES): - manifest = implementation.implementation_manifest(stage) + for stage in sorted(graph_implementation.STAGE_DEPENDENCIES): + manifest = graph_implementation.implementation_manifest(stage) stages[stage] = { "inventory_sha256": manifest["inventory_sha256"], "module_count": len(manifest["modules"]), @@ -104,8 +107,8 @@ def main() -> int: }, "inventory_contracts": { "generator": ( - "graph_implementation._dependency_contract(payload, name, " - "graph_implementation._covered_imports(name, inventory))" + "graph_graph_implementation._dependency_contract(payload, name, " + "graph_graph_implementation._covered_imports(name, inventory))" ), "checked": checked, "moved": moved, @@ -115,7 +118,7 @@ def main() -> int: "entries": accepted, }, "stage_manifests": { - "generator": "graph_implementation.implementation_manifest(stage)", + "generator": "graph_graph_implementation.implementation_manifest(stage)", "built": len(stages), "stages": stages, }, diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py index bc4ccc291..363cd5ca7 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py @@ -32,7 +32,7 @@ # Exact accepted direct AGEP -> A_AGE and AGEP -> age implementation, and owners. # A new transform/owner version requires explicit review of this successor. _ACCEPTED = { - "acs_pums.py": "6ecf79f0dfb0c0bc0ad0af6be5fa65bd8c2d1e1968009402e4c8dee347de70dc", + "acs_pums.py": "e79a2a4ecbc81e52b361918e7f391725a336274c74041387d71195cca46cb83c", "acs_inputs.py": "aa4a8aeaba63dfef2f3e04fb89de59766deb088ed7f4d290aeba0425739916da", "acs_housing_universe_source.py": "beb46a4a05a13580a868be423809a77946441dcafc3a0f57160561157e93e9a3", "acs_person_coverage_authentication.py": "9ec68721a4cf480ef412c51ab354db9000eb7a7989e35e6e09b1574d88e8e49f", diff --git a/pyproject.toml b/pyproject.toml index 5589dcfd6..2537e24bc 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -35,6 +35,13 @@ ignore = ["E501"] "experiments/native-scale-transport/harness19_base.py" = ["F401"] "experiments/native-scale-transport/harness19_tenth.py" = ["F401"] "experiments/native-scale-transport/harness19_after_with_required_replay.py" = ["F401"] +# The row-ceiling lane's pin tools put the worktree's own packages on +# sys.path before importing them, and assert beneath the import block that +# each module resolved inside this tree -- a pin re-derived from another +# checkout would be the wrong value. That ordering requires E402. +"experiments/native-row-ceilings/repin.py" = ["E402"] +"experiments/native-row-ceilings/regenerate_inventory_contract.py" = ["E402"] +"experiments/native-row-ceilings/origin_budget_size.py" = ["E402"] [tool.ruff.lint.isort] # microcosm is a PEP 420 namespace package (no top-level __init__.py), so From 2b25f95a31ac3639e481ab9d3b3725ecef5eccd5 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:53:29 -0400 Subject: [PATCH 08/22] Test every moved ceiling: its own boundary, and a full-source-sized count One new file carries the rule as an assertion -- the measured full-source counts, and for each moved ceiling that it is exactly four times its count rounded up to the next whole million. It also pins the bounds that must not move, with the reason each stays: the ACS source-file row bounds assert a file this build does not produce, the ASEC codec bounds count a source no fraction grows, 2**53 bounds a value not a row count, and the origin budget's byte transport takes the segmented-transport argument. Per-module boundary tests prove each refusal fires at exactly its own constant, so acceptance at full source follows from the two together without any test allocating a full-source roster: acs_pums.MAX_EXACT_HOUSEHOLDS patched to 3; 3 accepted, 4 and () refuse acs_pums.MAX_EXACT_PERSON_ROWS NP is a declared count, so a two-row household table declaring the whole 3,422,888-person ACS file drives the shipped ceiling for free -- it runs past and refuses later on the person archive; patched one below, it refuses there MAX_SELECTED_ROWS patched to 2; 2 proceeds to the archive, 3 refuses survey_observed_age.MAX_ROWS a real 3,422,888-row normalize, 0.02s and 27 MiB survey_origin_budget.MAX_GROUPS six accepted, five refuses GROUP_COUNT_BOUND. MAX_GROUPS is inside the module's loaded contract, so the test re-seals _LIVE beside the patch -- what a module shipped with a different number looks like -- leaving every other producer check in force. Co-Authored-By: Claude Opus 5 --- .../test_us_acs_person_coverage_columns.py | 20 +++ .../microcosm-build/tests/test_us_acs_pums.py | 60 +++++++ .../tests/test_us_native_row_ceilings.py | 152 ++++++++++++++++++ .../tests/test_us_survey_observed_age.py | 27 ++++ .../tests/test_us_survey_origin_budget.py | 27 ++++ 5 files changed, 286 insertions(+) create mode 100644 packages/microcosm-build/tests/test_us_native_row_ceilings.py diff --git a/packages/microcosm-build/tests/test_us_acs_person_coverage_columns.py b/packages/microcosm-build/tests/test_us_acs_person_coverage_columns.py index 5ac09804c..72403c3c3 100644 --- a/packages/microcosm-build/tests/test_us_acs_person_coverage_columns.py +++ b/packages/microcosm-build/tests/test_us_acs_person_coverage_columns.py @@ -494,3 +494,23 @@ def read(self, *args, **kwargs): ) assert opened == ["psam_pusa.csv"] assert sum(consumed) < 32_768 + + +def test_selected_row_ceiling_refuses_at_its_own_number(tmp_path, monkeypatch): + """The refusal is `<= MAX_SELECTED_ROWS`, whatever that number is. + + At exactly the ceiling the read proceeds and fails later, on the absent + archive it then opens. `test_us_native_row_ceilings.py` carries the separate + assertion that the shipped number admits the 3,422,888 persons a full-source + ACS selection requests. + """ + source = acs_pums.AcsPumsSource( + tmp_path / "absent-h.zip", tmp_path / "absent-p.zip" + ) + monkeypatch.setattr(module, "MAX_SELECTED_ROWS", 2) + at_ceiling = _keys([_row(), _row(order="2")]) + with pytest.raises(FileNotFoundError): + module.read_acs_person_coverage_columns(source, person_keys=at_ceiling) + over = _keys([_row(), _row(order="2"), _row(order="3")]) + with pytest.raises(ValueError, match="selected person count is outside the bound"): + module.read_acs_person_coverage_columns(source, person_keys=over) diff --git a/packages/microcosm-build/tests/test_us_acs_pums.py b/packages/microcosm-build/tests/test_us_acs_pums.py index d663d21ac..1dc4edb28 100644 --- a/packages/microcosm-build/tests/test_us_acs_pums.py +++ b/packages/microcosm-build/tests/test_us_acs_pums.py @@ -770,3 +770,63 @@ def test_household_axis_composition_refuses_misaligned_weights() -> None: with pytest.raises(ValueError, match="one DESIGN weight per loaded"): acs_pums._household_axis_composition(household, weights) + + +def _serialnos(count: int) -> tuple[str, ...]: + return tuple(f"2024HU{index:07d}" for index in range(1, count + 1)) + + +def test_exact_household_ceiling_refuses_at_its_own_number(monkeypatch) -> None: + """The refusal is `<= MAX_EXACT_HOUSEHOLDS`, whatever that number is. + + Driven at a patched-down ceiling so the boundary is exercised without + building a full-source key tuple; `test_us_native_row_ceilings.py` carries + the separate assertion that the shipped number admits the 1,531,614 + households a full-source ACS selection supplies. + """ + monkeypatch.setattr(acs_pums, "MAX_EXACT_HOUSEHOLDS", 3) + accepted = _serialnos(3) + assert AcsPumsSource.snapshot_serialnos(accepted) == accepted + with pytest.raises(ValueError, match="bounded unique raw native keys"): + AcsPumsSource.snapshot_serialnos(_serialnos(4)) + with pytest.raises(ValueError, match="bounded unique raw native keys"): + AcsPumsSource.snapshot_serialnos(()) + + +def test_exact_person_row_ceiling_counts_np_not_person_records( + tmp_path: Path, monkeypatch +) -> None: + """NP is a declared count, so the ceiling is driven at full-source scale free. + + The check sums NP over the selected households before the person archive is + opened, so a two-row household table declaring 3,422,888 people -- the whole + 2024 ACS person file -- exercises the shipped ceiling without materialising + a single person row. + """ + household_zip = tmp_path / "h.zip" + person_zip = tmp_path / "p.zip" + serialnos = ("2024HU0000001", "2024HU0000002") + _write_csv_zip( + household_zip, + { + "psam_hus.csv": [ + _household(serialnos[0], NP=3_422_887), + _household(serialnos[1], NP=1), + ] + }, + ) + _write_csv_zip(person_zip, {"psam_pus.csv": [_person(serialnos[0], 1, 20)]}) + source = AcsPumsSource(household_zip, person_zip) + + # The shipped ceiling admits the whole 2024 ACS person file: the load runs + # past this check and refuses later, on the person archive it then opens. + assert acs_pums.MAX_EXACT_PERSON_ROWS > 3_422_888 + with pytest.raises(ValueError, match="NP/person row-count mismatch"): + load_acs_pums_tables(source, serialnos=serialnos) + + # Patched below that roster, the same expression refuses it. + monkeypatch.setattr(acs_pums, "MAX_EXACT_PERSON_ROWS", 3_422_887) + with pytest.raises( + ValueError, match="selected complete roster exceeds native person budget" + ): + load_acs_pums_tables(source, serialnos=serialnos) diff --git a/packages/microcosm-build/tests/test_us_native_row_ceilings.py b/packages/microcosm-build/tests/test_us_native_row_ceilings.py new file mode 100644 index 000000000..f3bc56d38 --- /dev/null +++ b/packages/microcosm-build/tests/test_us_native_row_ceilings.py @@ -0,0 +1,152 @@ +"""The row-ceiling rule, as an assertion rather than only a paragraph. + +`docs/us-native-row-ceilings.md` states one rule: a bound moves only if a +full-source native build meets it, and a bound that moves becomes four times the +measured full-source count of exactly what it counts, rounded up to the next +whole million. This file holds the measured counts and checks every moved +ceiling against them, so a later edit that drifts from the rule fails here +rather than in a build. + +The counts are measurements, not extrapolations. They come from the catalogues +carried inside the recovered 1/1000 pilot preparation artifact +(sha256 34b362d8...), whose ACS occupied, institutional-GQ and +noninstitutional-GQ households plus its ASEC households equal that artifact's +own ``selection.supplied_households`` exactly -- so "the whole catalogue" and +"what a full-source selection supplies" are the same set. +``experiments/native-row-ceilings/roster_census.py`` re-derives them and asserts +that reconciliation; ``roster-census.json`` is its output. + +Nothing here reads a source, allocates a full-source roster or runs a build. +""" + +from __future__ import annotations + +import pytest + +from microcosm.build.us_runtime import ( + acs_native_coverage_binding, + acs_person_coverage_columns, + acs_pums, + asec_current_money, + survey_observed_age, + survey_origin_budget, +) + +# Measured full-source counts. See the module docstring for the derivation. +ACS_HOUSEHOLDS = 1_531_614 +ACS_PERSONS = 3_422_888 +ASEC_HOUSEHOLDS = 55_762 +ASEC_PERSONS = 142_125 +STACKED_HOUSEHOLDS = 1_587_376 +STACKED_PERSONS = 3_565_013 +COMBINED_CLONE_PERSONS = 7_130_026 + +HEADROOM = 4 + +# Each moved ceiling, with the measured full-source count of exactly what it +# counts. The third element is what the rule produces from the second. +MOVED = ( + (acs_pums, "MAX_EXACT_HOUSEHOLDS", ACS_HOUSEHOLDS, 7_000_000), + (acs_pums, "MAX_EXACT_PERSON_ROWS", ACS_PERSONS, 14_000_000), + (acs_person_coverage_columns, "MAX_SELECTED_ROWS", ACS_PERSONS, 14_000_000), + (survey_observed_age, "MAX_ROWS", ACS_PERSONS, 14_000_000), + (survey_origin_budget, "MAX_GROUPS", STACKED_HOUSEHOLDS, 7_000_000), +) + + +def _next_whole_million(value: int) -> int: + return -(-value // 1_000_000) * 1_000_000 + + +@pytest.mark.parametrize( + "module, name, measured, expected", + MOVED, + ids=[f"{m.__name__.rsplit('.', 1)[-1]}.{n}" for m, n, _, _ in MOVED], +) +def test_every_moved_ceiling_is_exactly_what_the_rule_produces( + module, name, measured, expected +): + """Four times the measured count, rounded up to the next whole million.""" + assert expected == _next_whole_million(HEADROOM * measured) + assert getattr(module, name) == expected + + +@pytest.mark.parametrize( + "module, name, measured", + [(m, n, c) for m, n, c, _ in MOVED], + ids=[f"{m.__name__.rsplit('.', 1)[-1]}.{n}" for m, n, _, _ in MOVED], +) +def test_every_moved_ceiling_admits_a_full_source_count(module, name, measured): + """A full-source count falls inside the accepted region, with headroom. + + The per-module boundary tests prove each refusal fires at exactly its own + constant; together with this, a full-source-sized count is accepted without + any test allocating a full-source roster. + """ + ceiling = getattr(module, name) + assert ceiling >= HEADROOM * measured + assert measured < ceiling + + +def test_the_acs_source_file_bound_does_not_move(): + """MAX_ROWS asserts the source file's own size, so it is not this rule's. + + The 2024 ACS person file holds 3,422,888 records. Raising this would weaken + a real structural check on a file this build does not produce, and the + selected roster it feeds is bounded separately by MAX_SELECTED_ROWS. + """ + assert acs_person_coverage_columns.MAX_ROWS == 6_000_000 + assert acs_person_coverage_columns.MAX_ROWS > ACS_PERSONS + assert acs_native_coverage_binding.MAX_SOURCE_ROWS == 6_000_000 + assert acs_native_coverage_binding.MAX_SOURCE_ROWS > ACS_PERSONS + + +def test_the_asec_codec_bounds_do_not_move_because_they_never_bind(): + """These count the ASEC source's own rows, which no selection fraction grows. + + The transport lane's census listed asec_current_money.MAX_PERSONS among the + bounds a full-source build meets. It does not: person_rows and + household_rows are the ASEC source scope's own sizes, and the catalogue + measures them at 142,125 persons in 55,762 households. + """ + assert asec_current_money.MAX_PERSONS == 1_000_000 + assert asec_current_money.MAX_HOUSEHOLDS == 400_000 + assert ASEC_PERSONS < asec_current_money.MAX_PERSONS + assert ASEC_HOUSEHOLDS < asec_current_money.MAX_HOUSEHOLDS + + +def test_the_float64_representation_limit_is_not_a_row_ceiling(): + """It bounds an observed age's value, not how many rows carry one.""" + assert survey_observed_age.MAX_EXACT_FLOAT64_INTEGER == 2**53 + + +def test_the_origin_budget_byte_transport_is_left_for_the_other_argument(): + """Deliberately unmoved, and the loudest thing this lane found. + + survey_origin_budget streams one origin record per allocation group into a + single bytearray under MAX_PAYLOAD_BYTES. Measured through the module's own + encoder at full-source household-id widths, a record plus its two header + entries costs 764 bytes, so 64 MiB admits 87,838 households -- 5.53% of + source, below one tenth, and below the 96,860-household preparation-receipt + ceiling the transport lane lifted. A full-source payload is 1.13 GiB. + + That is a byte transport, and the answer to a byte transport is the + segmented stream `survey_population_preparation` already carries, not a + larger single cap. It is a separate change with the transport lane's + argument, so this lane lifts MAX_GROUPS and leaves this where it is. Lifting + MAX_GROUPS alone is necessary and not sufficient: at full source the refusal + moves from GROUP_COUNT_BOUND to TRANSPORT_LIMIT. + + See experiments/native-row-ceilings/origin-budget-size.json. + """ + assert survey_origin_budget.MAX_PAYLOAD_BYTES == 64 * 1024**2 + admitted = survey_origin_budget.MAX_PAYLOAD_BYTES // 764 + assert admitted < STACKED_HOUSEHOLDS // 10 + assert survey_origin_budget.MAX_GROUPS > STACKED_HOUSEHOLDS + + +def test_the_measured_counts_reconcile(): + """The arithmetic the catalogue derivation rests on, kept honest here.""" + assert ACS_HOUSEHOLDS + ASEC_HOUSEHOLDS == STACKED_HOUSEHOLDS + assert ACS_PERSONS + ASEC_PERSONS == STACKED_PERSONS + assert STACKED_PERSONS * 2 == COMBINED_CLONE_PERSONS diff --git a/packages/microcosm-build/tests/test_us_survey_observed_age.py b/packages/microcosm-build/tests/test_us_survey_observed_age.py index 638603ac5..643a0fc66 100644 --- a/packages/microcosm-build/tests/test_us_survey_observed_age.py +++ b/packages/microcosm-build/tests/test_us_survey_observed_age.py @@ -118,3 +118,30 @@ def test_known_negative_zero_and_rule_document_are_preserved(): assert document["top_code_replacement"] is False document["relation"] = "changed" assert rule.rule_document()["relation"] == "numeric_identity" + + +def test_row_ceiling_admits_a_full_source_channel(): + """A full-source ACS channel really is normalized, not only asserted. + + `_normalized_source_copy` normalizes the ACS and ASEC native frames + separately, so the ceiling is met one channel at a time and the larger + channel is ACS: 3,422,888 persons at full source. That is one int64 column, + about 27 MiB, so the acceptance is driven for real rather than inferred. + """ + rows = 3_422_888 + assert rule.MAX_ROWS >= rows + raw = pd.Series(np.arange(rows, dtype="int64") % 101, name="A_AGE") + result = rule.normalize_observed_age(raw) + assert len(result) == rows + assert result.dtype == np.dtype("float64") + assert result.iloc[0] == 0.0 and result.iloc[100] == 100.0 + assert result.name == "age" + + +def test_row_ceiling_refuses_one_row_past_its_own_number(monkeypatch): + """The refusal is `<= MAX_ROWS`, whatever that number is.""" + monkeypatch.setattr(rule, "MAX_ROWS", 3) + at_ceiling = pd.Series([1, 2, 3], dtype="int64") + assert rule.normalize_observed_age(at_ceiling).tolist() == [1.0, 2.0, 3.0] + with pytest.raises(ValueError, match="ROW_BOUND"): + rule.normalize_observed_age(pd.Series([1, 2, 3, 4], dtype="int64")) diff --git a/packages/microcosm-build/tests/test_us_survey_origin_budget.py b/packages/microcosm-build/tests/test_us_survey_origin_budget.py index cc7c45817..38952bd4b 100644 --- a/packages/microcosm-build/tests/test_us_survey_origin_budget.py +++ b/packages/microcosm-build/tests/test_us_survey_origin_budget.py @@ -560,3 +560,30 @@ def restore(): assert owner.source._ISSUED[id(preparation)] is original _assert_final_mutation_refused(owner._initial, budget.checked_view, mutate, restore) + + +def test_group_ceiling_refuses_at_its_own_number(tmp_path, monkeypatch): + """The refusal is `<= MAX_GROUPS`, whatever that number is. + + One group per allocation instruction, and `allocation_instructions` requires + one instruction per selected household, so at full source the count is the + 1,587,376 households the catalogues supply. The invented fixture has six, so + the boundary is driven there; `test_us_native_row_ceilings.py` carries the + separate assertion that the shipped number admits a full-source selection. + + `MAX_GROUPS` is inside the module's own loaded contract, so a changed value + is PRODUCER_CHANGED before it can be GROUP_COUNT_BOUND. Re-sealing `_LIVE` + alongside the patch is what a module genuinely shipped with a different + number looks like, and leaves every other producer check in force. + """ + arguments = _actual(tmp_path, monkeypatch) + + monkeypatch.setattr(owner, "MAX_GROUPS", 6) + monkeypatch.setattr(owner, "_LIVE", owner._live()) + budget = owner.freeze_survey_origin_budget(**arguments) + assert budget.checked_view().document["group_count"] == 6 + + monkeypatch.setattr(owner, "MAX_GROUPS", 5) + monkeypatch.setattr(owner, "_LIVE", owner._live()) + with pytest.raises(owner.SurveyOriginBudgetError, match="GROUP_COUNT_BOUND"): + owner.freeze_survey_origin_budget(**arguments) From cb5a62cf05f92040a2e0ee03cfcb2406ea5c8d76 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:55:21 -0400 Subject: [PATCH 09/22] Record the rule, the measured counts, and what the census turned up Co-Authored-By: Claude Opus 5 --- PROGRESS-native-row-ceilings.md | 86 +++++++++++++++++++++++---------- 1 file changed, 61 insertions(+), 25 deletions(-) diff --git a/PROGRESS-native-row-ceilings.md b/PROGRESS-native-row-ceilings.md index 476a10c2e..74807bd59 100644 --- a/PROGRESS-native-row-ceilings.md +++ b/PROGRESS-native-row-ceilings.md @@ -7,39 +7,75 @@ This file is a session-handoff journal. Per the repo's CLAUDE.md, its "State"/"Next" sections are accurate when written and historical afterward — check git and GitHub for current truth. -## The brief +## The rule (the one argument) -The transport lane's census (`docs/us-native-scale-transport.md` §5) found that -after the 64 MiB transport ceilings were lifted, a full-source native US build -(1,587,376 source households; ~3,471,000 stacked and ~6,943,000 cloned persons) -still meets at least four row-count bounds: +**A bound moves only if a full-source native build meets it, and a bound that +moves becomes four times the measured full-source count of exactly what it +counts, rounded up to the next whole million.** -- `acs_person_coverage_columns.MAX_SELECTED_ROWS` — 1,000,000 -- `asec_current_money.MAX_PERSONS` — 1,000,000 -- `acs_pums.MAX_EXACT_HOUSEHOLDS` / `MAX_EXACT_PERSON_ROWS` — 1,000,000 -- `survey_observed_age.MAX_ROWS` — 2,000,000 +Byte transports are *not* this rule's. The transport lane's answer to a byte +ceiling is a segmented stream under `MAX_ROSTER_BYTES`, not a larger single cap, +and applying this rule to one would be the wrong argument. -and that `survey_origin_budget.MAX_GROUPS` (1,000,000) was not established. -Max decided (2026-09-17) they are lifted as **one change with one argument**, -not one build at a time. +## The counts are measured, not extrapolated -## State +The recovered 1/1000 pilot artifact +(`_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json`, +sha256 `34b362d8…`) carries the whole ACS and ASEC **catalogues**, which do not +scale with the selection fraction. Its ACS occupied + institutional-GQ + +noninstitutional-GQ households plus its ASEC households equal its own +`selection.supplied_households` exactly (1,348,408 + 84,422 + 98,784 + 55,762 = +1,587,376), so "the whole catalogue" and "what a full-source selection supplies" +are the same set: -Census in progress. Nothing implemented yet. +| | households | persons | +|---|---|---| +| ACS | 1,531,614 | 3,422,888 | +| ASEC | 55,762 | 142,125 | +| stacked | 1,587,376 | 3,565,013 | +| combined clone | 3,174,752 | 7,130,026 | + +The transport lane's ~3,471,000 stacked / ~6,943,000 cloned were the 1/1000 +per-household ratio extrapolated; these supersede them. +`experiments/native-row-ceilings/roster_census.py` re-derives them and asserts +the reconciliation. ## Done -- Worktree verified at `a64f7b733`; `uv sync --all-packages --locked --extra us` exit 0. -- Read the transport lane's §2c/§2e (the `MAX_ROSTER_BYTES` argument shape: - an explicit resource ceiling at a fixed multiple above the full-source count) - and its §4 (how it re-derived pins through `graph_implementation._dependency_contract`). +- Census of every `MAX_*` bound in `us_runtime/` reachable from the two graph + entry points: 146 bounds across five module families, each with its + enforcement site, refusal code, what it protects, and its counts at 1/10 and + full source. **14 bind at full source; 5 of those bind at 1/10 too.** +- `survey_origin_budget.MAX_GROUPS` **established**: `allocation_instructions` + requires one instruction per selected household, so a full-source budget has + 1,587,376 groups and the 1,000,000 bound binds. +- Five constants lifted under the rule (commit `14defbfc0`). +- Pins re-derived through their generators (`2ebd246f1`): one moved, + `acs_native_coverage_binding._ACCEPTED["acs_pums.py"]`. +- **An inherited break re-pinned** (`55ca820c7`): base-branch commit `b6081efcb` + added `path.read_bytes()` to `_spill_roster` without regenerating + `survey_population_preparation.py`'s `resource_accesses_sha256`, so + `implementation_manifest()` raised for every stage containing it — including + the one the nineteen-node path runs. Not this branch's file; re-pinned here so + the base is functional and this lane's own manifests can be built. +- Tests: the rule as an executable table, plus a boundary test per moved bound. + +## The loudest finding + +`survey_origin_budget.MAX_PAYLOAD_BYTES` (64 MiB) admits **87,838 households, +5.53% of source** — below 1/10, and below the 96,860-household ceiling the +transport lane lifted. Measured through the module's own encoder at full-source +id widths: 764 B per group, 1.13 GiB at full source, 18.07× the cap. It is a +byte transport, so this lane lifts `MAX_GROUPS` and leaves it: at full source the +refusal moves from `GROUP_COUNT_BOUND` to `TRANSPORT_LIMIT`. Necessary, not +sufficient, and the report says so. ## Next -1. Census every `MAX_*` bound in `us_runtime/` reachable from the 19-node - financial graph and the 45-node pilot graph, from code at this head. -2. Establish `survey_origin_budget.MAX_GROUPS`. -3. Write `docs/us-native-row-ceilings.md` — the one argument. -4. Implement; a test per moved bound. -5. Re-derive every moved pin through its generator. -6. Draft PR against `native-scale-transport`. +1. Finish the adversarial verification of the 14 binding verdicts. +2. Decide, on that evidence, whether the two pure row-count bounds the transport + lane's census missed (`asec_demographic_source._MAX_PERSONS`, + `current_child_property_income_source.MAX_ROWS`, both 600,000) move here. +3. `docs/us-native-row-ceilings.md`. +4. Tests as CI runs them; `ci_test_groups --verify`; `spec_engine_coverage --check`. +5. Draft PR against `native-scale-transport`. From 907e308c3ec2ec1b0ba7b94a35211484d9418d7a Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 17:59:10 -0400 Subject: [PATCH 10/22] Write the one argument down docs/us-native-row-ceilings.md: the rule and why each of its three clauses is load-bearing; the measured counts and why the artifact's catalogues answer the full-source question where its rosters would only estimate it; the five moved constants with their refusal codes unchanged; why a bound stays a bound rather than being deleted; the three classes that must not move at all (an upstream file's real size, a fixed-width encoding, a quantity no fraction grows); and what still binds and whose argument it is. Co-Authored-By: Claude Opus 5 --- docs/us-native-row-ceilings.md | 242 +++++++++++++++++++++++++++++++++ 1 file changed, 242 insertions(+) create mode 100644 docs/us-native-row-ceilings.md diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md new file mode 100644 index 000000000..b522d2d85 --- /dev/null +++ b/docs/us-native-row-ceilings.md @@ -0,0 +1,242 @@ +# US native build: the row-count ceilings, and the rule their new values follow + +The [transport lane](us-native-scale-transport.md) lifted the 64 MiB +whole-roster receipt ceilings and, in its §5, recorded a second family it did +not touch: row-count bounds that a full-source native US build meets. This note +is that family's argument. + +It is one rule, not five decisions. Max decided (2026-09-17) that these are +lifted as one change with one argument rather than one build at a time, and a +rule is the only thing that makes the sixth case decide itself. + +## 1. The rule + +> **A bound moves only if a full-source native build meets it. A bound that moves +> becomes four times the measured full-source count of exactly what it counts, +> rounded up to the next whole million.** + +Three clauses, each load-bearing. + +**"only if a full-source build meets it."** Most of the bounds censused in §3 do +not bind, and most of those never will, because what they count is an upstream +file's own size rather than a roster this build chooses. A rule that raised +everything uniformly would weaken real checks for no gain. This one leaves them +alone by construction, and §4 says why each must stay. + +**"of exactly what it counts."** The commonest way to get one of these wrong is +to bound the wrong roster. The 19-node path carries at least five distinct +quantities that differ by more than an order of magnitude — the upstream ACS +file's rows, the selected ACS roster, the selected ASEC roster, the stacked +roster, and the combined clone — and a ceiling raised against the wrong one is +either dead or still binding. §2 measures each, and every moved constant carries +a comment naming the one it counts. + +**"four times … rounded up to the next whole million."** The multiple is the +transport lane's: `MAX_ROSTER_BYTES` is `64 × MAX_SEGMENT_BYTES` = 4 GiB, which +that lane measured at 3.9× a full-source preparation receipt. Adopting the same +headroom keeps one law across both families. The rounding makes the constants +legible and the rule checkable at a glance: divide any moved constant by the +count its comment names and the answer is between 4 and 4.6. + +The rule is not only written here. `test_us_native_row_ceilings.py` holds the +measured counts and asserts that each moved ceiling is exactly what the rule +produces from its count, so a later edit that drifts fails a test rather than a +build. + +### 1a. What the rule is not for + +**Byte transports take the transport lane's argument, not this one.** That +lane's answer to a payload that outgrows its cap is a segmented stream under one +explicit total, not a larger single cap, and it proved the segmented bytes +identical to the unsegmented ones. Raising such a cap here would contradict the +neighbouring argument rather than extend it. §5 names the byte ceilings that +bind and leaves them to it — including one that binds harder than anything this +lane moved. + +## 2. The counts, measured rather than extrapolated + +Every number below comes from the recovered 1/1000 pilot preparation artifact +(`_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json`, +sha256 `34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0`), and +specifically from its **catalogues** rather than its rosters. + +That distinction is the point. A roster at 1/1000 has to be scaled to say +anything about full source, and scaling a stratified sample introduces error. +The catalogues do not: they are counts of the whole upstream ACS and ASEC files, +and they are identical in a 1/1000 artifact and a full-source one. What makes +them answer the full-source question is this identity, which +`experiments/native-row-ceilings/roster_census.py` asserts before reporting +anything: + +``` +ACS occupied_hu 1,348,408 + + institutional_gq 84,422 + + noninstitutional_gq 98,784 + + ASEC households 55,762 + = 1,587,376 == selection.supplied_households +``` + +"The whole catalogue" and "what a full-source selection supplies" are therefore +the same set, and the catalogue's counts *are* the full-source counts: + +| quantity | full source | at 1/10 | +|---|---:|---:| +| ACS households (selectable; excludes 100,355 vacancies) | 1,531,614 | 153,161 | +| ACS persons | 3,422,888 | 342,288 | +| ASEC households | 55,762 | 5,576 | +| ASEC persons | 142,125 | 14,212 | +| **stacked households** | **1,587,376** | 158,737 | +| **stacked persons** | **3,565,013** | 356,501 | +| combined-clone households | 3,174,752 | 317,475 | +| combined-clone persons | 7,130,026 | 713,002 | + +The transport lane's §5 quoted ~3,471,000 stacked and ~6,943,000 cloned persons. +Those were the 1/1000 artifact's 2.186869 persons per household extrapolated, +and the catalogue figures supersede them — about 2.7% higher, which changes no +verdict but is the number each new constant is derived from. + +One further split matters, because two of the moved bounds are met by one +channel rather than the total: **`_normalized_source_copy` normalizes the ACS +and ASEC native frames separately**, and the ACS channel carries 96.0% of the +persons. A bound met per channel is met by 3,422,888, not by 3,565,013. + +## 3. The moved ceilings + +| constant | was | now | counts | full source | × | +|---|---:|---:|---|---:|---:| +| `acs_pums.MAX_EXACT_HOUSEHOLDS` | 1,000,000 | **7,000,000** | exact ACS household keys in one snapshot | 1,531,614 | 4.57 | +| `acs_pums.MAX_EXACT_PERSON_ROWS` | 1,000,000 | **14,000,000** | `NP` summed over the selected ACS households | 3,422,888 | 4.09 | +| `acs_person_coverage_columns.MAX_SELECTED_ROWS` | 1,000,000 | **14,000,000** | requested ACS person keys | 3,422,888 | 4.09 | +| `survey_observed_age.MAX_ROWS` | 2,000,000 | **14,000,000** | one channel's person rows | 3,422,888 | 4.09 | +| `survey_origin_budget.MAX_GROUPS` | 1,000,000 | **7,000,000** | allocation instructions = selected households | 1,587,376 | 4.41 | + +Each refusal keeps its code, its exception type and its expression; only the +number moves. + +| constant | refusal | raises | +|---|---|---| +| `MAX_EXACT_HOUSEHOLDS` | `"ACS exact selection requires bounded unique raw native keys."` | `ValueError` | +| `MAX_EXACT_PERSON_ROWS` | `"ACS selected complete roster exceeds native person budget."` | `ValueError` | +| `MAX_SELECTED_ROWS` | `"ACS coverage selected person count is outside the bound"` | `ValueError` | +| `survey_observed_age.MAX_ROWS` | `"SURVEY_OBSERVED_AGE_ROW_BOUND"` | `ValueError` | +| `MAX_GROUPS` | `"GROUP_COUNT_BOUND"` | `SurveyOriginBudgetError` | + +### 3a. `survey_origin_budget.MAX_GROUPS`, which the transport lane left open + +Its §5 recorded this one as "not established — the group count was not derived +here". It is established now, and it binds. `_initial` builds its groups with + +```python +instructions = graph.allocation_instructions( + view.selection_plan, view.receipt["origins"]["households"] +) +_require(0 < len(instructions) <= MAX_GROUPS, "GROUP_COUNT_BOUND") +``` + +and `allocation_instructions` opens with +`_require(len(household_origins) == len(plan.selected), "ORIGIN_COUNT")`. One +instruction per selected household, so the group count *is* the selected +household count: 1,587,376 at full source against a 1,000,000 bound. + +Two things about where it runs, stated rather than implied. It is reachable by +import from `graph_atomic_survey_financial`, but **no node in the nineteen-node +graph executes it** — the recovered run's `graph.json` lists all nineteen and +none is a budget node — and none in the completion host does either. +`survey_age_calibration` and `graph_survey_budget` execute it, over the same +full-source selection, so the count and the verdict are unchanged. + +## 4. Why a bound stays a bound, and why some must not move at all + +**Deleting a bound is not the cheap version of raising it.** Each of these +stands between a malformed or runaway input and an unbounded allocation, and the +refusal is what makes the failure legible and early rather than an OOM kill deep +in a multi-hour build with no statement of which roster was wrong. A ceiling at +4× the real thing still refuses a 10× runaway, which is the shape a real mistake +takes: a fraction argument off by a decimal place, a selection that forgot to +filter, a key list built from the wrong catalogue. + +Three classes must not move at all, and this note names them so a later reader +does not mistake a low number for an oversight. + +**An upstream file's real size.** `acs_person_coverage_columns.MAX_ROWS` and +`acs_native_coverage_binding.MAX_SOURCE_ROWS`, both 6,000,000, bound how many +records the reader will stream out of the ACS archive. The 2024 ACS person file +holds 3,422,888 of them. That bound is a structural assertion about a file this +build does not produce and cannot check any other way, and the roster it feeds +is separately bounded by `MAX_SELECTED_ROWS`. Raising it would trade a real +check for nothing: a selection cannot exceed the source it is drawn from. + +**A fixed-width encoding.** The `PAYLOAD_MAX_BYTES = len(MAGIC) + 4 + +HEADER_MAX_BYTES + ROW_BYTES * MAX_HOUSEHOLDS + 32` family in +`asec_housing_status`, `asec_housing_universe` and `asec_income_observations` +computes a byte budget *from* a row ceiling, so moving the row ceiling moves the +wire format a reader depends on. `survey_observed_age.MAX_EXACT_FLOAT64_INTEGER` += 2**53 is the same kind of thing one level down: the module's own docstring +calls it "a representation limit, not a scientifically permitted maximum age", +and it bounds an age's value, not how many rows carry one. + +**A quantity no selection fraction grows.** The ASEC codec's +`asec_current_money.MAX_PERSONS` (1,000,000) and `MAX_HOUSEHOLDS` (400,000) +count the ASEC source scope's own `person_rows` and `household_rows`. The +catalogue measures those at 142,125 persons in 55,762 households, and a +full-source survey selection does not change them. **The transport lane's §5 +listed `asec_current_money.MAX_PERSONS` among the bounds a full-source build +meets; it does not, and this note corrects that.** + +## 5. What still binds, and whose argument it is + +The census in §6 found fourteen bounds a full-source build meets. Five are the +row counts §3 moved. The rest are byte transports and byte-derived row +pre-checks, and they belong to the transport lane's argument — a segmented +stream under one explicit total — not to this one. The loudest is in a module +this lane did touch, so it gets said plainly: + +> **`survey_origin_budget.MAX_PAYLOAD_BYTES` (64 MiB) admits 87,838 households — +> 5.53% of source.** That is below one tenth, and below the 96,860-household +> preparation-receipt ceiling the transport lane lifted. A full-source payload is +> 1.13 GiB, 18.07× the cap. + +Measured, not estimated: `experiments/native-row-ceilings/origin_budget_size.py` +builds one faithful origin record through the module's own `_reference` and +`_json` at full-source household-id widths and gets 764 bytes per group, +including the two `household_ids` and two `group_indices` entries the header +carries for each group's two clone roles. + +The consequence for this lane is stated rather than glossed: **lifting +`MAX_GROUPS` is necessary and not sufficient.** `GROUP_COUNT_BOUND` is checked +in `_initial` before the payload is built, so at full source it fired first; with +it lifted, the refusal moves to `TRANSPORT_LIMIT` in the same module. That is the +correct outcome of one argument applied to one class of bound, and the remaining +ceiling is a real piece of work for whoever carries the segmented transport into +this module. + +## 6. The census + +`experiments/native-row-ceilings/` carries the receipts. Every `MAX_*` row, +household, person, group or byte bound in +`packages/microcosm-build/src/microcosm/build/us_runtime/` reachable from the +19-node financial graph and the 45-node pilot graph was read at this head — 146 +bounds across five module families — with its enforcement site, refusal code, +what it protects, and its counts at 1/10 and at full source. The lane report +holds the table. + +## 7. Pins + +These modules are hashed into stage implementation identities, so the pins were +re-derived through their generators rather than assumed. + +| pin | generator | result | +|---|---|---| +| `graph_implementation_inventory.json` → `contracts` (all 124) | `graph_implementation._dependency_contract(payload, name, _covered_imports(name, inventory))` | unchanged by these edits — a constant's value is not an import, an unbound use or a resource access | +| `acs_native_coverage_binding._ACCEPTED["acs_pums.py"]` | `sha256(.read_bytes())`, the same call the check makes | **moved**, `6ecf79f0…` → `e79a2a4e…` | +| the other three `_ACCEPTED` entries | same | unchanged; this branch does not edit those modules | +| all ten `implementation_manifest(stage)` | `graph_implementation.implementation_manifest` | built | + +`experiments/native-row-ceilings/repin.py` reports all of this in one run and +`regenerate_accepted_pin.py` produced the moved value. Neither pin was +hand-edited, and both generators are idempotent. + +What moves and is not a committed pin: each US stage's `implementation_hash` is +over its whole module roster, so editing any inventoried module moves it and +with it every node key and store address. #935 recorded that this is true of any +change to these files including a comment. From 3aa1802a88d80672944cb6906c07ce6486163dce Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:01:40 -0400 Subject: [PATCH 11/22] Apply the rule to a sixth ceiling the transport lane's census did not reach current_survey_geography.MAX_HOUSEHOLDS was 64 * 1024**2 // 128 = 524,288 against the 1,587,376 households a full-source build selects, so it binds -- harder than any of the five, at 33% of source. It is a pure row ceiling despite its byte-derived form: _projection_digest streams one bounded row encoding at a time into a hashlib.sha256 and materialises no per-household payload, and the module's only bytes are the 64 KiB summary receipt that MAX_RECEIPT_BYTES already bounds. So no byte transport sits behind it, the rule applies with no judgment left, and the expression's borrowed 64 MiB never described anything in this module. MAX_HOUSEHOLDS 524,288 -> 7,000,000 (4x 1,587,376, rounded up) Refusals unchanged: HOUSEHOLD_COUNT at :68 and PROJECTION_STORAGE at :188, both ValueError("CURRENT_SURVEY_GEOGRAPHY_" + reason). Boundary test drives PROJECTION_STORAGE over a three-row projection at a patched-down ceiling. Nothing in the tree referenced the old value. No pin moves: the module is not in graph_implementation_inventory.json and not in _ACCEPTED. Co-Authored-By: Claude Opus 5 --- .../us_runtime/current_survey_geography.py | 9 +++++++- .../tests/test_us_current_survey_geography.py | 23 +++++++++++++++++++ .../tests/test_us_native_row_ceilings.py | 2 ++ 3 files changed, 33 insertions(+), 1 deletion(-) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/current_survey_geography.py b/packages/microcosm-build/src/microcosm/build/us_runtime/current_survey_geography.py index f463b6412..535927a85 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/current_survey_geography.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/current_survey_geography.py @@ -26,7 +26,14 @@ "survey_observed_state", "survey_observed_puma", ) -MAX_HOUSEHOLDS = 64 * 1024**2 // 128 +# The selected/stacked household roster, which a full-source build supplies at +# 1,587,376. Four times that, rounded up to the next whole million. The old +# 64 * 1024**2 // 128 borrowed a byte budget this module does not have: its +# only bytes are the 64 KiB summary receipt below, because _projection_digest +# streams one bounded row encoding at a time into a hash and materialises no +# per-household payload. So there is no byte transport behind this bound and +# it is an explicit row ceiling. See docs/us-native-row-ceilings.md. +MAX_HOUSEHOLDS = 7_000_000 MAX_RECEIPT_BYTES = 64 * 1024 diff --git a/packages/microcosm-build/tests/test_us_current_survey_geography.py b/packages/microcosm-build/tests/test_us_current_survey_geography.py index e8ea63dab..a6f49d436 100644 --- a/packages/microcosm-build/tests/test_us_current_survey_geography.py +++ b/packages/microcosm-build/tests/test_us_current_survey_geography.py @@ -190,3 +190,26 @@ def trace(frame, event, result): finally: sys.setprofile(previous) assert fired == [True] + + +def test_household_ceiling_refuses_at_its_own_number(monkeypatch): + """The refusal is `<= MAX_HOUSEHOLDS`, whatever that number is. + + `_projection_digest` streams one bounded row encoding at a time into a hash, + so the bound guards a row count and not a materialised payload. Driven here + at a patched-down ceiling over a three-row projection; + `test_us_native_row_ceilings.py` carries the separate assertion that the + shipped number admits a full-source selection's 1,587,376 households. + """ + table = pd.DataFrame( + { + column: pd.array(["x", "y", "z"], dtype=dtype_for_token("string")) + for column in geography.COLUMNS + }, + index=pd.Index([1, 2, 3], name="household_id", dtype="int64"), + ) + monkeypatch.setattr(geography, "MAX_HOUSEHOLDS", 3) + assert len(geography._projection_digest(table)) == 64 + monkeypatch.setattr(geography, "MAX_HOUSEHOLDS", 2) + with pytest.raises(ValueError, match="PROJECTION_STORAGE"): + geography._projection_digest(table) diff --git a/packages/microcosm-build/tests/test_us_native_row_ceilings.py b/packages/microcosm-build/tests/test_us_native_row_ceilings.py index f3bc56d38..4e8f7452b 100644 --- a/packages/microcosm-build/tests/test_us_native_row_ceilings.py +++ b/packages/microcosm-build/tests/test_us_native_row_ceilings.py @@ -28,6 +28,7 @@ acs_person_coverage_columns, acs_pums, asec_current_money, + current_survey_geography, survey_observed_age, survey_origin_budget, ) @@ -51,6 +52,7 @@ (acs_person_coverage_columns, "MAX_SELECTED_ROWS", ACS_PERSONS, 14_000_000), (survey_observed_age, "MAX_ROWS", ACS_PERSONS, 14_000_000), (survey_origin_budget, "MAX_GROUPS", STACKED_HOUSEHOLDS, 7_000_000), + (current_survey_geography, "MAX_HOUSEHOLDS", STACKED_HOUSEHOLDS, 7_000_000), ) From 3fb1111e282602fff3cbb334a5d49b2f9a912cc4 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:02:13 -0400 Subject: [PATCH 12/22] Add the changelog fragment Co-Authored-By: Claude Opus 5 --- changelog.d/native-row-ceilings.changed.md | 1 + 1 file changed, 1 insertion(+) create mode 100644 changelog.d/native-row-ceilings.changed.md diff --git a/changelog.d/native-row-ceilings.changed.md b/changelog.d/native-row-ceilings.changed.md new file mode 100644 index 000000000..60a5e5c71 --- /dev/null +++ b/changelog.d/native-row-ceilings.changed.md @@ -0,0 +1 @@ +Raise the six US native-build row-count ceilings a full-source build meets to four times their measured full-source counts, keeping every refusal code and expression, and re-pin the one implementation digest that moves. From d0ead6d03683eddbe8f0f399e81eec5340f03439 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:05:54 -0400 Subject: [PATCH 13/22] Record why the origin budget's byte cap cannot be raised here at all graph._bounded_json opens with _require(type(limit) is int and 0 < limit <= 64 * 1024**2, "TRANSPORT_LIMIT") so the shared encoder refuses any cap above 64 MiB before it encodes a byte. A larger MAX_PAYLOAD_BYTES in survey_origin_budget would not loosen the bound; it would refuse the module. That turns "this takes the transport argument" from a preference into a structural fact, and the test now drives it. Also names, in section 5a, the test the rule applies to a bound that looks like a row ceiling -- does the binding site stream, or materialise a per-row payload -- and the three bounds it decides: current_survey_geography streams (moved), current_survey_household_roles serialises its whole per-person table to JSON under a 64 MiB artifact cap, and current_child_property_income_source does one json.dumps under a 64 MiB projection cap. Both of those are the transport's. Co-Authored-By: Claude Opus 5 --- docs/us-native-row-ceilings.md | 55 +++++++++++++++++-- .../tests/test_us_native_row_ceilings.py | 8 +++ 2 files changed, 58 insertions(+), 5 deletions(-) diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md index b522d2d85..c32ef2e1b 100644 --- a/docs/us-native-row-ceilings.md +++ b/docs/us-native-row-ceilings.md @@ -185,11 +185,44 @@ meets; it does not, and this note corrects that.** ## 5. What still binds, and whose argument it is -The census in §6 found fourteen bounds a full-source build meets. Five are the -row counts §3 moved. The rest are byte transports and byte-derived row -pre-checks, and they belong to the transport lane's argument — a segmented -stream under one explicit total — not to this one. The loudest is in a module -this lane did touch, so it gets said plainly: +The census in §6 found fourteen bounds a full-source build meets. Six are the +row counts §3 moved. The rest belong to the transport lane's argument — a +segmented stream under one explicit total — not to this one. + +### 5a. The test the rule applies + +Three of the remaining bounds *look* like row ceilings, and one of them was +moved here after that test was applied to it. The test is a single question, +answered from code rather than from the constant's name: + +> **Does the binding site materialise a per-row payload, or does it stream?** + +A bound in front of a materialised payload cannot be usefully raised on its own, +because the payload's own byte cap refuses first and at a smaller number. A +bound in front of a streaming digest has nothing behind it, and the rule applies. + +| bound | value | what is behind the binding site | verdict | +|---|---:|---|---| +| `current_survey_geography.MAX_HOUSEHOLDS` | 524,288 | `_projection_digest` streams one `_encode(values, maximum=4096)` at a time into a `hashlib.sha256`; the module's only bytes are a 64 KiB summary receipt | **streams — moved in §3** | +| `current_survey_household_roles.MAX_PERSONS` | 2,097,152 | `_projection_bytes` is `table.reset_index().to_json(orient="table").encode()` — the whole per-person table in one string — under `graph_current_survey_household_roles.MAX_ARTIFACT_BYTES` = 64 MiB | materialises — transport's | +| `current_child_property_income_source.MAX_ROWS` | 600,000 | `_json` is one `json.dumps(value)` under `MAX_PROJECTION_BYTES` = 64 MiB | materialises — transport's | + +The middle row is the clearest case for why the test matters: at roughly 200 +bytes of JSON per person, that 64 MiB cap admits a few hundred thousand persons, +so raising the 2,097,152 row bound would move nothing at all. + +One further bound binds and is deliberately **not** decided here. +`asec_demographic_source._MAX_PERSONS` (600,000) is enforced at five sites that +count two different quantities — a fixed ASEC three-cohort roster of 432,523 +rows, and the retained ACS person roster, which a full-source build grows past +three million — and its `_encode` builds a `struct`-packed binary payload rather +than a streamed digest. One constant serving two quantities, in front of a +binary encoding, is not a case the rule settles by itself. The lane report puts +it to the owner rather than guessing. + +### 5b. The loudest one + +It is in a module this lane did touch, so it gets said plainly: > **`survey_origin_budget.MAX_PAYLOAD_BYTES` (64 MiB) admits 87,838 households — > 5.53% of source.** That is below one tenth, and below the 96,860-household @@ -202,6 +235,18 @@ builds one faithful origin record through the module's own `_reference` and including the two `household_ids` and two `group_indices` entries the header carries for each group's two clone roles. +**It also cannot be raised on its own, and that is a structural fact rather +than a preference.** `survey_origin_budget._json` is +`graph._bounded_json(value, MAX_PAYLOAD_BYTES)`, and `_bounded_json` opens with + +```python +_require(type(limit) is int and 0 < limit <= 64 * 1024**2, "TRANSPORT_LIMIT") +``` + +so the shared encoder refuses any cap above 64 MiB before it encodes a byte. A +larger number in this module would not loosen the bound; it would refuse the +module. The 64 MiB is the encoder's, and moving it is the transport change. + The consequence for this lane is stated rather than glossed: **lifting `MAX_GROUPS` is necessary and not sufficient.** `GROUP_COUNT_BOUND` is checked in `_initial` before the payload is built, so at full source it fired first; with diff --git a/packages/microcosm-build/tests/test_us_native_row_ceilings.py b/packages/microcosm-build/tests/test_us_native_row_ceilings.py index 4e8f7452b..add31a681 100644 --- a/packages/microcosm-build/tests/test_us_native_row_ceilings.py +++ b/packages/microcosm-build/tests/test_us_native_row_ceilings.py @@ -29,6 +29,7 @@ acs_pums, asec_current_money, current_survey_geography, + graph_survey_population, survey_observed_age, survey_origin_budget, ) @@ -145,6 +146,13 @@ def test_the_origin_budget_byte_transport_is_left_for_the_other_argument(): admitted = survey_origin_budget.MAX_PAYLOAD_BYTES // 764 assert admitted < STACKED_HOUSEHOLDS // 10 assert survey_origin_budget.MAX_GROUPS > STACKED_HOUSEHOLDS + # And it could not be raised here even if this lane wanted to: the shared + # encoder refuses any cap above 64 MiB before encoding a byte, so a larger + # number in this module would refuse the module rather than loosen it. + with pytest.raises( + graph_survey_population.SurveyPopulationGraphError, match="TRANSPORT_LIMIT" + ): + graph_survey_population._bounded_json({"a": 1}, 64 * 1024**2 + 1) def test_the_measured_counts_reconcile(): From f2428e929543239ec4d38e62a7520c16489a4d6f Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:09:10 -0400 Subject: [PATCH 14/22] Apply the rule to a seventh ceiling, after reading what stood behind it asec_demographic_source._MAX_PERSONS was 600,000 over two rosters: a fixed ASEC three-cohort 432,523 that no fraction grows, and the retained ACS persons, which a full-source build grows to 3,422,888. One constant over two rosters takes the larger, so the rule produces 14,000,000. Three things had to hold and each was read. It is not a fixed-width encoding: rows is a JSON header integer, the only struct.pack is " --- docs/us-native-row-ceilings.md | 33 ++++++++++++++----- .../us_runtime/asec_demographic_source.py | 10 +++++- .../tests/test_us_asec_demographic_source.py | 23 +++++++++++++ .../tests/test_us_native_row_ceilings.py | 2 ++ 4 files changed, 59 insertions(+), 9 deletions(-) diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md index c32ef2e1b..eec1cf329 100644 --- a/docs/us-native-row-ceilings.md +++ b/docs/us-native-row-ceilings.md @@ -109,6 +109,8 @@ persons. A bound met per channel is met by 3,422,888, not by 3,565,013. | `acs_person_coverage_columns.MAX_SELECTED_ROWS` | 1,000,000 | **14,000,000** | requested ACS person keys | 3,422,888 | 4.09 | | `survey_observed_age.MAX_ROWS` | 2,000,000 | **14,000,000** | one channel's person rows | 3,422,888 | 4.09 | | `survey_origin_budget.MAX_GROUPS` | 1,000,000 | **7,000,000** | allocation instructions = selected households | 1,587,376 | 4.41 | +| `current_survey_geography.MAX_HOUSEHOLDS` | 524,288 | **7,000,000** | selected households | 1,587,376 | 4.41 | +| `asec_demographic_source._MAX_PERSONS` | 600,000 | **14,000,000** | retained ACS persons (the larger of its two rosters) | 3,422,888 | 4.09 | Each refusal keeps its code, its exception type and its expression; only the number moves. @@ -120,6 +122,8 @@ number moves. | `MAX_SELECTED_ROWS` | `"ACS coverage selected person count is outside the bound"` | `ValueError` | | `survey_observed_age.MAX_ROWS` | `"SURVEY_OBSERVED_AGE_ROW_BOUND"` | `ValueError` | | `MAX_GROUPS` | `"GROUP_COUNT_BOUND"` | `SurveyOriginBudgetError` | +| `current_survey_geography.MAX_HOUSEHOLDS` | `"CURRENT_SURVEY_GEOGRAPHY_HOUSEHOLD_COUNT"`, `…_PROJECTION_STORAGE` | `ValueError` | +| `asec_demographic_source._MAX_PERSONS` | `"ACS_ROWS"`, `"MEMBERSHIP_ROWS"`, `"CLASSIFY_ROWS"`, `"DEMOGRAPHIC_ROWS"` | `ValueError` | ### 3a. `survey_origin_budget.MAX_GROUPS`, which the transport lane left open @@ -185,7 +189,7 @@ meets; it does not, and this note corrects that.** ## 5. What still binds, and whose argument it is -The census in §6 found fourteen bounds a full-source build meets. Six are the +The census in §6 found fourteen bounds a full-source build meets. Seven are the row counts §3 moved. The rest belong to the transport lane's argument — a segmented stream under one explicit total — not to this one. @@ -211,14 +215,27 @@ The middle row is the clearest case for why the test matters: at roughly 200 bytes of JSON per person, that 64 MiB cap admits a few hundred thousand persons, so raising the 2,097,152 row bound would move nothing at all. -One further bound binds and is deliberately **not** decided here. +A fourth bound needed the same test plus one more question, and passed both. `asec_demographic_source._MAX_PERSONS` (600,000) is enforced at five sites that -count two different quantities — a fixed ASEC three-cohort roster of 432,523 -rows, and the retained ACS person roster, which a full-source build grows past -three million — and its `_encode` builds a `struct`-packed binary payload rather -than a streamed digest. One constant serving two quantities, in front of a -binary encoding, is not a case the rule settles by itself. The lane report puts -it to the owner rather than guessing. +count **two different quantities**: a fixed ASEC three-cohort roster of 432,523 +rows, and the retained ACS person roster, which a full-source build grows to +3,422,888. One constant over two rosters takes the larger, so the rule produces +14,000,000 from the ACS arm. Three things had to hold before it could move, and +each was read rather than assumed: + +- **It is not a fixed-width encoding.** `rows` lives as a JSON integer in the + header; the module's only `struct.pack` is `" Date: Thu, 17 Sep 2026 18:12:46 -0400 Subject: [PATCH 15/22] Measure the two ceilings that refuse before anything this lane moved Both are byte transports, so neither is this lane's to move. Both are measured and pinned in a test, because a reader who meets them in a build has lost hours. ACS coverage authentication charges every selected row 6 * len(raw) + 1024 against MAX_BODY_BYTES before the reader allocates, refusing SELECTED_BODY_BUDGET. Measured over 200,000 real records of the pilot's captured public ACS PUMS archive: a person record averages 695.57 bytes, so the charge is 5,197 bytes and 64 MiB admits 12,911 selected persons -- 0.38% of source, 265x under at full source. It is the tightest ceiling censused, and it refuses 265x before acs_person_coverage_columns.MAX_SELECTED_ROWS, which this lane lifted. And the preparation-receipt ceiling the transport lane reported as moved from 96,860 households to 6,206,000 is still enforced one module downstream: _roster_payload returns one joined payload under MAX_ROSTER_BYTES = 4 GiB, and graph_survey_population._checked_preparation checks those same bytes against PREPARATION_MAX_BYTES, still 64 MiB, refusing PREPARATION_BYTES. From the transport lane's own committed receipt: a 1/10 roster measured 109,804,304 bytes and was recorded accepted by the producer, 1.64x this cap; a full-source roster measured 1,099,892,722 bytes, of which this cap admits 96,839 households. That is the number the transport lane lifted, still standing. Co-Authored-By: Claude Opus 5 --- experiments/native-row-ceilings/census.json | 2881 +++++++++++++++++ .../native-row-ceilings/consumer-gap.json | 26 + .../native-row-ceilings/consumer_gap.py | 92 + .../selected-body-budget.json | 37 + .../selected_body_budget.py | 104 + .../tests/test_us_native_row_ceilings.py | 59 + pyproject.toml | 2 + 7 files changed, 3201 insertions(+) create mode 100644 experiments/native-row-ceilings/census.json create mode 100644 experiments/native-row-ceilings/consumer-gap.json create mode 100644 experiments/native-row-ceilings/consumer_gap.py create mode 100644 experiments/native-row-ceilings/selected-body-budget.json create mode 100644 experiments/native-row-ceilings/selected_body_budget.py diff --git a/experiments/native-row-ceilings/census.json b/experiments/native-row-ceilings/census.json new file mode 100644 index 000000000..e4dd1af2f --- /dev/null +++ b/experiments/native-row-ceilings/census.json @@ -0,0 +1,2881 @@ +{ + "scope": "Census of every MAX_* row/household/person/group/byte bound in packages/microcosm-build/src/microcosm/build/us_runtime/ reachable from the 19-node financial graph and the 45-node pilot graph, read at this head, with an adversarial verification pass over every verdict that claimed a full-source build meets the bound or that the bound guards an encoding width or an upstream file's size. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "bounds_censused": 146, + "verdicts_verified": 41, + "binds_at_full_source": [ + "acs_housing_universe_source.py:43 ACS_HU_RECEIPT_MAX_BYTES", + "acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES", + "acs_person_coverage_authentication.py:37 MAX_BODY_BYTES", + "asec_demographic_source.py:67 _MAX_PERSONS", + "current_child_property_income_source.py:34 MAX_ROWS", + "current_survey_geography.py:29 MAX_HOUSEHOLDS", + "current_survey_household_roles.py:86 MAX_PERSONS", + "graph_survey_age_artifact.py:33 MAX_BYTES", + "graph_survey_calibration.py:44 MAX_BYTES", + "graph_survey_calibration.py:45 MAX_ROWS", + "graph_survey_population.py:61 PREPARATION_MAX_BYTES", + "puf55_survey_recipients.py:22 RAW_BYTES_MAX_BYTES (module-level name bound by `from microcosm.graph.codecs import RAW_BYTES_MAX_BYTES`; defined at packages/microcosm-graph/src/microcosm/graph/codecs.py:78)", + "puf_diagnostic_consumer.py:671 CURRENT_SURVEY_MAX_BYTES", + "survey_origin_budget.py:54 MAX_PAYLOAD_BYTES" + ], + "binds_at_one_tenth": [ + "acs_housing_universe_source.py:43 ACS_HU_RECEIPT_MAX_BYTES", + "acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES", + "acs_person_coverage_authentication.py:37 MAX_BODY_BYTES", + "graph_current_survey_household_roles.py:70 MAX_ARTIFACT_BYTES", + "graph_survey_population.py:61 PREPARATION_MAX_BYTES", + "survey_origin_budget.py:54 MAX_PAYLOAD_BYTES" + ], + "by_family": { + "PUF and downstream transfer": 24, + "Survey population, origins and age": 27, + "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)": 18, + "ASEC codec and prepared source": 46, + "ACS source and coverage": 31 + }, + "bounds": [ + { + "constant": "puf_growth.py:197 MAX_ROWS", + "value": "2_000_000 (literal)", + "enforced_at": "puf_growth.py:1207 (_checked_source_table); puf_growth.py:1441 (apply_puf_design_weight_growth)", + "refusal_code": "\"SOURCE_TABLE_ROWS\" at puf_growth.py:1207 and \"WEIGHT_SHAPE\" at puf_growth.py:1441; both raised by _require (puf_growth.py:219-221) as PufGrowthRefusalError, defined puf_growth.py:207 as a subclass of ValueError", + "protects": "memory-or-time", + "protects_detail": "The module comment at puf_growth.py:195-196 states it outright: \"Largest table this transform will accept, so a mistake is a refusal rather than an unbounded allocation.\" Nothing in the receipt, the contract digest or any wire format depends on the value; it is only compared against len(table) and values.size.", + "counted_quantity": "Rows of the PUF source money-view pandas DataFrame handed to apply_puf_growth (puf_growth.py:1205-1207), and the length of the design-weight array (puf_growth.py:1441). Both are at IRS PUF 2015 source-return grain \u2014 NOT the selected, stacked or cloned survey.", + "count_at_tenth": "207,696 (invariant: the PUF delivery's data_records, pinned at puf_2015_raw_source_definition.json sources.main.data_records = 207696 and restated in puf_raw_source.py:5-6)", + "count_at_full_source": "207,696 (same; the PUF file does not scale with the survey selection fraction)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at any survey scale: the counted roster is the fixed 207,696-row PUF delivery, 9.6x below the bound. It is also unreachable in the release lane for a second reason \u2014 REVIEWED_FACTOR_TABLE_DIGESTS is empty (puf_growth.py:155) and RELEASE_ELIGIBLE_AUTHORITIES admits only REVIEWED_RESOURCE (puf_growth.py:326), so no release-eligible contract can compile today.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_growth.py:200 MAX_DOCUMENT_BYTES", + "value": "512 * 1024 = 524,288", + "enforced_at": "puf_growth.py:234 (_parse); puf_growth.py:710 (FactorTable.__post_init__)", + "refusal_code": "\"DOCUMENT_SIZE_OR_TYPE\" at both sites; PufGrowthRefusalError (puf_growth.py:207, subclass of ValueError)", + "protects": "memory-or-time", + "protects_detail": "Comment at puf_growth.py:199: \"Largest factor document this transform will parse.\" It bounds one JSON growth-factor table and the receipt re-parse; it is a parser allocation guard, not part of any encoded layout.", + "counted_quantity": "Bytes of a single growth factor-table JSON document (and of the emitted receipt when re-parsed through the same _parse). One document per contract; it does not grow with any row roster.", + "count_at_tenth": "not established \u2014 no factor table document exists in-tree to measure (REVIEWED_FACTOR_TABLE_DIGESTS is the empty tuple at puf_growth.py:155); the quantity is row-independent either way", + "count_at_full_source": "not established, same reason; row-independent", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Row-independent by construction: the document holds declared factor series and provenance keys, never per-return data (_parse at puf_growth.py:231-252 parses a plain JSON object). Scaling the survey cannot move it.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_detail_transfer.py:33 MAX_ROWS", + "value": "4096 (literal); used only as the default of the decode_price_arrays keyword max_rows at puf_detail_transfer.py:106", + "enforced_at": "puf_detail_transfer.py:113 (domain check on the supplied max_rows: 0 < max_rows <= 250_000); puf_detail_transfer.py:158-165 (per-column: 0 < entry[\"rows\"] <= max_rows)", + "refusal_code": "\"PRICE_ROW_BOUND\" at puf_detail_transfer.py:113 and \"PRICE_COLUMN_TYPE_ROWS\" at puf_detail_transfer.py:165; raised by require (puf_detail_transfer.py:78-81) as a bare ValueError", + "protects": "memory-or-time", + "protects_detail": "It caps how many rows a locally decoded price-restatement envelope may declare per column before np.frombuffer allocates; the body offsets are computed from rows * itemsize (puf_detail_transfer.py:167-172), so a large declared row count is an allocation, not a format, question. The envelope is explicitly fixture-only (SCOPE = \"invented_fixture_machinery_only\", puf_detail_transfer.py:30; release_eligible must be False at :133).", + "counted_quantity": "Rows per array column in the local price-restatement envelope = the ordinary (non-disclosure-aggregate) PUF 2015 returns carried as donor rows. PUF grain, not survey grain.", + "count_at_tenth": "207,692 ordinary returns (207,696 delivered minus the four PUF_AGGREGATE_RECIDS at puf_raw_source.py:141); invariant to the survey fraction", + "count_at_full_source": "207,692 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "The 4096 default never applies on the qualified route: puf_diagnostic_consumer.py:355 passes max_rows=rows, where rows is config[\"ordinary_rows\"] already checked against MAX_DONORS. Only a direct caller using the default would hit it, and at the real 207,692-row PUF that call would refuse with PRICE_COLUMN_TYPE_ROWS. The operative production ceiling on the same quantity is the inline 250_000 at puf_detail_transfer.py:113 and MAX_DONORS.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_detail_transfer.py:34 MAX_BYTES", + "value": "128 * 1024 * 1024 = 134,217,728", + "enforced_at": "puf_detail_transfer.py:114-119 (len(payload) <= MAX_BYTES inside the envelope check)", + "refusal_code": "\"PRICE_ENVELOPE\" at puf_detail_transfer.py:119; require -> ValueError (puf_detail_transfer.py:78-81)", + "protects": "memory-or-time", + "protects_detail": "A whole-payload byte ceiling checked before any parsing or hashing of the envelope. It is not consulted by any reader offset computation (offsets come from the header's per-column rows and dtypes at :167-172), so it is a resource ceiling, not a layout fact.", + "counted_quantity": "Total bytes of the price-restatement envelope: MAGIC + 8-byte header length + header (<= 65,536 by the inline PRICE_HEADER_BOUND at :128) + 155 bytes of column body per row (22 columns in ARRAY_DTYPES at puf_detail_transfer.py:50-61: 13 money f8 + RECID i8 + puf_source_recid i8 + puf_source_year_agi f8 + S006 f8 + FLPDYR i4 + demographic_status i4 + S006_delivered i8 + FLPDYR_delivered i2 + demographic_status_delivered i1).", + "count_at_tenth": "~32.2 MB (207,692 ordinary PUF rows x 155 bytes = 32,192,260 bytes + header); invariant to the survey fraction", + "count_at_full_source": "~32.2 MB (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: the implied row ceiling is 134,217,728 / 155 = 866,000 rows against a fixed 207,692-row PUF donor roster. MAX_DONORS (250,000) and the inline 250_000 PRICE_ROW_BOUND both bite long before this byte cap.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_diagnostic_consumer.py:32 MAX_DONORS", + "value": "250_000 (literal)", + "enforced_at": "puf_diagnostic_consumer.py:331-340 (0 < rows <= MAX_DONORS, with rows == config[\"ordinary_rows\"] == source/acceptance/completed ordinary_rows and excluded_aggregate_rows == 4)", + "refusal_code": "\"PRICE_ORDINARY_ROSTER\" at puf_diagnostic_consumer.py:340; require is detail.require (aliased at puf_diagnostic_consumer.py:27), which raises a bare ValueError", + "protects": "memory-or-time", + "protects_detail": "It is the ceiling the qualified transport then hands to the array decoder as max_rows (puf_diagnostic_consumer.py:355), i.e. the allocation budget for np.frombuffer over 22 columns. The structural facts in the same require are the exact equalities across the four receipts and the exactly-4 aggregate rows; the 250,000 is the resource half of that check.", + "counted_quantity": "config[\"ordinary_rows\"]: the PUF 2015 returns inside the individual-return amount universe, i.e. delivered main records minus the four disclosure-aggregate records. PUF grain.", + "count_at_tenth": "207,692 (invariant)", + "count_at_full_source": "207,692 (invariant)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind, but the headroom is only ~20% (207,692 of 250,000) and it is fixed by the PUF delivery, not by the survey fraction. A future PUF vintage with more returns, not a larger survey, is what would move it.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_diagnostic_consumer.py:671 CURRENT_SURVEY_MAX_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "puf_diagnostic_consumer.py:682-684 (host-projection payload bound); puf_diagnostic_consumer.py:833 (encoded recipient-matrix bound); graph_puf_diagnostic_consumer.py:403 (encode side, via survey_graph._bounded_json); graph_puf_diagnostic_consumer.py:648-652 (kernel-side re-check of the projection payload)", + "refusal_code": "\"SURVEY_HOST_PROJECTION_BOUND\" at puf_diagnostic_consumer.py:684 and graph_puf_diagnostic_consumer.py:651 (require -> bare ValueError); \"SURVEY_MATRIX_BOUND\" at puf_diagnostic_consumer.py:833 (require -> bare ValueError); \"TRANSPORT_LIMIT\" at graph_survey_population.py:273 and :275, raised as SurveyPopulationGraphError (graph_survey_population.py:112-114), which is a ValueError subclass", + "protects": "memory-or-time", + "protects_detail": "Two payload ceilings on one development diagnostic (SCOPE = \"development_conditional_support_not_tax_inputs\", release_eligible False at :846). The projection bound guards a JSON decode; the matrix bound guards the encoded recipient matrix. Neither value is read back by any decoder: model_input.decode_recipient_matrix reads rows from the header (microcosm-fit/src/microcosm/fit/model_input.py:97-107) and bounds only the header with its own MATRIX_HEADER_MAX_BYTES = 64 * 1024. So it is a deliberate resource ceiling, movable on its own terms.", + "counted_quantity": "Two different quantities. (a) puf_diagnostic_consumer.py:683 / graph:403 / graph:650 \u2014 the canonical JSON host-projection document, whose asec[\"rows\"] carries exactly one 7-element row per NATIVE ASEC-channel person in the frame (required equal to len(native_asec) at :756-758). (b) puf_diagnostic_consumer.py:833 \u2014 the encoded recipient matrix from detail.recipient_matrix, which is 48 bytes per CLONE-ONE tax unit (8-byte int64 id + 5 float64 FEATURES, encode_recipient_matrix at microcosm-fit/src/microcosm/fit/model_input.py:88-94), over the mask units[clone_index] == 1 (puf_detail_transfer.py:374-375).", + "count_at_tenth": "(a) ~14,200 native ASEC persons (1/1000 pilot recorded 140; preparation.json /native/asec/persons = 140) -> roughly 0.8 MB of JSON. (b) 213,251 clone-one tax units (158,738 hh x 1.343434) x 48 = 10,236,048 bytes.", + "count_at_full_source": "(a) ~142,000 native ASEC persons \u2014 the whole pinned ASEC person universe (preparation.json /catalogues/asec/counts/persons = 142,125, which equals the pppub25.csv row pin at education_assistance_source.py:163) -> roughly 8 MB of JSON, an estimate at ~55 bytes per 7-element row, not a measured figure. (b) 2,132,510 clone-one tax units x 48 = 102,360,480 bytes.", + "binds_at_tenth": false, + "binds_at_full_source": true, + "binding_note": "Arm (b) BINDS at full source: 102,360,480 > 67,108,864, so SURVEY_MATRIX_BOUND refuses. The threshold is 67,108,864 / 48 = 1,398,101 tax units, i.e. about 1,040,000 selected households (~66% of the 1,587,376-household source), so it first binds somewhere between the 1/10 and full runs. Arm (a) never binds, and importantly it is an ASEC-roster quantity capped by the real upstream file (142,125 person rows), not by the survey fraction \u2014 it plateaus at ~8 MB. No tighter upstream batching limit intervenes: recipient_matrix builds the whole matrix in one pass over the frame.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "puf_monetary_source.py:180 BODY_MAX_BYTES", + "value": "64 * 1024 * 1024 = 67,108,864", + "enforced_at": "puf_monetary_source.py:1388-1389 (encode_monetary_projection, on payload_bound total); puf_monetary_source.py:1567-1568 (decode-side running column total)", + "refusal_code": "\"BODY_OVER_BOUND\" at puf_monetary_source.py:1389 and \"ENVELOPE_BODY_OVER_BOUND\" at puf_monetary_source.py:1568; _refuse (puf_monetary_source.py:349) returns PufMonetaryRefusalError, defined at puf_monetary_source.py:341 as a subclass of ValueError", + "protects": "memory-or-time", + "protects_detail": "The docstring at puf_monetary_source.py:177-179 gives the sizing argument explicitly: \"Twelve projected columns over the delivered 207,696 rows is 262 bytes a row, so about 51.9 MiB; a thirteenth column of the same shape would still fit and a much wider selection would not.\" The per-column sizes are recomputed from rows and dtypes (payload_bound, :1385-1387), so the reader does not rely on the cap for offsets.", + "counted_quantity": "Total encoded column-body bytes of the monetary projection artifact: 262 bytes per PUF source return (structural RECID i8 + disclosure_aggregate u1 + demographic_status i1 = 10 bytes, plus 12 projected columns x (8 typed + 13 lexical) = 252). PUF grain.", + "count_at_tenth": "54,416,352 bytes (207,696 x 262) = 51.9 MiB; invariant to the survey fraction", + "count_at_full_source": "54,416,352 bytes (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind, but the implied row ceiling is only 256,140 rows (67,108,864 / 262) against 207,696 delivered rows \u2014 23% headroom, fixed by the PUF file. Adding projected columns, not survey rows, is what would move it.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_monetary_source.py:184 HEADER_MAX_BYTES", + "value": "64 * 1024 = 65,536", + "enforced_at": "puf_monetary_source.py:1455-1456 (encode, canonical_json header); puf_monetary_source.py:1643 (decode, 0 < length <= HEADER_MAX_BYTES)", + "refusal_code": "\"HEADER_OVER_BOUND\" at puf_monetary_source.py:1456 and \"ENVELOPE_HEADER_LENGTH\" at puf_monetary_source.py:1644; PufMonetaryRefusalError (ValueError subclass)", + "protects": "memory-or-time", + "protects_detail": "Comment at :182-183: the header carries per-column facts for three row classes, so it is bigger than the status artifact's, but it is still one fixed-shape JSON object (schema keys enumerated in _ENVELOPE_HEADER_KEYS, :267-282). The 4-byte big-endian length prefix (:1459) could express far more, so the cap is policy, not format.", + "counted_quantity": "Bytes of the canonical JSON envelope header: per-column digests plus aggregate/return-class fact blocks. The number of columns (12) and row classes (3) is fixed; per-row data never enters the header.", + "count_at_tenth": "not established (I did not build an artifact to measure the header); structurally row-independent \u2014 the header holds 12 column entries and three fixed fact blocks", + "count_at_full_source": "not established, same; row-independent", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Row-independent: nothing in the header grows with the row roster, so no survey scale can move it. The only growth vector is more projected columns or more row classes.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_monetary_source.py:187 LEXICAL_WIDTH_MAX", + "value": "64 (literal)", + "enforced_at": "puf_monetary_source.py:511-512 (declared projection column width); puf_monetary_source.py:1550-1551 (envelope column width, plus exact equality with the declared width)", + "refusal_code": "\"PROJECTION_LEXICAL_WIDTH\" at puf_monetary_source.py:512 and \"ENVELOPE_COLUMN_WIDTH\" at puf_monetary_source.py:1551; PufMonetaryRefusalError (ValueError subclass)", + "protects": "encoding-width", + "protects_detail": "Comment at :186: \"The widest lexical allocation a projected column may declare.\" The lexical bodies are fixed-width NUL-padded ASCII (_LEXICAL_KIND at :204) and every reader slices body[position*width:(position+1)*width] (_read_fixed_ascii, :1440-1451) and requires bytes == rows * width (:1553, :1565). Changing a width changes the stored artifact layout, so this is a wire-format bound, not a resource one.", + "counted_quantity": "Bytes allocated per lexical cell for one projected monetary column. The packaged projection declares 13 for all twelve columns (puf_2015_monetary_source_projection.json, projection.columns[*].lexical_width).", + "count_at_tenth": "13 (declared width for every one of the 12 columns); invariant", + "count_at_full_source": "13 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and is not a scale bound: it caps per-cell width (13 used of 64), not any roster. Listed because the name matches the MAX criterion and it is a byte ceiling.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "puf_monetary_source.py:338 _MAX_AMOUNT_DIGITS", + "value": "12 (literal)", + "enforced_at": "puf_monetary_source.py:515-516 (digits > _MAX_AMOUNT_DIGITS or width < digits + 1)", + "refusal_code": "\"PROJECTION_WIDTH_BELOW_GRAMMAR\" at puf_monetary_source.py:516; PufMonetaryRefusalError (ValueError subclass)", + "protects": "encoding-width", + "protects_detail": "Comment at :337: \"The publisher's amount fields are twelve digits wide.\" It validates the declared field_width_digits of each projected column against the publisher's delivered grammar and forces the lexical allocation to leave room for the sign (width >= digits + 1). Per-token grammar, checked again per value at :879-880 against column.max_digits.", + "counted_quantity": "Declared decimal digits per delivered amount token for one projected column (12 in the packaged projection).", + "count_at_tenth": "12; invariant", + "count_at_full_source": "12; invariant", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a roster bound at all \u2014 a token grammar width. Included only because the name contains MAX and it is a character ceiling.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": true + } + }, + { + "constant": "puf_raw_source.py:117 PUF_RAW_SOURCE_MAX_BYTES", + "value": "128 * 1024 * 1024 = 134,217,728", + "enforced_at": "puf_raw_source.py:318-319 (declared source bytes in a definition); puf_raw_source.py:502-503 (codec pin before opening the file); puf_raw_source.py:1219-1220 (envelope-declared source bytes)", + "refusal_code": "\"DEFINITION_SOURCE_OVER_CAP\" (:319), \"CODEC_PIN_OVER_CAP\" (:503), \"ENVELOPE_SOURCE_OVER_CAP\" (:1220); _refuse (puf_raw_source.py:240) returns PufRawSourceRefusalError, defined puf_raw_source.py:233 as a subclass of ValueError", + "protects": "memory-or-time", + "protects_detail": "The module docstring at puf_raw_source.py:29-31 states the rationale: \"The private cap is private. PUF_RAW_SOURCE_MAX_BYTES is 128 MiB because the pinned main delivery is 126,034,649 bytes. microcosm.graph.codecs.RAW_BYTES_MAX_BYTES stays at 64 MiB.\" It is a read/allocation ceiling deliberately sized just above the real delivery. The STRUCTURAL assertion about the real file is separate and exact: load_pinned_puf_bytes requires before.st_size == pin.bytes (:516-517) plus SHA-256 and git-blob-SHA-1 equality, and decode_full_puf_source requires len(data) == pin.bytes (puf_full_source.py:177). Those, not this cap, are what protect the upstream file's true size.", + "counted_quantity": "Declared/actual byte size of one delivered PUF file. Two files: main = 126,034,649 bytes, demographic = 2,225,050 bytes (puf_2015_raw_source_definition.json sources.*.bytes).", + "count_at_tenth": "126,034,649 bytes (main); invariant to the survey fraction", + "count_at_full_source": "126,034,649 bytes (main); same", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind, with only 6.5% headroom over the real main delivery (126,034,649 of 134,217,728). It is the one bound in this family whose value was chosen FROM the upstream file's true size; lowering it would refuse the real file, and raising it weakens nothing because exact identity is enforced by pin.bytes + sha256 elsewhere.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_raw_source.py:121 BODY_MAX_BYTES", + "value": "64 * 1024 * 1024 = 67,108,864", + "enforced_at": "puf_raw_source.py:1088-1089 (encode_return_status, on _payload_bound total); puf_raw_source.py:1283-1284 (decode-side running column total)", + "refusal_code": "\"BODY_OVER_BOUND\" (:1089) and \"ENVELOPE_BODY_OVER_BOUND\" (:1284); PufRawSourceRefusalError (ValueError subclass)", + "protects": "memory-or-time", + "protects_detail": "Comment at :119-120: \"Bound on the encoded artifact's column bodies, below the 64 MiB a shared opaque payload is expected to stay under. The full slice is ~35 MB.\" _payload_bound (:1062-1073) recomputes every column size from rows and the declared widths, so the cap governs allocation only.", + "counted_quantity": "Encoded column-body bytes of the return-status artifact: 169 bytes per PUF source return (29 bytes of typed columns from _TYPED_DTYPES at :158-173, plus 140 bytes of fixed-width lexical columns \u2014 RECID 64, S006 32, FLPDYR 8, and nine 4-byte fields, from puf_2015_raw_source_definition.json fields[*].lexical_width).", + "count_at_tenth": "35,100,624 bytes (207,696 x 169) = 33.5 MiB; invariant to the survey fraction", + "count_at_full_source": "35,100,624 bytes (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; implied row ceiling 397,093 (67,108,864 / 169) against 207,696 delivered rows. Matches the docstring's \"~35 MB\".", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_raw_source.py:124 HEADER_MAX_BYTES", + "value": "64 * 1024 = 65,536", + "enforced_at": "puf_raw_source.py:1154-1155 (encode, canonical_json header); puf_raw_source.py:1395 (decode, 0 < length <= HEADER_MAX_BYTES)", + "refusal_code": "\"HEADER_OVER_BOUND\" (:1155) and \"ENVELOPE_HEADER_LENGTH\" (:1396); PufRawSourceRefusalError (ValueError subclass)", + "protects": "memory-or-time", + "protects_detail": "Comment at :123: \"Bound on the canonical JSON header, as in the ACS rent placement envelope.\" The header (built at :1135-1152) holds 26 fixed column entries, two source pins and the fixed facts dict (_ENVELOPE_FACT_KEYS, :211-228). It carries no per-row data. The 4-byte big-endian length prefix could express far more, so the cap is policy.", + "counted_quantity": "Bytes of the canonical JSON envelope header of the return-status artifact.", + "count_at_tenth": "not established (I did not encode an artifact to measure it); structurally row-independent", + "count_at_full_source": "not established, same; row-independent", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Row-independent: only column count, pin strings and the fixed facts block feed it, so no survey scale can move it.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_raw_source.py:185 _LEXICAL_WIDTH_MAX", + "value": "64 (literal)", + "enforced_at": "puf_raw_source.py:1272-1273 (envelope column width check)", + "refusal_code": "\"ENVELOPE_COLUMN_WIDTH\" at puf_raw_source.py:1273 (and again at :1276 for the exact-width mismatch); PufRawSourceRefusalError (ValueError subclass)", + "protects": "encoding-width", + "protects_detail": "Comment at :183-184: \"The widest lexical allocation any field may declare, matching the closed definition's own 0 < lexical_width <= 64 bound.\" Lexical bodies are fixed-width NUL-padded ASCII read back with body[position*width:(position+1)*width] (_read_fixed_ascii, :1046-1060) and sizes must equal rows * width (:1277, :1281). A different width is a different stored artifact.", + "counted_quantity": "Bytes allocated per lexical cell. Declared widths in use: RECID 64, S006 32, FLPDYR 8, and 4 for FLPDMO, MARS, DSI, AGEDP1-3, AGERANGE, EARNSPLIT, GENDER.", + "count_at_tenth": "64 (RECID, the widest in use \u2014 exactly at the bound); invariant", + "count_at_full_source": "64 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a roster bound. Worth noting that the widest declared field (RECID, 64) sits exactly at the ceiling, so a wider lexical field would refuse \u2014 but nothing about the survey scale touches it.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "puf_full_source.py:27 FULL_SOURCE_MAX_BYTES", + "value": "128 * 1024 * 1024 = 134,217,728", + "enforced_at": "puf_full_source.py:194-196 (pre-decode projection body check); puf_full_source.py:290 (encode, sum of all blocks); puf_full_source.py:297-299 (decode, whole-payload check); puf_full_source.py:346-349 (decode, declared status_bytes)", + "refusal_code": "\"FULL_PROJECTION_BODY_LIMIT\" (:196), \"FULL_BODY_LIMIT\" (:290), \"FULL_ARTIFACT_LIMIT\" (:299), \"FULL_STATUS_LENGTH\" (:349); _require (puf_full_source.py:97-99) raises a bare ValueError", + "protects": "memory-or-time", + "protects_detail": "Pure allocation ceilings: :195 pre-checks rows * 59 columns * 8 before allocating the 59 int64 value arrays at :199-202; :290 caps the concatenated blocks at encode; :298 caps the whole payload before parsing. Column offsets on decode come from the per-column declared bytes (== rows * 8, :368-373), not from this cap.", + "counted_quantity": "Two different sums. At :195 \u2014 rows * len(PROJECTED_COLUMNS=59) * 8 = 472 bytes per PUF source return. At :290 \u2014 the SAME 472 bytes/row PLUS the embedded return-status payload from puf_raw_source (169 bytes/row), i.e. 641 bytes per source return.", + "count_at_tenth": "at :195, 98,032,512 bytes (207,696 x 472); at :290, 133,133,136 bytes (207,696 x 641). Invariant to the survey fraction.", + "count_at_full_source": "same: 98,032,512 and 133,133,136 bytes", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind, but this is the tightest bound in the family by a wide margin: the encode-side FULL_BODY_LIMIT sits at 133,133,136 of 134,217,728 bytes \u2014 1,084,592 bytes of headroom, 0.8%. Its implied row ceiling is 209,388 rows against the 207,696 actually delivered. It cannot be moved by survey scale, but a PUF vintage with ~1,700 more returns, or one more projected column, refuses. The looser :195 arm implies 284,359 rows and is not the operative one.", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf_full_source.py:94 AGGREGATE_LEXEME_MAX_CHARACTERS", + "value": "64 (literal)", + "enforced_at": "Not referenced by name in any comparison. The value is duplicated into the regex _AGGREGATE = re.compile(r\"[\\x20-\\x7e]{0,64}\\Z\", re.ASCII) at puf_full_source.py:95, which is enforced at puf_full_source.py:211-214 (decode of aggregate rows) and puf_full_source.py:386-391 (artifact decode of aggregate_tokens)", + "refusal_code": "\"FULL_AGGREGATE_LEXICAL_GRAMMAR\" at puf_full_source.py:214 and puf_full_source.py:391; _require -> bare ValueError (puf_full_source.py:97-99)", + "protects": "encoding-width", + "protects_detail": "Comment at :92-93: \"Operational preservation bound for four out-of-universe source rows. Empty, scientific notation and literal formatting are retained, not parsed.\" It caps the character length of a retained disclosure-aggregate token so those four rows' delivered literals can be preserved verbatim in the JSON header without parsing them.", + "counted_quantity": "Characters in one retained disclosure-aggregate lexeme, for the 4 aggregate rows x 59 projected columns = 236 tokens.", + "count_at_tenth": "not established \u2014 I did not read the delivered aggregate tokens (they live in the gated PUF delivery, not in-tree). The token count is fixed at 236; each must be <= 64 characters. Invariant to the survey fraction.", + "count_at_full_source": "not established, same", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a roster bound, and the constant is dead as a name: only the regex literal at :95 enforces it, so changing the constant alone would change nothing. Flagging that duplication because it is a maintenance trap.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": false, + "protects": "structural-invariant", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "puf_full_source.py:254 _HEADER_LIMIT", + "value": "128 * 1024 = 131,072", + "enforced_at": "puf_full_source.py:289 (encode); puf_full_source.py:297-299 (decode, inside FULL_ARTIFACT_LIMIT); puf_full_source.py:304 (decode, declared header size)", + "refusal_code": "\"FULL_HEADER_LIMIT\" at puf_full_source.py:289 and :304; also a term of \"FULL_ARTIFACT_LIMIT\" at :299; _require -> bare ValueError", + "protects": "memory-or-time", + "protects_detail": "Bounds the canonical JSON header before json.loads. The header (built :272-288) carries per-column digests for 59 columns plus the aggregate_tokens block for the 4 disclosure-aggregate rows (4 x 59 tokens). Its 8-byte little-endian length prefix (:291) could express far more, so the cap is policy.", + "counted_quantity": "Bytes of the full-source artifact's JSON header: 59 column entries + 4 x 59 retained aggregate tokens + source pins. Fixed shape; no per-return data.", + "count_at_tenth": "not established (I did not encode an artifact); row-independent \u2014 only the 4 aggregate rows contribute row-derived content", + "count_at_full_source": "not established, same; row-independent", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Row-independent, so no survey scale can move it. Its only growth vectors are more projected columns, more aggregate rows, or longer retained aggregate lexemes (each <= 64 chars by _AGGREGATE).", + "family": "PUF and downstream transfer" + }, + { + "constant": "puf59_canonical_artifact.py:17 MAX_BODY", + "value": "128 * 1024 * 1024 = 134,217,728; used both directly and as the derived row ceiling MAX_BODY // (len(NAMES) * 8) = 134,217,728 // 512 = 262,144, where len(NAMES) = 5 PREFIX_NAMES (puf59_canonical.py:29-35) + 59 OUTPUTS (puf_target2024_growth.py:30, ordered_targets) = 64", + "enforced_at": "puf59_canonical_artifact.py:39-44 (receipt rows); puf59_canonical_artifact.py:66-69 (_validate rows); puf59_canonical_artifact.py:150 (encode, len(body) <= MAX_BODY); puf59_canonical_artifact.py:167-170 (decode, whole-payload bound); puf59_canonical_artifact.py:212-215 (decode, header rows)", + "refusal_code": "\"PUF59_ARTIFACT_RECEIPT_CONTRACT\" (:44), \"PUF59_ARTIFACT_ROWS\" (:68 and :214), \"PUF59_ARTIFACT_SIZE\" (:150 and :169); _require (puf59_canonical_artifact.py:23-25) raises a bare ValueError", + "protects": "memory-or-time", + "protects_detail": "The derived form MAX_BODY // (len(NAMES) * 8) converts a byte budget into a row ceiling for a dense 64-column float64/int64 body; the body is then required to be exactly n * 8 * len(NAMES) bytes (:216-219) and columns are np.frombuffer views at offset i * n * 8 (:221-223). The row count itself travels as a JSON integer, not a fixed-width field (only the header LENGTH uses struct \" bare ValueError", + "protects": "memory-or-time", + "protects_detail": "Bounds the JSON header (schema, rows, the 64 names, the 64 dtypes, body digest and the growth receipt \u2014 built at :138-146) before json.loads. The struct \" _INT64_MAX)", + "refusal_code": "No short code \u2014 a prose ValueError: \"PUF support ID remap would overflow int64: max source ID ... exceeds {_INT64_MAX}. Assembled source IDs must not exceed {PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID}.\" Exception type: bare ValueError.", + "protects": "encoding-width", + "protects_detail": "The comment at puf_support.py:2318-2324 says it directly: \"Assembly enforces this bound; _remap_ids re-checks it so a violation is a governed ValueError, never an OverflowError.\" It is the int64 storage width of the id columns, re-proved at the point the clone shift is applied.", + "counted_quantity": "clone_index * id_multiplier + max entity id for the frame being cloned, where id_multiplier = 10**digits(max_id) (_id_multiplier_for_values, puf_support.py:2334-2345).", + "count_at_tenth": "about 1.7e6 (shift 10**6 + max id ~694,277)", + "count_at_full_source": "about 1.7e7 (shift 10**7 + max id ~6,942,766)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; twelve orders of magnitude of headroom at full source. It is the backstop for PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID, which itself is the operative bound.", + "family": "PUF and downstream transfer", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "graph_survey_population.py:61 PREPARATION_MAX_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "graph_survey_population.py:305 (payload), graph_survey_population.py:309 (context), graph_survey_population.py:470 (_bounded_json of the outgoing frame context), graph_survey_population.py:577 (kernel ctor, preparation bytes), graph_survey_population.py:582 (kernel ctor, context bytes)", + "refusal_code": "\"PREPARATION_BYTES\" and \"CONTEXT_BYTES\" (lines 305/309/577/582); \"TRANSPORT_LIMIT\" from _bounded_json (line 470, via graph_survey_population.py:263). Exception type: SurveyPopulationGraphError(ValueError), graph_survey_population.py:105.", + "protects": "memory-or-time", + "protects_detail": "It is a transport ceiling on two distinct byte strings. The context string (graph_context.encode_us_frame_context, graph_context.py:123-160) is O(1) in rows \u2014 per entity it emits only {\"columns\", \"rows\", \"id_dtype\", \"ordered_ids_sha256\"} (graph_context.py:104-120) \u2014 so CONTEXT_BYTES can never bind. The payload is the whole-roster preparation receipt built by survey_population_preparation._roster_payload (survey_population_preparation.py:2010), which carries one record per selected household, per person and per entity row. Its own producer bounds it at MAX_ROSTER_BYTES = 4 GiB; this adapter then re-bounds the same bytes at 64 MiB, 64x tighter, so this is the number that actually stops a run. Nothing in the encoding depends on it.", + "counted_quantity": "Bytes of the canonical preparation receipt for the SELECTED/stacked survey (not the clone, not an upstream file); separately, bytes of the US frame-context document.", + "count_at_tenth": "Preparation receipt ~109-128 MB at 158,738 selected households. Measured basis: the recovered 1/1000 artifact preparation.json is 1,136,063 B for 1,584 selected households (sha256 34b362d8...), of which ~45,922 B is fraction-invariant (producer 21,135, storage_transitions 13,268, selection.excluded 7,250 for all 65 full-source exclusions, selection.cells 923, source_files 1,004, catalogues 726, native 413, digests/protocol/request ~1,203); the remaining 1,090,141 B is 688 B per selected household at 1/1000 id widths, rising to ~730-806 B at 6-7 digit ids. Frame context: a few kB, unchanged.", + "count_at_full_source": "Preparation receipt ~1.2-1.3 GB at 1,587,376 selected households (the module's own comment at survey_population_preparation.py:51-53 states 1.02 GiB, citing docs/us-native-scale-transport.md \u00a71). Frame context: a few kB.", + "binds_at_tenth": true, + "binds_at_full_source": true, + "binding_note": "BINDS at both 1/10 and full source, and it is the FIRST refusal in the whole family: run_authenticated_survey_population calls _checked_preparation (graph_survey_population.py:915) before allocation_instructions, before the clone, before the budget. It refuses at roughly 83,000-97,000 selected households (survey_population_preparation.py:52 says \"a single bounded encode refuses at 6.1% of the source\", i.e. ~97k). No tighter upstream limit exists: the preparation module's own MAX_ROSTER_BYTES is 64x looser, and survey_catalogue_selection's batching caps only the classification working set, not the receipt. At 1/100 (15,874 households, ~11 MB) it does not bind.", + "family": "Survey population, origins and age", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "survey_origin_budget.py:54 MAX_PAYLOAD_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "survey_origin_budget.py:612 (the streaming append of the budget document), survey_origin_budget.py:86 (_json -> graph._bounded_json(value, MAX_PAYLOAD_BYTES), enforced at graph_survey_population.py:263), survey_origin_budget.py:931 (candidate bytes)", + "refusal_code": "\"TRANSPORT_LIMIT\" (line 612, and line 86 via _bounded_json) and \"CANDIDATE_LIMIT\" (line 931). Exception types: SurveyOriginBudgetError(ValueError) at survey_origin_budget.py:66 for lines 612/931; SurveyPopulationGraphError(ValueError) for the _bounded_json path.", + "protects": "memory-or-time", + "protects_detail": "An explicit byte ceiling on the sampling-origin budget document, accumulated one record at a time into a bytearray (survey_origin_budget.py:608-620). Deliberately left at 64 MiB by commit 14defbfc0, whose message says it \"is a byte transport, not a row count, and is deliberately left alone: it takes the segmented transport argument, not this one.\" No reader depends on the value; the artifact is plain canonical JSON re-decoded by graph_survey_budget.py:141-156.", + "counted_quantity": "Bytes of the budget document: one ~720-730 B origin record per SELECTED household, plus two header lists (\"household_ids\", \"group_indices\") each holding one entry per CLONED household row (2x selected).", + "count_at_tenth": "~124 MB at 158,738 selected / 317,476 cloned rows. Basis: I encoded a representative record with the exact field set built at survey_origin_budget.py:564-580 plus the eleven _reference fields (survey_origin_budget.py:114-126) at full-source id widths \u2014 728 B canonical, ~723 B at 6-digit ids \u2014 times 158,738, plus 2 x 317,476 ids/groups at ~7 B each in two lists (~8.9 MB).", + "count_at_full_source": "~1.21 GB at 1,587,376 selected / 3,174,752 cloned rows (1.156 GB of origin records + 50.8 MB of id and group-index lists).", + "binds_at_tenth": true, + "binds_at_full_source": true, + "binding_note": "BINDS at both 1/10 and full source; refuses at roughly 88,000 selected households. It does NOT bind at 1/100 (~12 MB). It is not the first refusal in a full run \u2014 graph_survey_population.py:61 PREPARATION_MAX_BYTES stops the same run slightly earlier, at ~83k-97k households \u2014 but it is independent of that one: freeze_survey_origin_budget takes the preparation through survey_population_preparation's own checked_view (survey_origin_budget.py:933-935), which does NOT apply the 64 MiB preparation bound, so a caller that reaches the budget without the graph adapter hits this ceiling on its own. Sibling MAX_GROUPS in the same module was lifted to 7,000,000 by commit 14defbfc0 and no longer binds, which leaves this constant as the module's only live blocker.", + "family": "Survey population, origins and age", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "graph_survey_age_artifact.py:33 MAX_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "graph_survey_age_artifact.py:57-59 (the HEADER + ROW_BYTES * rows budget in _shape, reached from _validate at :73, decode at :117 and the kernel at :172), graph_survey_age_artifact.py:100 (whole payload on decode), graph_survey_age_artifact.py:211 (whole payload on encode)", + "refusal_code": "\"LIMIT\" (lines 58 and 211) and \"PAYLOAD\" (line 100), raised as ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason) \u2014 plain ValueError, graph_survey_age_artifact.py:43-45.", + "protects": "encoding-width", + "protects_detail": "Line 58 is exactly the HEADER + ROW_BYTES * N shape the task calls out: len(MAGIC)=9 + 4 (big-endian header length) + MAX_HEADER_BYTES=65,536 + rows * _WIDTH * 8 <= MAX_BYTES, with _WIDTH = 1 + 18 = 19 (graph_survey_age_artifact.py:36-37; national_age_activation.py:309 asserts exactly 18 bands). The reader at :119-122 relies on that layout: it asserts len(payload) - offset == rows * _WIDTH * 8 and then np.frombuffer(..., \"/selection-request.json.", + "count_at_tenth": "168 bytes (measured: the recovered artifact's source_files[0] records [\"selection-request.json\", 168, \"33f8ec12...\"]). Identical at every fraction \u2014 only the two integers in \"fraction\" change width.", + "count_at_full_source": "168 bytes, give or take a few characters for the fraction literal.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds, at any scale; 24x headroom over a document whose shape is fixed at four fields.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_population_domains.py:19 MAX_HOUSEHOLDS", + "value": "100_000", + "enforced_at": "survey_population_domains.py:551; also asserted as an upper bound on the batch size at survey_catalogue_selection.py:129", + "refusal_code": "\"HOUSEHOLD_BATCH\" (line 551). Exception type: DomainInputError(ValueError), survey_population_domains.py:23. The :129 assertion raises CatalogueSelectionError(ValueError) with \"BATCH_LIMITS\".", + "protects": "memory-or-time", + "protects_detail": "A ceiling on one classify_households call's working set \u2014 the function builds key and person-key sets over the batch (:552-576) and then materializes one HouseholdDecision per row, each carrying its full observation and annotation tuples. Purely a resource ceiling; nothing downstream reads it.", + "counted_quantity": "len(rows) in ONE classify_households batch \u2014 not the supplied catalogue and not the selected roster.", + "count_at_tenth": "<= 10,000. The single production caller is survey_catalogue_selection.plan_catalogue_selection's consume() (:139), fed from a batch flushed at survey_catalogue_selection.py:196-201 whenever len(batch) == _BATCH_HOUSEHOLDS = 10,000.", + "count_at_full_source": "<= 10,000, unchanged \u2014 the batch size is independent of the catalogue size; the full source is simply classified in ~159 batches.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "NEVER BINDS, at any fraction, because of a tighter upstream streaming limit \u2014 this is the pattern the task asked me to verify, and I confirmed it: survey_catalogue_selection.py:19 _BATCH_HOUSEHOLDS = 10,000 caps every batch at exactly one tenth of this ceiling, survey_catalogue_selection.py:129 statically asserts _BATCH_HOUSEHOLDS <= min(10_000, MAX_HOUSEHOLDS), and a repo-wide grep finds no other production caller of classify_households (only two references in microcosm-build/tests/test_us_acs_population_catalogue.py).", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_population_domains.py:20 MAX_TOTAL_MEMBERS", + "value": "1_000_000", + "enforced_at": "survey_population_domains.py:563; also asserted as an upper bound on the people batch at survey_catalogue_selection.py:131", + "refusal_code": "\"BATCH_MEMBER_BOUND\" (line 563). Exception type: DomainInputError(ValueError), survey_population_domains.py:23. The :131 assertion raises CatalogueSelectionError(ValueError) with \"BATCH_LIMITS\".", + "protects": "memory-or-time", + "protects_detail": "A running person-count ceiling over one classification batch, checked incrementally as members accumulate. Guards the per-person decision objects _person builds for every row in the batch. Purely a resource ceiling.", + "counted_quantity": "Sum of len(row.persons) over ONE classify_households batch \u2014 not the stacked or selected person roster.", + "count_at_tenth": "<= 100,000. plan_catalogue_selection flushes whenever batch_people + len(row.persons) would exceed _BATCH_PEOPLE = 100,000 (survey_catalogue_selection.py:198), and the per-household member bound at :186-189 caps a single household at min(MAX_MEMBERS, _BATCH_PEOPLE) = 20, so a batch can never overshoot.", + "count_at_full_source": "<= 100,000, unchanged. (For scale, the full supplied catalogue holds 3,565,013 persons \u2014 ACS 3,422,888 plus ASEC 142,125 \u2014 but it is never presented to this function at once.)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "NEVER BINDS, at any fraction, for the same verified reason as MAX_HOUSEHOLDS: survey_catalogue_selection.py:20 _BATCH_PEOPLE = 100,000 is exactly one tenth of this ceiling and is enforced by the flush at survey_catalogue_selection.py:196-201, with the static relation asserted at :131. The transport lane's claim about this constant and MAX_HOUSEHOLDS is correct as stated.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_population_domains.py:18 MAX_MEMBERS", + "value": "20", + "enforced_at": "survey_population_domains.py:244 (reduce_qualifiers roster), survey_population_domains.py:358-361 (_members), survey_population_domains.py:558-561 (classify_households); also survey_catalogue_selection.py:186-189 as min(domains.MAX_MEMBERS, _BATCH_PEOPLE)", + "refusal_code": "\"QUALIFIER_ROSTER\" (line 244), \"MEMBER_TUPLE_BOUND\" (lines 360 and 560). Exception type: DomainInputError(ValueError), survey_population_domains.py:23. The selection-side site raises CatalogueSelectionError(ValueError) with \"HOUSEHOLD_MEMBER_BOUND\".", + "protects": "structural-invariant", + "protects_detail": "A domain invariant on one household's roster, not a scale bound. It matches the source coding: the ACS NP field is parsed with two digits and a maximum of 20 at survey_population_domains.py:459-464, and ASEC H_NUMPER with a maximum of 16 at the same site; the ACS SPORDER line number is likewise bounded at 20 (:369-373). Moving it would contradict the published field ranges.", + "counted_quantity": "Persons in ONE household.", + "count_at_tenth": "<= 20 by the source's own field width; identical at every fraction.", + "count_at_full_source": "<= 20.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a scale bound and cannot bind: the NP/H_NUMPER parse at :459-464 already rejects anything above 20/16 before this check, and the value is the publisher's own top code.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_population_domains.py:17 MAX_TOKEN_CHARS", + "value": "128", + "enforced_at": "survey_population_domains.py:187 (_token, called from _key at :204, _person at :291/:296/:329, _unsigned at :192, and classify_households at :569); also used as the digit width for HSUP_WGT at survey_population_domains.py:500", + "refusal_code": "\"LITERAL_TYPE_SIZE\". Exception type: DomainInputError(ValueError), survey_population_domains.py:23.", + "protects": "structural-invariant", + "protects_detail": "A per-literal character ceiling applied to every raw source token before it is parsed or re-emitted. It bounds one field, never a count of rows. The same 128 reappears as the admitted key width at graph_survey_population.py:143 (0 < len(row.key.native_id) <= 128).", + "counted_quantity": "Characters of one raw source literal (SERIALNO, H_SEQ, PERIDNUM, AGEP, WGTP, ...).", + "count_at_tenth": "<= 22 in practice: ACS SERIALNO is fixed at 13 by the regex at :208, ASEC H_SEQ at <= 5 digits (:219, maximum 99999), PERIDNUM at exactly 22 (:381). Identical at every fraction.", + "count_at_full_source": "<= 22.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; ~6x headroom over the widest literal the source defines, and literal widths do not grow with the roster.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_catalogue_selection.py:19 _BATCH_HOUSEHOLDS", + "value": "10_000", + "enforced_at": "survey_catalogue_selection.py:127-133 (static self-consistency assertion), survey_catalogue_selection.py:196-201 (the flush that actually caps the batch); no data-driven refusal of its own", + "refusal_code": "\"BATCH_LIMITS\" at line 132 \u2014 but that assertion is about the constants themselves (0 < _BATCH_HOUSEHOLDS <= min(10_000, domains.MAX_HOUSEHOLDS)), not about any supplied data. Exception type: CatalogueSelectionError(ValueError), survey_catalogue_selection.py:23.", + "protects": "memory-or-time", + "protects_detail": "A streaming batch size, not a ceiling: when the batch reaches 10,000 households the loop calls consume() and resets rather than refusing (:196-201). It exists so that classify_households sees a bounded working set, and it is precisely what makes survey_population_domains.MAX_HOUSEHOLDS unreachable.", + "counted_quantity": "Households held in memory between two consume() calls.", + "count_at_tenth": "Exactly <= 10,000 at any fraction; the full catalogue is classified in ceil(1,587,376 / 10,000) ~ 159 batches regardless of the selection fraction, because the draw happens after classification.", + "count_at_full_source": "Exactly <= 10,000; ~159 batches.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Cannot bind on data \u2014 it flushes rather than refuses. Reported because it is the upstream limit that makes two nominal ceilings in this family dead: survey_population_domains.MAX_HOUSEHOLDS (100,000) and, with its sibling, MAX_TOTAL_MEMBERS (1,000,000). Raising it would make those two live.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_catalogue_selection.py:20 _BATCH_PEOPLE", + "value": "100_000", + "enforced_at": "survey_catalogue_selection.py:127-133 (static self-consistency assertion), survey_catalogue_selection.py:186-189 (per-household member bound, as min(domains.MAX_MEMBERS, _BATCH_PEOPLE)), survey_catalogue_selection.py:196-201 (the flush)", + "refusal_code": "\"BATCH_LIMITS\" at line 132 (constants only) and \"HOUSEHOLD_MEMBER_BOUND\" at line 188 (which resolves to min(20, 100_000) = 20, i.e. MAX_MEMBERS does the work). Exception type: CatalogueSelectionError(ValueError), survey_catalogue_selection.py:23.", + "protects": "memory-or-time", + "protects_detail": "The person-side streaming batch size. Like its sibling it flushes rather than refuses; its only refusal expression collapses to the structural 20-member bound.", + "counted_quantity": "Persons held in memory between two consume() calls.", + "count_at_tenth": "Exactly <= 100,000 at any fraction. The full supplied catalogue's 3,565,013 persons (ACS 3,422,888 + ASEC 142,125, both from the recovered artifact's catalogues block) are streamed through in batches.", + "count_at_full_source": "Exactly <= 100,000.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Cannot bind on data. Reported because it is the upstream limit that makes survey_population_domains.MAX_TOTAL_MEMBERS (1,000,000) dead at 10x headroom.", + "family": "Survey population, origins and age" + }, + { + "constant": "graph_survey_population.py:62 ALLOCATION_MAX_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "graph_survey_population.py:385-388 (the segment-closing test), graph_survey_population.py:391-392, graph_survey_population.py:401", + "refusal_code": "\"ALLOCATION_LIMIT\". Exception type: SurveyPopulationGraphError(ValueError), graph_survey_population.py:105.", + "protects": "memory-or-time", + "protects_detail": "A per-segment accumulation ceiling, explicitly documented as such at :63-67 and :356-366: \"ALLOCATION_MAX_BYTES stays the per-segment ceiling, unchanged; the total is a separate, explicit resource ceiling.\" A segment closes before it would pass the bound (:385-390), so the refusal at :391 can only fire if one instruction document plus the tail exceeds a whole segment. Segment boundaries are invisible downstream \u2014 the segments are joined at :406 and the artifact digest is unchanged.", + "counted_quantity": "Bytes of one open allocation segment; equivalently ~211,929 selected households per segment at the measured 316.665 B per instruction document (figure from the comment at :63-65, citing docs/us-native-scale-transport.md \u00a71).", + "count_at_tenth": "One instruction document ~317 B against a 64 MiB segment; the 158,738-household roster fits in a single segment (~50 MB).", + "count_at_full_source": "~503 MB of instruction documents across ~8 segments at 1,587,376 households; no segment exceeds 64 MiB.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds on roster scale \u2014 it partitions rather than refuses. Its refusal expression is only reachable for a single pathological instruction document, and every field in one is bounded (native_id <= 128 chars at :143, all other members are integer pairs).", + "family": "Survey population, origins and age" + }, + { + "constant": "graph_survey_population.py:68 ALLOCATION_ROSTER_BYTES", + "value": "64 * ALLOCATION_MAX_BYTES = 64 * 67,108,864 = 4,294,967,296", + "enforced_at": "graph_survey_population.py:162 (record-count form: len(plan.selected) <= ALLOCATION_ROSTER_BYTES // 128 = 33,554,432), graph_survey_population.py:394-396, graph_survey_population.py:405, graph_survey_population.py:407", + "refusal_code": "\"ALLOCATION_LIMIT\" (line 162) and \"ALLOCATION_ROSTER_LIMIT\" (lines 396, 405, 407). Exception type: SurveyPopulationGraphError(ValueError), graph_survey_population.py:105.", + "protects": "memory-or-time", + "protects_detail": "The explicit total-resource ceiling for the allocation artifact, introduced alongside the segmented encoder. The :158-163 comment states the record form's purpose precisely: the implied 128-byte row budget \"is below the payload's 220-byte structural overhead alone, so this never binds the payload; it bounds the two maps below, and its ceiling is the roster one.\"", + "counted_quantity": "(a) at :162, selected households, capped at 33,554,432; (b) at :396/:405/:407, total bytes of the allocation payload at ~316.665 B per selected household.", + "count_at_tenth": "(a) 158,738 selected households; (b) ~50.3 MB.", + "count_at_full_source": "(a) 1,587,376 selected households; (b) ~502.6 MB, 11.7% of the ceiling. The byte form would refuse at ~13,563,442 selected households.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at either scale, on either form \u2014 21x headroom on the record count and 8.5x on the bytes. It is also unreachable as a first refusal: run_authenticated_survey_population calls _checked_preparation (which enforces the 64 MiB PREPARATION_MAX_BYTES) at :915, before allocation_instructions at :922.", + "family": "Survey population, origins and age" + }, + { + "constant": "graph_survey_age_artifact.py:35 MAX_PEOPLE", + "value": "10_000_000", + "enforced_at": "graph_survey_age_artifact.py:56 (inside _shape, reached from _validate at :73, decode at :117, and the kernel at :172), graph_survey_age_artifact.py:168", + "refusal_code": "\"SHAPE\" (line 56) and \"PEOPLE_LIMIT\" (line 168), raised as ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason) \u2014 plain ValueError, graph_survey_age_artifact.py:43-45.", + "protects": "memory-or-time", + "protects_detail": "An explicit ceiling on the person table handed to the age-count kernel, and on the \"people\" total declared in the artifact header. The comment at :84 notes it is also what keeps the per-cell count sum well inside int64, but the int64 headroom is enormous and is not the operative reason; the bound itself is a resource ceiling.", + "counted_quantity": "len(context.tables[\"person\"]) and header[\"people\"] for the CLONED population (the count node runs on clone.COMBINED_CLONE_NODE; survey_age_calibration.py:102 and :317).", + "count_at_tenth": "694,277 cloned persons (347,138 stacked x 2).", + "count_at_full_source": "~6,942,766 cloned persons (~3,471,383 stacked x 2), 69% of the ceiling.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind at either scale \u2014 but with only 1.44x headroom at full source it is the closest non-binding person-side bound in the family. It is in any case shadowed: MAX_BYTES at :33 refuses the same artifact at 441,074 cloned household rows, i.e. at roughly 220,537 selected households, long before the person count matters.", + "family": "Survey population, origins and age" + }, + { + "constant": "graph_survey_age_artifact.py:34 MAX_HEADER_BYTES", + "value": "65_536", + "enforced_at": "graph_survey_age_artifact.py:57-59 (reserved in full inside the row budget), graph_survey_age_artifact.py:105-108 (declared header length on decode), graph_survey_age_artifact.py:204 (produced header on encode)", + "refusal_code": "\"LIMIT\" (line 58), \"HEADER_LIMIT\" (lines 106 and 204), raised as ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason) \u2014 plain ValueError, graph_survey_age_artifact.py:43-45.", + "protects": "encoding-width", + "protects_detail": "Part of the artifact's wire format. The payload is MAGIC (9 B) + a 4-byte big-endian header length + the canonical header JSON + the int64 matrix (:205-210), and the decoder at :104-108 refuses any declared length above this before slicing. It is also charged in full against the row budget at :58 whether or not the real header uses it, so it costs 65,536 B = 431 rows of capacity. Changing it changes what a reader will accept.", + "counted_quantity": "Bytes of the canonical header JSON: {protocol, population, columns (18 band names), rows, people, age_convention}.", + "count_at_tenth": "~600 B \u2014 six fields, of which only the 18 fixed column names are large; \"rows\" and \"people\" grow by a handful of digits. Same at every fraction.", + "count_at_full_source": "~610 B.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds directly (over 100x headroom over a fixed-shape header). It does, however, consume 431 rows of the MAX_BYTES budget at :58 unconditionally, which is why the effective row ceiling is 441,074 rather than 441,505.", + "family": "Survey population, origins and age", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "survey_atomic_geography.py:63 MAX_SUPPORT_BYTES", + "value": "RAW_BYTES_MAX_BYTES, defined at packages/microcosm-graph/src/microcosm/graph/codecs.py:78 as 64 * 1024 * 1024 = 67,108,864", + "enforced_at": "survey_atomic_geography.py:151-154", + "refusal_code": "\"SUPPORT_SIZE\", raised as ValueError(\"ATOMIC_SURVEY_RECONSTRUCTION_\" + reason) \u2014 plain ValueError, survey_atomic_geography.py:67-69.", + "protects": "upstream-file-size", + "protects_detail": "This bounds a real on-disk publisher-derived input \u2014 the national atomic geography support archive named by AtomicSurveyReconstruction.support_path and pinned by support_sha256 \u2014 read through an O_NOFOLLOW descriptor with before/after stat identity checks (:143-159). It is an alias of the shared raw-bytes codec limit, so moving it moves every raw-bytes source in the repo at once (atomic_block_sources.py:365, atomic_block_api_sources.py:264, puf55_survey_recipients.py:422/431 all use the same constant). It is NOT a survey roster bound and must not be moved to make a larger survey fit.", + "counted_quantity": "Bytes of the atomic-support .npz file (national block/tract/PUMA/district geography), which is invariant to the survey selection fraction.", + "count_at_tenth": "28,862,508 bytes \u2014 measured directly on the real file used by this lane's pilots (national-atomic-support.npz, sha256 pinned in execution-config.proposed.json as 5edc0e77471ba31d550a1eed416d5b46ada0a35425718eb87cfabe4d66fe4960; identical copies at /Users/maxghenis/PolicyEngine/_recovered/microcosm-native-sources-20260909/ and four pilot run roots). 43% of the ceiling.", + "count_at_full_source": "28,862,508 bytes \u2014 unchanged. National geography does not scale with the survey fraction.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds, and never will as a function of survey scale: the counted file is national geography, not a survey roster. Do not move it as part of a row-ceiling lift.", + "family": "Survey population, origins and age", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "survey_atomic_geography.py:64 MAX_SUPPORT_EXPANDED_BYTES", + "value": "2 * 1024**3 = 2,147,483,648", + "enforced_at": "survey_atomic_geography.py:173-177", + "refusal_code": "\"SUPPORT_EXPANDED_SIZE\", raised as ValueError(\"ATOMIC_SURVEY_RECONSTRUCTION_\" + reason) \u2014 plain ValueError, survey_atomic_geography.py:67-69.", + "protects": "memory-or-time", + "protects_detail": "The comment at :168-169 is explicit: \"Bound declared decompressed bytes before numpy opens any array.\" It sums ZipInfo.file_size over the archive's members and refuses before any np.lib.format header is read, guarding against a small archive declaring a huge allocation (the per-member cross-check at :185-190 then proves the declared shape matches the physical member length). It is an allocation ceiling, not a statement about the file.", + "counted_quantity": "Sum of declared uncompressed sizes of the support archive's ZIP members; invariant to the survey fraction.", + "count_at_full_source": "1,061,673,840 bytes \u2014 measured from the real archive: area.npy 346,196,648 + tract.npy 253,877,576 + puma.npy 161,558,504 + county.npy 115,398,968 + district.npy 92,319,200 + population.npy 46,159,664 + state.npy 46,159,664 + metadata_json.npy 3,616. That is 49.4% of the ceiling.", + "count_at_tenth": "1,061,673,840 bytes \u2014 unchanged.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind, but it is the tightest non-binding bound in the whole family in relative terms: the real archive already occupies 49.4% of the ceiling. It has nothing to do with the survey roster, so a row-ceiling lift must not touch it \u2014 but it is worth flagging separately, because a doubling of national block geography would trip it.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_atomic_geography.py:65 MAX_SUPPORT_MEMBERS", + "value": "32", + "enforced_at": "survey_atomic_geography.py:173-177 (same require as MAX_SUPPORT_EXPANDED_BYTES)", + "refusal_code": "\"SUPPORT_EXPANDED_SIZE\", raised as ValueError(\"ATOMIC_SURVEY_RECONSTRUCTION_\" + reason) \u2014 plain ValueError, survey_atomic_geography.py:67-69.", + "protects": "structural-invariant", + "protects_detail": "Bounds how many ZIP members the support archive may declare, so that the per-member NPY header walk at :178-190 is a bounded loop. The archive's member set is a fixed schema of named geography arrays, so this is a shape bound rather than a scale one.", + "counted_quantity": "Number of members in the support archive; invariant to the survey fraction.", + "count_at_full_source": "8 members (area, county, district, metadata_json, population, puma, state, tract) \u2014 measured on the real file.", + "count_at_tenth": "8 members \u2014 unchanged.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; 4x headroom over a fixed eight-array schema. Unrelated to survey scale.", + "family": "Survey population, origins and age" + }, + { + "constant": "survey_financial_successor.py:26 MAX_PAYLOAD_BYTES", + "value": "256 * 1024 = 262,144", + "enforced_at": "survey_financial_successor.py:211-215 (candidate bytes), survey_financial_successor.py:232 (the produced document)", + "refusal_code": "\"CANDIDATE_TYPE_OR_BOUND\" (line 213) and \"PAYLOAD_BOUND\" (line 232), raised as SurveyFinancialSuccessorError(\"SURVEY_FINANCIAL_SUCCESSOR_\" + reason) \u2014 SurveyFinancialSuccessorError(ValueError), survey_financial_successor.py:30.", + "protects": "memory-or-time", + "protects_detail": "A byte ceiling on the financial-successor admission document built at survey_financial_successor.py:116-131. I traced every field: protocol, four sha256 digests, two version strings, owned_columns, three booleans, and the nested \"financial_run\" document. That nested document (graph_atomic_survey_financial.py:561-657) is O(graph nodes and frame columns), not O(households) \u2014 its largest members are node_keys, artifact_payload_sha256 (one row per node artifact) and financial_owners (one entry per (entity, column) cell). Nothing in it carries a per-household row.", + "counted_quantity": "Bytes of the successor document. Its dominant term is the frame's cell inventory, not its row count.", + "count_at_tenth": "~20-40 kB. Basis: the recovered pilot manifest records 256 artifact cells for survey_population.allocate and survey_population.create and 259 for combined_survey_puf_support_clone, so financial_owners is ~260 entries; at ~60 B each that is ~16 kB, plus ~20 graph nodes of keys and digests.", + "count_at_full_source": "~20-40 kB \u2014 unchanged. The cell inventory and node count do not vary with the selection fraction.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds as a function of survey scale; the document is row-independent. It would only become live if the frame's column inventory grew roughly tenfold.", + "family": "Survey population, origins and age" + }, + { + "constant": "native_household_origin.py:48 MAX_SOURCE_BYTES", + "value": "512 * 1024**2 = 536,870,912", + "enforced_at": "native_household_origin.py:125-131 (AsecNativeMemberPin.__post_init__, on a declared upstream member file's size) and native_household_origin.py:196 (_source_document, on the derived origin JSON)", + "refusal_code": "\"ASEC_MEMBER_BOUNDS\" (line 130) and \"SOURCE_SIZE\" (line 196). Exception type: NativeOriginError(ValueError), native_household_origin.py:51.", + "protects": "upstream-file-size", + "protects_detail": "Two sites with genuinely different characters, which is why I flag it. At :127 it bounds pin.size_bytes \u2014 the declared byte size of a real pinned Census household CSV, cross-checked against the archive digest at :113-120 and re-proved by hashing the captured file at :353-355. That arm is a structural assertion about the publisher's file and MUST NOT be moved. At :196 it bounds the derived AuthenticatedNativeOriginSource JSON, which is a memory-or-time ceiling on a document whose row count is nevertheless fixed by the upstream files, not by the survey fraction. I classify the constant as upstream-file-size because the :127 arm is the one that constrains what may be changed.", + "counted_quantity": "(a) at :127, bytes of one pinned ASEC hhpub CSV; (b) at :196, bytes of the origin JSON \u2014 one [period, native_id, member_index, origin_key] record per ACS housing record or per ASEC household row across all three pinned cohorts. Both are full upstream rosters, NOT the selected survey.", + "count_at_tenth": "Identical to full source at both sites \u2014 neither depends on the survey fraction. (a) 30,259,450 / 30,448,089 / 33,298,479 bytes for hhpub23/24/25, read straight off the pins at native_household_origin.py:145-166 and corroborated by the recovered artifact's source_files entry for asec/hhpub25.csv (33,298,479). (b) ACS ~150 MB, ASEC ~22 MB.", + "count_at_full_source": "(a) 33,298,479 bytes maximum, 6.2% of the ceiling. (b) ACS document ~150 MB: 1,631,969 ACS housing records (the recovered artifact's catalogues.acs.counts.households) at ~92 B per canonical record ([2024,\"2024HU1234567\",0,\"<64 hex>\"]); ASEC document ~22 MB: 88,978 + 89,473 + 88,932 = 267,383 household rows (the pins' own rows fields) at ~82 B. The ACS arm is ~28% of the ceiling.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at either site, at either scale. Critically, neither counted quantity scales with the survey selection fraction: produce_acs_native_origins (:264-292) runs over the full ACS housing universe with serialnos=None, and produce_asec_native_origins (:294-357) over all three complete pinned hhpub CSVs. A row-ceiling lift has no reason to touch this constant, and the :127 arm must not be touched at all.", + "family": "Survey population, origins and age", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "current_survey_geography.py:29 MAX_HOUSEHOLDS", + "value": "64 * 1024**2 // 128 = 524,288", + "enforced_at": "current_survey_geography.py:68 (_project: `0 < len(origins) <= MAX_HOUSEHOLDS` and `len(origins) == len(households)`); current_survey_geography.py:188 (_projection_digest: `0 < len(household) <= MAX_HOUSEHOLDS`)", + "refusal_code": "\"HOUSEHOLD_COUNT\" (current_survey_geography.py:70) and \"PROJECTION_STORAGE\" (current_survey_geography.py:190); _require at current_survey_geography.py:33-35 raises ValueError(\"CURRENT_SURVEY_GEOGRAPHY_\" + reason)", + "protects": "memory-or-time", + "protects_detail": "Written as a 64 MiB byte budget divided by an assumed 128 bytes per household row. Nothing reads it as a field width: both sites only compare len() of a Python list and of a pandas DataFrame, and the receipt written at :255-283 carries only aggregate counts. No consumer of the projection re-derives this number, so it is an arbitrary-but-deliberate resource ceiling and is movable.", + "counted_quantity": "The selected/stacked household roster of the prepared survey population: `origins[\"households\"]` decoded from the preparation payload (current_survey_geography.py:228) and `state.frame.table(\"household\")` (:239), which carries both the acs and asec support channels (:80). Pre-clone.", + "count_at_tenth": "158,738 households", + "count_at_full_source": "1,587,376 households", + "binds_at_tenth": false, + "binds_at_full_source": true, + "binding_note": "Binds at full source (1,587,376 > 524,288); first binds around sample fraction 0.33. No tighter upstream preempts it: the same origins roster is bounded upstream only in bytes, by ORIGIN_LIMIT against MAX_ROSTER_BYTES = 64 * MAX_SEGMENT_BYTES = 4 GiB (survey_population_preparation.py:58, enforced at :1149 and :259), which at roughly 120 bytes per origin row is about 190 MB at full source, far from 4 GiB. I verified the transport lane's batching claim for the neighbouring domains bounds: survey_population_domains.MAX_HOUSEHOLDS = 100,000 is enforced on one batch only (survey_population_domains.py:551, \"HOUSEHOLD_BATCH\") and MAX_TOTAL_MEMBERS = 1,000,000 likewise (:563, \"BATCH_MEMBER_BOUND\"), while survey_catalogue_selection.py:129-131 pins _BATCH_HOUSEHOLDS = 10,000 and _BATCH_PEOPLE = 100,000 at or below those \u2014 so neither can preempt this module's whole-roster ceiling.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_survey_geography.py:30 MAX_RECEIPT_BYTES", + "value": "64 * 1024 = 65,536", + "enforced_at": "current_survey_geography.py:282 (`maximum=MAX_RECEIPT_BYTES` passed to source._encode for the geography receipt)", + "refusal_code": "\"PAYLOAD_LIMIT\", raised inside survey_population_preparation._encode at survey_population_preparation.py:117; _require at :82-84 raises SurveyPopulationPreparationError (a ValueError subclass, survey_population_preparation.py:78)", + "protects": "memory-or-time", + "protects_detail": "A byte ceiling on one receipt document. The document built at current_survey_geography.py:255-283 has a fixed key roster of scalars, digests and aggregate counts (households, acs_households, asec_households, state_unknown_households, puma_observed_households); no row, id or per-household value is written into it, so its size does not scale with the roster.", + "counted_quantity": "Bytes of the encoded geography receipt JSON (a fixed ~26-key aggregate document).", + "count_at_tenth": "roughly 1 KB (fixed-size document; the only varying parts are five integer counts)", + "count_at_full_source": "roughly 1 KB (same document; integers grow by a few digits)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at any survey scale, because the counted document is fixed-shape. It would only bind if row-level detail were added to the receipt, which the module's scope fence forbids.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_survey_household_roles.py:86 MAX_PERSONS", + "value": "64 * 1024**2 // 32 = 2,097,152", + "enforced_at": "current_survey_household_roles.py:667 (_origins: `0 < len(original) <= MAX_PERSONS`)", + "refusal_code": "\"ORIGIN_ROSTER\" (current_survey_household_roles.py:668); require at :238-240 raises ValueError(\"CURRENT_SURVEY_HOUSEHOLD_ROLES_\" + reason)", + "protects": "memory-or-time", + "protects_detail": "Again a 64 MiB budget divided by an assumed 32 bytes per person row. The only use is a len() comparison on the rebuilt origins DataFrame; no serialized artifact encodes this number and no reader re-derives it. Deliberate resource ceiling, movable.", + "counted_quantity": "The stacked person roster: one row per selected original person in `document[\"origins\"][\"persons\"][\"rows\"]`, checked row-for-row against `preparation_frame.person` (:673-686). Both acs and asec channels, pre-clone.", + "count_at_tenth": "347,138 persons", + "count_at_full_source": "3,471,383 persons", + "binds_at_tenth": false, + "binds_at_full_source": true, + "binding_note": "Binds at full source (3,471,383 > 2,097,152); first binds around sample fraction 0.60. It is the FIRST scale refusal in this qualifier: qualify_current_survey_household_roles calls _origins at :1127, before _acs_roster at :1136 and before any projection bytes are built at :1151. Consequently it preempts both MAX_ROSTER_ROWS (:799) and the graph fragment's MAX_ARTIFACT_BYTES at full source.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": true + } + }, + { + "constant": "current_survey_household_roles.py:88 MAX_HOUSEHOLD_MEMBERS", + "value": "20 (literal; the comment at :87 calls it \"at most one ACS household roster of the printed line-number domain\")", + "enforced_at": "current_survey_household_roles.py:717 (ACS origin line number); current_survey_household_roles.py:751 (ASEC origin line number); current_survey_household_roles.py:794 (SPORDER of a row read from the pinned ACS person archive)", + "refusal_code": "\"ACS_LINE_KEY\" (:719), \"ASEC_LINE_KEY\" (:753), \"ACS_SOURCE_LINE\" (:796); require at :238-240 raises ValueError(\"CURRENT_SURVEY_HOUSEHOLD_ROLES_\" + reason)", + "protects": "structural-invariant", + "protects_detail": "It bounds a within-household line number (ACS SPORDER, ASEC A_LINENO), i.e. household size in the published line-number domain, not any roster, file or buffer. Every site pairs it with `1 <= int(line)` and a 1-2 digit regex, so it is a domain check on one person's coordinate. The same domain appears independently as 16 in current_child_property_income_source.py:216 and :392 (`_integer_literal(..., \"A_LINENO\", 16)`).", + "counted_quantity": "The printed line number of one person inside one household (equivalently, the maximum number of members a household roster may enumerate).", + "count_at_tenth": "at most 20 by the printed domain, independent of sample size", + "count_at_full_source": "at most 20 by the printed domain, independent of sample size", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a scale bound at all; it cannot bind as the survey grows. Moving it would change what counts as a well-formed Census line number, which is a source-semantics decision, not a capacity one.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_survey_household_roles.py:89 MAX_ROSTER_ROWS", + "value": "MAX_ROSTER_ROWS = MAX_PERSONS, i.e. 64 * 1024**2 // 32 = 2,097,152", + "enforced_at": "current_survey_household_roles.py:799 (_scan_acs_relationships: `require(len(selected) < MAX_ROSTER_ROWS, ...)` checked before each insertion, so at most 2,097,152 rows are retained)", + "refusal_code": "\"ACS_ROSTER_BOUND\" (current_survey_household_roles.py:799); require at :238-240 raises ValueError(\"CURRENT_SURVEY_HOUSEHOLD_ROLES_ACS_ROSTER_BOUND\")", + "protects": "memory-or-time", + "protects_detail": "It caps the in-memory `roster` dict built while streaming the pinned ACS PUMS person member (_acs_roster, :1054-1100). It is an alias of MAX_PERSONS, so it inherits the same 64 MiB / 32 B-per-row derivation; nothing serializes or re-derives it.", + "counted_quantity": "Retained ACS person rows: every member of every SELECTED ACS household, read from the pinned ACS person archive. Selection is by household (whole rosters retained, :760-766), so this is the ACS arm of the stacked roster plus any unselected members of those households.", + "count_at_tenth": "about 332,400 (the pilot's native ACS arm at 1/1000 is 3,324 persons \u2014 /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json, /native/acs/persons)", + "count_at_full_source": "3,422,888 \u2014 every ACS person in the source file (same artifact, /catalogues/acs/counts/people)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds, because a tighter upstream check in the same qualifier fires first: MAX_PERSONS with the identical value 2,097,152 is enforced at current_survey_household_roles.py:667 on the STACKED roster (3,471,383 at full), which is a strict superset of this ACS-only roster (3,422,888), and _origins runs at :1127 before _acs_roster at :1136. A second, non-constant upstream also applies inside the scan: `maximum=document[\"counts\"][\"people\"] - count` (:1088) pins the row budget to the ACS catalogue's own declared person count, refusing with \"ACS_SOURCE_ROW_SHAPE\" (:781) \u2014 that one is a class-(c) assertion about the real PUMS file (3,422,888 rows) and must not be moved.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_survey_household_roles.py:90 MAX_RECEIPT_BYTES", + "value": "64 * 1024 = 65,536", + "enforced_at": "current_survey_household_roles.py:1178 (`maximum=MAX_RECEIPT_BYTES` passed to source._encode for the roles receipt)", + "refusal_code": "\"PAYLOAD_LIMIT\", raised inside survey_population_preparation._encode at survey_population_preparation.py:117; SurveyPopulationPreparationError (ValueError subclass, survey_population_preparation.py:78)", + "protects": "memory-or-time", + "protects_detail": "Bounds one aggregate receipt document (built at :1151-1179). PUBLIC_RECEIPT_KEYS (:205-235) admits only counts, digests, contract references and closed flags. The one embedded sub-document, `acs_relationship_diagnostics`, is a fixed set of scalar counters computed in asec_demographic_source.acs_household_reference_states (asec_demographic_source.py:1033-1048: housing_units, gq_households, unbound_housing_units, bad_reference_cardinality, mixed_households, invalid_households), so the receipt does not grow with rows.", + "counted_quantity": "Bytes of the encoded household-roles receipt JSON (fixed key roster of aggregates and digests).", + "count_at_tenth": "roughly 1-2 KB (fixed-shape document)", + "count_at_full_source": "roughly 1-2 KB (same document, larger integers)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at any survey scale; the document it bounds is fixed-shape by contract.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "graph_current_survey_household_roles.py:70 MAX_ARTIFACT_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "graph_current_survey_household_roles.py:274 (_check_artifacts: `len(value.payload) <= MAX_ARTIFACT_BYTES` for every artifact_input of the node; the bind node's artifact_inputs are the ordering edge and the projection edge, declared at :239)", + "refusal_code": "\"ARTIFACT_TYPE\" (graph_current_survey_household_roles.py:277); `require` is rebound to roles.require at :56, so it raises ValueError(\"CURRENT_SURVEY_HOUSEHOLD_ROLES_ARTIFACT_TYPE\")", + "protects": "memory-or-time", + "protects_detail": "A flat byte ceiling applied to whole artifact payloads as bytes objects. There is no length field, no framing arithmetic and no reader that re-derives it (compare graph_composed_asec_binding.py:956, which does compute a framed budget). Deliberate, movable resource ceiling.", + "counted_quantity": "Bytes of the source-projection artifact: `qualified.projection` = `table.reset_index().to_json(orient=\"table\", index=False)` (current_survey_household_roles.py:637), one JSON object per STACKED original person over the 13 declared COLUMN_TOKENS plus person_id.", + "count_at_tenth": "about 215 MiB (347,138 persons at a computed 648 bytes per row \u2014 I measured the row width by json.dumps of one ACS-arm row with the real column names and the real ORIGIN/UNIVERSE literals; the 64 MiB ceiling corresponds to about 103,600 persons)", + "count_at_full_source": "about 2.1 GiB (3,471,383 persons x 648 bytes)", + "binds_at_tenth": true, + "binds_at_full_source": false, + "binding_note": "This is the binding refusal at 1/10: the row-count ceilings in the qualifier (MAX_PERSONS = 2,097,152) pass at 347,138 persons, but the projection artifact is about 215 MiB against a 64 MiB ceiling. It does NOT bind at full source only because current_survey_household_roles.py:667 (\"ORIGIN_ROSTER\", MAX_PERSONS) refuses earlier in qualify (:1127) before the projection is even built (:1151). At 1/100 (34,714 persons, about 21.5 MiB) it passes. The 648 B/row figure is computed from the declared columns, not measured on a real artifact; the ASEC arm's literals are a few bytes shorter and are only ~4% of rows (pilot: 140 of 3,464).", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "graph_current_survey_person_status.py:59 MAX_ARTIFACT_BYTES", + "value": "2**31 = 2,147,483,648", + "enforced_at": "graph_current_survey_person_status.py:387 (projection_payload: `len(payload) <= MAX_ARTIFACT_BYTES`); graph_current_survey_person_status.py:566 (_check_artifacts: `len(value.payload) <= MAX_ARTIFACT_BYTES` for every artifact_input)", + "refusal_code": "\"PROJECTION_BYTES\" (:387) and \"ARTIFACT_TYPE\" (:569); `require` is rebound to status.require at :46, which raises ValueError(\"SURVEY_PERSON_STATUS_\" + reason) (current_survey_person_status.py:126-128)", + "protects": "memory-or-time", + "protects_detail": "A flat ceiling on a whole bytes payload produced by canonical_json. I found no fixed-width length field anywhere in the framing: the payload is a bare JSON document, unlike the length-prefixed arm-rows format in graph_composed_asec_binding.py. The value 2**31 sits exactly on the signed-32-bit boundary, which is suggestive, but no consumer in the code I read stores this length in a 32-bit field, so I classify it as a deliberate 2 GiB resource ceiling rather than an encoding width.", + "counted_quantity": "Bytes of the person-status source projection: canonical_json of `_table_document(qualified.origins)` (10 columns) plus `_table_document(qualified.raw)` (45 RAW_COLUMNS: 24 ASEC + 21 ACS, current_survey_person_status.py:87-116), one row in each per STACKED original person; integers are encoded as {\"integer_literal\":\"...\"} (:365-368).", + "count_at_tenth": "about 0.14 GiB (347,138 persons at a computed 447 bytes per person: 206 B origins row + 220 B raw row + 2 index entries for an ACS-arm person, weighted with the pilot's 3,324 ACS / 140 ASEC mix)", + "count_at_full_source": "about 1.44 GiB (3,471,383 persons x 447 bytes); the ceiling corresponds to about 4.80M persons", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Does not bind on my computed width, but the margin at full source is only about 1.5x (1.44 GiB against 2 GiB), and the 447 B/row figure is an estimate from the declared columns and literal encodings, not a measurement \u2014 the real ASEC-arm literal widths could move it by tens of percent, and a 1.4 GB single bytes object must be materialized in memory regardless. Treat full-source headroom here as unverified rather than safe. I found no tighter upstream in this lane: current_survey_health_source._origins (current_survey_health_source.py:108-149), which supplies the origins table, carries no row ceiling at all, and the reader budgets at current_survey_person_status_source.py:395 and :423 are pinned source-file counts (ASEC pinned rows, ACS catalogue people), not module ceilings.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_child_property_income_source.py:34 MAX_ROWS", + "value": "600_000", + "enforced_at": "current_child_property_income_source.py:186 (_raw_axis: `0 < len(frame) <= MAX_ROWS`); current_child_property_income_source.py:203 (_catalogue_members: `len(households) <= MAX_ROWS`); current_child_property_income_source.py:217 (reused as the numeric range cap for the H_NUMPER literal); current_child_property_income_source.py:393 (project_child_property_recipients: `len(table) <= MAX_ROWS`)", + "refusal_code": "\"FULL_SOURCE_AXIS\" (:188), \"CATALOGUE_TYPE\" (:203), \"H_NUMPER_RANGE\" (via _integer_literal at :175, name + \"_RANGE\"), \"RECIPIENT_AXIS\" (:395); require at :41-43 raises ValueError(\"CHILD_PROPERTY_\" + code)", + "protects": "memory-or-time", + "protects_detail": "One constant serving four different quantities. At :186 and :203 it ceilings two upstream ASEC-source rosters that are read WHOLE regardless of sample fraction \u2014 the module states this in its own evidence: donor_scope = \"complete_original_current_year_source\" and donor_sampling_fraction_applied = False (:636-637). At :217 it is merely a generous numeric cap on a household-size literal (a domain check, not a roster). At :393 it ceilings the selected/stacked person roster. Nothing serializes it and no reader re-derives it, so as a capacity ceiling it is movable \u2014 but see the binding note about the two ASEC-source sites.", + "counted_quantity": "Binding site (:393): the stacked person roster, `origins[\"persons\"][\"rows\"]` from the preparation payload (:377-380, fed from :612-613). Non-binding sites: the full ASEC person literal frame (:186) and the full ASEC household catalogue tuple (:203).", + "count_at_tenth": "347,138 stacked persons (:393). Upstream sites are fraction-independent: 142,125 ASEC persons (:186) and 55,762 ASEC households (:203).", + "count_at_full_source": "3,471,383 stacked persons (:393). Upstream sites unchanged: 142,125 ASEC persons, 55,762 ASEC households.", + "binds_at_tenth": false, + "binds_at_full_source": true, + "binding_note": "Binds at full source through :393 (3,471,383 > 600,000); first binds around sample fraction 0.17, and it is the first refusal in the child-property lane (qualify calls project_child_property_donors at :606 then project_child_property_recipients at :613, so the ASEC-source sites pass first). The ASEC-source sizes are established, not assumed: the pinned 2025 ASEC person member pppub25.csv carries rows = 142,125 and member_size_bytes = 277,882,549 (education_assistance_source.py, the income_year 2024 AsecEducationArchive entry), and the pilot artifact records /catalogues/asec/counts/households 55762 and /catalogues/asec/counts/persons 142125. Those two sites therefore have 4.2x and 10.8x headroom and never bind \u2014 but MAX_ROWS must not be lowered below the real ASEC file sizes, because at :186/:203 it is guarding an upstream roster's true extent. The retained ASEC catalogue tuple is NOT protected by survey_population_domains.MAX_HOUSEHOLDS = 100,000: that bound is per batch (survey_population_domains.py:551) under _BATCH_HOUSEHOLDS = 10,000 (survey_catalogue_selection.py:129), so :203 is the only whole-tuple ceiling on it.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "current_child_property_income_source.py:35 MAX_PROJECTION_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "current_child_property_income_source.py:50 (_json: `len(result) <= MAX_PROJECTION_BYTES`, applied to every document this module encodes)", + "refusal_code": "\"PROJECTION_SIZE\" (current_child_property_income_source.py:50); require at :41-43 raises ValueError(\"CHILD_PROPERTY_PROJECTION_SIZE\")", + "protects": "memory-or-time", + "protects_detail": "_json is called on (a) one coordinate tuple at a time inside _table_stamp (:550), tens of bytes each, and (b) the evidence dict at :573 and :650, which is aggregate-only (:629-651: protocol strings, digests, three nested evidence dicts and four integer counts). No row-carrying document passes through it, so the ceiling does not track roster size.", + "counted_quantity": "Bytes of one encoded JSON document: a single key coordinate, or the aggregate evidence dict.", + "count_at_tenth": "a few KB (aggregate evidence document)", + "count_at_full_source": "a few KB (same document)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds at any survey scale; nothing row-shaped is routed through _json in this module.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "graph_child_property_income.py:79 MAX_ARTIFACT_BYTES", + "value": "64 * 1024**2 = 67,108,864", + "enforced_at": "graph_child_property_income.py:86 (_json: `len(payload) <= MAX_ARTIFACT_BYTES`), which every graph artifact passes through \u2014 donor projection (:357), recipient projection (:372), scenario (:510, :728), draw (:659), model metadata (:662), completion (:696), verification (:713), receipt stamps (:1144, :1158)", + "refusal_code": "\"GRAPH_ARTIFACT_SIZE\" (graph_child_property_income.py:86); `require` is rebound to child.require at :81, raising ValueError(\"CHILD_PROPERTY_GRAPH_ARTIFACT_SIZE\")", + "protects": "memory-or-time", + "protects_detail": "A flat byte ceiling on each serialized artifact. Most documents it guards are aggregate-only (donor and recipient projections carry table_sha256 digests and counts, :357-378), but two are row-carrying: the draw payload built from graph_joint_empirical._draw_document (microcosm-fit/src/microcosm/fit/graph_joint_empirical.py:470-496), which emits one row per recipient with support_id, coordinate, values, pattern and donor_key, and the completion document at :696, which re-embeds those same rows through original_draws (:704). No framing arithmetic, no reader re-derivation: movable.", + "counted_quantity": "Bytes of one serialized graph artifact. The scale-sensitive one is the draw/completion document: one row per ELIGIBLE child recipient (`recipient = recipient.loc[recipient.eligible]`, :635), at a computed 295 bytes per row.", + "count_at_tenth": "about 18 MiB, on an ESTIMATED 64,220 eligible children (I estimated eligible children as ~18.5% of persons; the exact eligible fraction is not established from any artifact in this lane)", + "count_at_full_source": "about 180 MiB, on an ESTIMATED 642,205 eligible children; the ceiling corresponds to about 227,500 recipient rows", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "It would bind at full source on my estimate, but it never gets the chance: the kernel calls child.qualify_child_property_sources at graph_child_property_income.py:923, whose current_child_property_income_source.py:393 ceiling (MAX_ROWS = 600,000 on the 3,471,383-row stacked roster) refuses first. Two further fit-side ceilings sit between this one and the data and would also have to be checked before moving anything here: graph_joint_empirical.MAX_RECIPIENTS = 1,048,576 (microcosm-fit/.../graph_joint_empirical.py:45, enforced :526 and :610, \"RECIPIENT_COUNT\") and MAX_DRAW_BYTES = 64 MiB (:44, enforced :565, \"DRAW_BYTES\"). The eligible-recipient count is an estimate, not established: no recovered artifact in this lane records eligible_child_rows.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "graph_child_property_income.py:80 MAX_ORIGINALS", + "value": "600_000", + "enforced_at": "graph_child_property_income.py:308 (_projections: `0 < len(recipients) <= MAX_ORIGINALS`, on a roster whose index is required to be identical to the origins index at :306)", + "refusal_code": "\"ORIGINAL_RECIPIENT_ROSTER\" (graph_child_property_income.py:311); `require` is child.require, raising ValueError(\"CHILD_PROPERTY_ORIGINAL_RECIPIENT_ROSTER\")", + "protects": "memory-or-time", + "protects_detail": "A duplicate of the source-side MAX_ROWS ceiling, re-asserted at the graph boundary on the same roster. Pure len() comparison, nothing serialized.", + "counted_quantity": "The stacked person roster again: `qualified.recipients`, index-identical to the origins table rebuilt from the preparation payload (:306, :224-240).", + "count_at_tenth": "347,138 stacked persons", + "count_at_full_source": "3,471,383 stacked persons", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds, because an identically valued upstream check on the identical count fires first: current_child_property_income_source.py:393 (\"RECIPIENT_AXIS\", MAX_ROWS = 600,000) runs inside child.qualify_child_property_sources, which the kernel calls at graph_child_property_income.py:923 (and again at :971), before _projections is ever reached. Raising MAX_ORIGINALS alone would change nothing.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "current_asec_income_routing_source.py:50 TOKEN_MAX_CHARS", + "value": "64", + "enforced_at": "current_asec_income_routing_source.py:476 (literal_code: `len(token) <= TOKEN_MAX_CHARS`); current_asec_income_routing_source.py:1106 (_read_capture: every non-coordinate, non-amount field of each CSV record); current_asec_income_routing_source.py:1228 (routing token contract over RECEIPT_ENTRIES, ACCOUNT_ENTRIES and OI_OFF)", + "refusal_code": "\"TOKEN_BOUND\" (:476), \"TOKEN_BOUND:\" + name (:1106), \"ROUTING_TOKEN_CONTRACT:\" + name (:1230); require at :435-437 raises ValueError(\"CURRENT_ASEC_INCOME_ROUTING_\" + reason)", + "protects": "structural-invariant", + "protects_detail": "It bounds the character length of ONE ASEC CSV field literal, alongside the printed-width regexes built from each dictionary entry's printed_length (:455-465) and COORDINATE_WIDTHS = {PH_SEQ: 5, A_LINENO: 2, A_AGE: 2} (:49). The printed entries are single-digit to eight-character codes, so 64 is a generous malformed-field guard, not a capacity ceiling. It scales with nothing.", + "counted_quantity": "Characters in one ASEC source field literal.", + "count_at_tenth": "at most the printed field width (1-8 characters for the routing and amount entries), independent of sample size", + "count_at_full_source": "same as at 1/10; independent of sample size", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Reported only because the name contains MAX; it is not a scale bound and cannot bind as the survey grows. Note also that this module's real per-run row ceiling is not a constant of its own: _read_capture enforces `len(records) < rows` (:1081, \"ROW_SHAPE\") and `len(records) == rows` (:1110, \"ROW_COUNT\") against `rows` taken from the pinned coverage member (:1303-1310), i.e. 142,125 for pppub25.csv \u2014 a class-(c) assertion about the real upstream file that must not be moved.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "asec_coverage_authentication.py:39 _BODY_MAX (imported into my family only as an enforcement argument: current_asec_income_routing_source.py:1319 passes `budget=[coverage._BODY_MAX]` to coverage._capture)", + "value": "268_435_456 (256 MiB)", + "enforced_at": "asec_coverage_authentication.py:287 (_CsvBounds.feed, decremented per record and per PRPERTYP token byte at :281-286); also asec_coverage_authentication.py:414, :424 and :519 on the assembled body; entered from current_asec_income_routing_source.py:1319", + "refusal_code": "\"COVERAGE_BODY_BYTES\"; _require at asec_coverage_authentication.py:86-88 raises AsecCoverageAuthenticationError (a ValueError subclass, :82)", + "protects": "memory-or-time", + "protects_detail": "It is a byte budget for the ENCODED coverage body, charged per CSV record while the pinned ASEC member is copied: 4 + _NUMBERS.size + 12 + 22 + 32 bytes per record plus one byte per PRPERTYP token byte (asec_coverage_authentication.py:281-286). It is not a ceiling on the 277,882,549-byte CSV itself \u2014 that file's exact size and digest are pinned separately and checked at :316-341 and at current_asec_income_routing_source.py:1303-1310.", + "counted_quantity": "Encoded coverage-body bytes, about 74 bytes per ASEC source person plus the PRPERTYP token bytes.", + "count_at_tenth": "about 10.5 MB \u2014 fraction-independent: the ASEC member is read whole (142,125 persons) whatever the survey sample fraction", + "count_at_full_source": "about 10.5 MB for 142,125 persons; the budget's capacity is roughly 3.5M persons", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: it guards an upstream ASEC file that does not scale with the US survey fraction, with roughly 25x headroom. Lowering it would weaken a real check on the pinned member; raising it is pointless. Listed here because it is the only capacity bound the routing module actually enforces on a roster.", + "family": "Current-survey and property/status graphs (us_runtime: current_survey_geography, current_survey_household_roles, graph_current_survey_household_roles, graph_current_survey_person_status, current_child_property_income_source, graph_child_property_income, current_asec_income_routing_source, graph_composed_asec_binding, graph_composed_population)" + }, + { + "constant": "graph_composed_asec_binding.py:956 ARM_ROWS_PAYLOAD_MAX_BYTES", + "value": "len(ARM_ROWS_MAGIC) + 4 + HEADER_MAX_BYTES + 16 * (MAX_PERSONS + MAX_HOUSEHOLDS) + 32 = 9 + 4 + 65,536 + 16 * (1,000,000 + 400,000) + 32 = 22,465,581", + "enforced_at": "graph_composed_asec_binding.py:1033 (encode_composed_asec_arm_rows: `len(payload) + 32 <= ARM_ROWS_PAYLOAD_MAX_BYTES`); graph_composed_asec_binding.py:1049 (bind_composed_asec_arm_rows: `len(ARM_ROWS_MAGIC) + 4 + 32 < len(payload) <= ARM_ROWS_PAYLOAD_MAX_BYTES`)", + "refusal_code": "\"ARM_ROWS_SIZE\" at both sites (:1033 and :1052); _require at graph_composed_asec_binding.py:236-238 raises PreparedGraphError (a ValueError subclass, graph_asec_prepared.py:145)", + "protects": "encoding-width", + "protects_detail": "This is the HEADER + ROW_BYTES * MAX_N form, and both the producer and the reader depend on the arithmetic. The payload is ARM_ROWS_MAGIC (9 bytes) + a uint32 little-endian header length (struct.pack(\" HEADER_MAX_BYTES` before slicing (_asec_current_money_codec.py:115-119). HEADER_MAX_BYTES is also the HEADER term of PAYLOAD_MAX_BYTES (:27) and SELECTED_PAYLOAD_MAX_BYTES (:86), so a reader's accept/refuse threshold depends on it.", + "counted_quantity": "bytes of the canonical money header JSON \u2014 a closed key set (asec_current_money.py:742-760) of digests plus two row counts; it carries no per-row content", + "count_at_tenth": "not established exactly; the key set is fixed and row-independent, so on the order of 1-2 kB", + "count_at_full_source": "same as at tenth \u2014 the header does not grow with rows", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Cannot bind: the header holds a fixed key set of sha256 strings and integers, so its size is independent of person_rows/household_rows.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "asec_current_money.py:48 MAX_PERSONS", + "value": "1_000_000", + "enforced_at": "asec_current_money.py:521 (\"SOURCE_SIZE\"), asec_current_money.py:802-805 (\"HEADER_ROW_BOUND\"); asec_current_money_selection.py:295 (\"BODY_ROW_BOUND\"), :433 (\"SELECTED_ROW_BOUND\")", + "refusal_code": "\"SOURCE_SIZE\", \"HEADER_ROW_BOUND\", \"BODY_ROW_BOUND\", \"SELECTED_ROW_BOUND\" \u2014 all MoneyRefusalError (ValueError), asec_current_money.py:59-70", + "protects": "memory-or-time", + "protects_detail": "An explicit deliberate ceiling on the person axis of the ASEC current-money body. Nothing downstream of it has a fixed width: person_rows is a plain JSON integer, and the payload length check at _asec_current_money_codec.py:122-125 recomputes from the declared rows. It is a resource ceiling, movable.", + "counted_quantity": "person rows of the three-cohort ASEC prepared current-money source (income years 2022+2023+2024) \u2014 an ASEC-side roster, NOT the selected/stacked/cloned survey. The selection sites bound a strictly-increasing subset of that same body.", + "count_at_tenth": "432523 (146133 + 144265 + 142125, the exactly pinned pppub23/24/25 row counts; does not scale with the survey sample fraction)", + "count_at_full_source": "432523 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and never scales with the survey fraction. The real upstream sizes are pinned exactly at education_assistance_source.py:117 (146133), :139 (144265), :161 (142125), and the readers force per-cohort equality (asec_income_observations.py:409 \"INCOME_COHORT_ROWS\", asec_student_controls.py:260 \"STUDENT_COHORT_ROWS\", asec_person_income_source.py:171 \"RESTORATION_COHORT_ROWS\"), so the count is fixed at 432523. 2.31x headroom.", + "family": "ASEC codec and prepared source" + }, + { + "constant": "asec_current_money.py:49 MAX_HOUSEHOLDS", + "value": "400_000", + "enforced_at": "asec_current_money.py:521 (\"SOURCE_SIZE\"), asec_current_money.py:803-805 (\"HEADER_ROW_BOUND\"); asec_current_money_selection.py:295 (\"BODY_ROW_BOUND\"), :433 (\"SELECTED_ROW_BOUND\")", + "refusal_code": "\"SOURCE_SIZE\", \"HEADER_ROW_BOUND\", \"BODY_ROW_BOUND\", \"SELECTED_ROW_BOUND\" \u2014 all MoneyRefusalError (ValueError), asec_current_money.py:59-70", + "protects": "memory-or-time", + "protects_detail": "Deliberate ceiling on the household axis of the ASEC money body; same argument as MAX_PERSONS \u2014 no fixed-width field depends on it.", + "counted_quantity": "household rows of the three-cohort ASEC prepared source. AsecMoneyScope.__post_init__ (asec_current_money.py:531-541) requires set(person_household_ids) == set(household_ids), so these are ASEC households that contain at least one person, not all hhpub rows.", + "count_at_tenth": "~169699 (derived: the 2024 leg is exactly 55762 from the pilot ASEC catalogue, scaled to 2022/2023 by their pinned person-row ratios; approximate, not measured). Fixed, does not scale with sample fraction.", + "count_at_full_source": "~169699 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and never scales. For reference, the whole three-cohort hhpub file set is only 267383 rows (88978 + 89473 + 88932, pinned at native_household_origin.py:147, :158, :169), which is still under 400000. ~2.36x headroom on the represented count.", + "family": "ASEC codec and prepared source" + }, + { + "constant": "_asec_current_money_codec.py:26 PAYLOAD_MAX_BYTES", + "value": "len(MAGIC) + 4 + HEADER_MAX_BYTES + 11 * (32 * MAX_PERSONS + MAX_HOUSEHOLDS) + 32 = 9 + 4 + 65536 + 11*(32000000 + 400000) + 32 = 356465581", + "enforced_at": "_asec_current_money_codec.py:56 (\"PAYLOAD_SIZE\"), :99 (\"PAYLOAD_SIZE\"); asec_current_money_selection.py:266 (imported as BODY_PAYLOAD_MAX_BYTES, \"PAYLOAD_SIZE\")", + "refusal_code": "\"PAYLOAD_SIZE\" raised as MoneyRefusalError (ValueError), asec_current_money.py:59-70", + "protects": "encoding-width", + "protects_detail": "A HEADER + ROW_BYTES*MAX_N byte budget the decoder relies on before it parses anything (_asec_current_money_codec.py:97-101). The 11 bytes/field-row are the amount(8)+status(1)+validity(1)+zero_origin(1) lanes written at :46-55. It is not an independent knob: it is derived from MAX_PERSONS/MAX_HOUSEHOLDS and moves automatically with them, and the exact-length check at :122-125 is what actually validates a payload.", + "counted_quantity": "bytes of the encoded ASEC current-money body = 13 + header + 11*(32*person_rows + household_rows) + 32", + "count_at_tenth": "~154.1 MB (11 * (32*432523 + 169699) = 154114785, plus header and framing); fixed, does not scale with sample fraction", + "count_at_full_source": "~154.1 MB (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; ~2.3x headroom. Moves automatically if MAX_PERSONS/MAX_HOUSEHOLDS move.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "asec_current_money_selection.py:83 SELECTED_PAYLOAD_MAX_BYTES", + "value": "len(SELECTED_MAGIC) + 4 + HEADER_MAX_BYTES + 11 * (32 * MAX_PERSONS + MAX_HOUSEHOLDS) + 32 = 356465581", + "enforced_at": "asec_current_money_selection.py:461 (\"PAYLOAD_SIZE\"), :485 (\"PAYLOAD_SIZE\")", + "refusal_code": "\"PAYLOAD_SIZE\" raised as MoneyRefusalError (ValueError), asec_current_money.py:59-70", + "protects": "encoding-width", + "protects_detail": "Same HEADER + ROW_BYTES*MAX_N form as the body budget, for the selected-subset artifact. Derived from the shared MAX_PERSONS/MAX_HOUSEHOLDS imported at :32-33.", + "counted_quantity": "bytes of the encoded selected-subset payload. The selection is a strictly-increasing, duplicate-free slice of the body (`_positions`, asec_current_money_selection.py:311-325), so selected rows can never exceed the body's rows \u2014 the combined clone cannot double it (graph_composed_asec_binding.py:313 states a non-increasing mapping cannot be a whole-household slice).", + "count_at_tenth": "at most the body size, ~154.1 MB; the ASEC arm selected into the composed population at full source is 55762 households / 142125 persons, i.e. ~52 MB", + "count_at_full_source": "~52 MB (55762 households + 142125 persons selected); upper bound still the body's ~154.1 MB", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: bounded above by the body's own row counts, which are fixed at the ASEC file sizes. ~2.3x headroom even at the theoretical maximum.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "asec_engine_evaluation.py:42 EVALUATION_HEADER_MAX_BYTES", + "value": "64 * 1024 = 65536", + "enforced_at": "asec_engine_evaluation.py:391 (\"HEADER_SIZE\"), :425 (\"HEADER_SIZE\")", + "refusal_code": "\"HEADER_SIZE\" raised as EngineAdmissionError (ValueError), asec_engine_evaluation.py:47-53", + "protects": "encoding-width", + "protects_detail": "The header length is a ` EVALUATION_HEADER_MAX_BYTES before slicing (:424-427); it is also the HEADER term the payload framing relies on.", + "counted_quantity": "bytes of the engine-evaluation header JSON: engine package/version, period, roots, closures and aggregates \u2014 a fixed roster (ADMITTED_ENGINE_OUTPUT_CONTRACTS, :69-114), row-independent", + "count_at_tenth": "not established exactly; row-independent, on the order of a few kB", + "count_at_full_source": "same as at tenth", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Cannot bind on row count: the header carries per-root metadata for a closed root roster, not per-row data.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "asec_engine_evaluation.py:43 EVALUATION_MAX_ROWS", + "value": "1_000_000", + "enforced_at": "asec_engine_evaluation.py:345 (\"ENGINE_ROW_BOUND\"), :437 (\"ENGINE_ROW_BOUND\")", + "refusal_code": "\"ENGINE_ROW_BOUND\" raised as EngineAdmissionError (ValueError), asec_engine_evaluation.py:47-53", + "protects": "memory-or-time", + "protects_detail": "An explicit ceiling on how many person rows are handed to the PolicyEngine-US engine and then written as 8*rows*len(roots) float64 bytes (:398-404, :439). Deliberate resource ceiling; the payload length is recomputed from the declared rows, so no fixed-width field depends on it.", + "counted_quantity": "frame.n(\"person\") of the ASEC-prepared engine frame (graph_asec_prepared.py:824 `_engine_frame`, :836), i.e. the ASEC-prepared person roster after the seeded whole-household selection kernel (graph_asec_prepared.py:339, fraction/seed params at :1016)", + "count_at_tenth": "~43252 if the ASEC-prepared selection kernel is run at fraction 0.1; 432523 at fraction 1.0", + "count_at_full_source": "432523 (the full three-cohort ASEC prepared roster)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: this is the ASEC-prepared slice, not the stacked ACS+ASEC survey. 2.31x headroom at fraction 1.0. This bound would bind if it were ever pointed at the stacked (~3.47M) or cloned (~6.94M) survey roster \u2014 it is not.", + "family": "ASEC codec and prepared source" + }, + { + "constant": "asec_housing_status.py:52 HEADER_MAX_BYTES", + "value": "65536", + "enforced_at": "asec_housing_status.py:102 (\"HEADER_SIZE\"), :445 (\"HEADER_SIZE\")", + "refusal_code": "\"HEADER_SIZE\" raised as HousingStatusRefusalError (ValueError), asec_housing_status.py:86-92", + "protects": "encoding-width", + "protects_detail": "Header length is a ` HEADER_MAX_BYTES before slicing; it is also the HEADER term of PAYLOAD_MAX_BYTES at :55.", + "counted_quantity": "bytes of the housing-status header JSON: codebook, dictionaries, zero-origin evidence and one household_rows integer \u2014 row-independent", + "count_at_tenth": "not established exactly; row-independent, on the order of a few kB (the embedded codebook and DICTIONARIES at :67 dominate)", + "count_at_full_source": "same as at tenth", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Cannot bind on row count.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "asec_housing_status.py:53 MAX_HOUSEHOLDS", + "value": "400_000", + "enforced_at": "asec_housing_status.py:174 (\"ROW_COUNT\"), :294 (\"ROW_COUNT\")", + "refusal_code": "\"ROW_COUNT\" raised as HousingStatusRefusalError (ValueError), asec_housing_status.py:86-92", + "protects": "memory-or-time", + "protects_detail": "Deliberate ceiling on the household axis of the ASEC housing-status evidence buffers (11 int64 raw columns + 14 uint8 derived, :19-50). Nothing fixed-width depends on it; the buffer sizes are checked against the declared n at :299-302.", + "counted_quantity": "ASEC household rows carried by the authenticated current-money source's household table (asec_housing_status_source.load_authenticated_housing_status reads source.frame.table(\"household\")) \u2014 an ASEC-side roster across income years 2022/2023/2024", + "count_at_tenth": "~169699 (derived; see asec_current_money.py:49). Does not scale with the survey sample fraction.", + "count_at_full_source": "~169699 (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and never scales. ~2.36x headroom.", + "family": "ASEC codec and prepared source" + }, + { + "constant": "asec_housing_status.py:54 ROW_BYTES", + "value": "8 * len(RAW_COLUMNS) + len(DERIVED_COLUMNS) = 8*11 + 14 = 102", + "enforced_at": "not directly enforced; it is the ROW_BYTES factor of PAYLOAD_MAX_BYTES at asec_housing_status.py:55. Per-column buffer widths are checked at asec_housing_status.py:299-302.", + "refusal_code": "n/a \u2014 not itself a refusal site; the derived budget it feeds refuses with \"PAYLOAD_SIZE\" (HousingStatusRefusalError, asec_housing_status.py:86-92)", + "protects": "encoding-width", + "protects_detail": "A per-row wire width, not a ceiling. It is the literal serialized width of one household row (11 int64 + 14 uint8). Changing it changes the wire format; it is reported here only because the task calls out the HEADER + ROW_BYTES*N construction and nothing should be silently dropped.", + "counted_quantity": "bytes per household row on the wire (a width, not a count)", + "count_at_tenth": "102 (constant)", + "count_at_full_source": "102 (constant)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Not a scale bound \u2014 it is a width factor. Included so it is not silently dropped.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "encoding-width", + "safe_to_move": false, + "agrees_with_census": true + } + }, + { + "constant": "asec_housing_status.py:55 PAYLOAD_MAX_BYTES", + "value": "len(MAGIC) + 4 + HEADER_MAX_BYTES + ROW_BYTES * MAX_HOUSEHOLDS + 32 = 9 + 4 + 65536 + 102*400000 + 32 = 40865581", + "enforced_at": "asec_housing_status.py:438 (\"PAYLOAD_SIZE\"), and as the bounded disk read at :473 (`handle.read(PAYLOAD_MAX_BYTES + 1)`)", + "refusal_code": "\"PAYLOAD_SIZE\" raised as HousingStatusRefusalError (ValueError), asec_housing_status.py:86-92", + "protects": "encoding-width", + "protects_detail": "Textbook HEADER + ROW_BYTES*MAX_N budget that the reader relies on before decoding and that caps the on-disk read at :473. Derived: moves automatically with MAX_HOUSEHOLDS.", + "counted_quantity": "bytes of the encoded housing-status artifact = 13 + header + 102*household_rows + 32", + "count_at_tenth": "~17.4 MB (102 * 169699 = 17309298 plus header/framing); fixed, does not scale with sample fraction", + "count_at_full_source": "~17.4 MB (same)", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds; ~2.36x headroom.", + "family": "ASEC codec and prepared source", + "verified": { + "binds_at_full_source": false, + "protects": "memory-or-time", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "asec_housing_universe.py:40 HEADER_MAX_BYTES", + "value": "65536", + "enforced_at": "asec_housing_universe.py:72 (\"HEADER_SIZE\"), reached from every `_parse` caller (e.g. the header_data property and :382-390 decode path)", + "refusal_code": "\"HEADER_SIZE\" raised as HousingUniverseRefusalError (ValueError), asec_housing_universe.py:52-58", + "protects": "encoding-width", + "protects_detail": "Header length is a ` MAX_ROWS: raise`); acs_person_coverage_authentication.py:373 (`_require(rows <= literal.MAX_ROWS, \"SOURCE_ROWS\")`). Recorded but not enforced at acs_person_coverage_authentication.py:173.", + "refusal_code": "acs_person_coverage_columns.py:199 raises bare `ValueError(\"ACS coverage source row bound exceeded\")` \u2014 this module has no _require/typed error. acs_person_coverage_authentication.py:373 raises `ACSCoverageAuthenticationError(\"SOURCE_ROWS\")` (ACSCoverageAuthenticationError subclasses ValueError, declared at acs_person_coverage_authentication.py:56).", + "protects": "upstream-file-size", + "protects_detail": "The counted rows are read straight off the upstream Census archive members and nothing per-row is retained on this path (in auth._inventory the selected dict stays empty when serialnos is frozenset()). It is a structural assertion that a genuine psam_*.csv roster is near its real size, which is exactly how the lane's own commit 14defbfc0 justified leaving it alone: 'MAX_ROWS stays at 6,000,000 because it asserts the source file's own size rather than a roster this build chooses.' MUST NOT be moved down.", + "counted_quantity": "Data rows (header excluded) accumulated across every psam_hus*/psam_pus* CSV member of ONE archive, per _inventory call / per _scan_acs_person_coverage call. Upstream archive rows, not the selected, stacked or cloned roster \u2014 so the count is identical at every selection fraction.", + "count_at_tenth": "person role 3,422,888; household role 1,631,969 (unchanged by fraction)", + "count_at_full_source": "person role 3,422,888; household role 1,631,969 \u2014 MEASURED this session by streaming the pilot's captured csv_pus.zip / csv_hus.zip at /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/snapshots/acs-hu-capture-e5pg8w1z/ (psam_pusa 1,743,751 + psam_pusb 1,679,137; psam_husa 827,133 + psam_husb 804,836). Both totals match the pilot receipt's catalogue counts exactly.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 3,422,888 of 6,000,000 is 57% used, 1.75x headroom. It is also the tightest gate of its kind, so it makes three other 6,000,000 constants dead code \u2014 acs_native_coverage_binding.py:315/:330/:344 (\"SOURCE_ROW_BUDGET\") and acs_population_catalogue.py:156 (\"SOURCE_ROW_LIMIT\") both test `sum(m[\"rows\"] for m in inventory) <= 6_000_000` on inventories that coverage._inventory already refused above 6,000,000 at acs_person_coverage_authentication.py:373.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": false, + "protects": "upstream-file-size", + "safe_to_move": false, + "agrees_with_census": true + } + }, + { + "constant": "acs_person_coverage_columns.py:40 MAX_SELECTED_ROWS", + "value": "14_000_000 at HEAD 14defbfc0. WAS 1_000_000 at 1d5cfafb1, the HEAD this census was scoped to; commit 14defbfc0 'Lift the four binding row-count ceilings under one rule' raised it mid-census.", + "enforced_at": "acs_person_coverage_columns.py:239-240; acs_person_coverage_authentication.py:396 (strict `<`, so the effective cap there is value-1); acs_person_coverage_authentication.py:419; acs_native_coverage_binding.py:317-322 (code literal at :321) \u2014 whole-source route only; acs_native_coverage_binding.py:336; acs_population_catalogue.py:312-314 (inside `min(_BATCH_PEOPLE, literal.MAX_SELECTED_ROWS)`).", + "refusal_code": "columns.py:240 bare `ValueError(\"ACS coverage selected person count is outside the bound\")`; auth.py:396 `ACSCoverageAuthenticationError(\"SELECTED_ROWS\")`; auth.py:419 `ACSCoverageAuthenticationError(\"NATIVE_ROWS\")`; binding.py:321 and :336 `ACSNativeCoverageBindingError(\"NATIVE_ROW_BUDGET\")` (declared binding.py:53, subclasses ValueError); catalogue.py:314 `ACSSourceCatalogueError(\"BATCH_PERSON_LIMIT\")` (declared catalogue.py:35).", + "protects": "memory-or-time", + "protects_detail": "Every live site bounds a roster this build chooses to materialize, not a source file. columns.py:239 gates `len(expected)` before the reader accumulates `pieces` DataFrames in RAM (columns.py:242-259); auth:396 gates the `selected` dict that _inventory builds; auth:419 gates the live Frame's person table; binding:336 gates `expected_rows = sum(households.values())` before construction. Deliberate, arbitrary, movable \u2014 which is what the lift commit acted on.", + "counted_quantity": "Selected/native ACS person rows \u2014 the requested roster, not the stacked or cloned roster. EXCEPTION: at binding.py:317-322 the same constant is applied to the whole upstream person archive's row total, but only on the `serialnos is None` route.", + "count_at_tenth": "342,289 selected ACS persons (3,422,888/10). Combined ACS+ASEC at 1/10 is 347,138, but these modules only ever see ACS keys \u2014 survey_population_preparation.py:1946-1948 filters `plan.selected` to `Source.ACS` before passing `serialnos=acs_keys` at :1959.", + "count_at_full_source": "3,422,888 selected ACS persons (every one of the 1,531,614 selectable non-vacant ACS households). Measured, not scaled.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "At HEAD it no longer binds anywhere: 3,422,888 of 14,000,000 is 24% used. At the pre-lift 1,000,000 it bound hard at full source (3.42x over) at all five live sites and did not bind at 1/10. It never binds at catalogue.py:314 at either value \u2014 `min(_BATCH_PEOPLE=100_000, ...)` makes _BATCH_PEOPLE the operative term, and the batching loop at catalogue.py:329 caps `batch_people` at 100,000 before the check runs. It is now moot regardless: MAX_BODY_BYTES at acs_person_coverage_authentication.py:381 refuses this same lane at ~12,805 selected person rows, 267x earlier, and was not lifted.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_person_coverage_columns.py:41 MAX_CSV_RECORD_CHARS", + "value": "100_000", + "enforced_at": "acs_person_coverage_columns.py:138 (bounds the readline request) and :140-141 (the refusal). Recorded at acs_person_coverage_authentication.py:172.", + "refusal_code": "bare `ValueError(\"ACS coverage CSV record character bound exceeded\")` \u2014 acs_person_coverage_columns.py:141", + "protects": "memory-or-time", + "protects_detail": "Per-record allocation ceiling: `record_chars` resets to 0 at every record start (columns.py:150), so it bounds one logical CSV record's decoded characters including its physical line ending, never a cumulative total. Movable, but there is no reason to.", + "counted_quantity": "Decoded characters in ONE logical CSV record of an ACS PUMS member. Not a row, household, person or roster count \u2014 it does not scale with the survey at all.", + "count_at_tenth": "1,882 (worst case, unchanged by fraction)", + "count_at_full_source": "1,882 \u2014 MEASURED this session by scanning every logical record of both archives: the longest line in csv_pus.zip is 1,882 bytes and in csv_hus.zip 1,570 bytes, and in both cases that longest line IS the header, so every data record is strictly shorter. Neither archive contains a single `\"` character, so the quote-parity paths never engage.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 1,882 of 100,000 is 1.9% used, 53x headroom.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_pums.py:55 MAX_EXACT_HOUSEHOLDS", + "value": "7_000_000 at HEAD 14defbfc0. WAS 1_000_000 at 1d5cfafb1; raised mid-census by commit 14defbfc0.", + "enforced_at": "acs_pums.py:194 (condition inside AcsPumsSource.snapshot_serialnos), raising at acs_pums.py:206-208. Reached from acs_native_coverage_binding.py:561 and acs_housing_universe_source.py:829.", + "refusal_code": "bare `ValueError(\"ACS exact selection requires bounded unique raw native keys.\")` \u2014 acs_pums.py:207. Inside issue_acs_native_coverage this is not an ACSNativeCoverageBindingError, so binding.py:735-736 converts it to `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")`; inside prepare_acs_housing_population, acs_housing_universe_source.py:918-928 converts it to `ACSHousingSourceError(\"FRAME_PREPARATION_REFUSED\")`.", + "protects": "memory-or-time", + "protects_detail": "Bounds the caller's requested key tuple before it is frozen and used to build frozensets and `.isin` masks over the whole archive. Not a source-file claim \u2014 the lift commit says so explicitly: 'Neither is a statement about the source file, which MAX_ROWS in acs_person_coverage_columns still makes at 6,000,000.' Movable.", + "counted_quantity": "`len(serialnos)` \u2014 the number of exact requested ACS household SERIALNOs, i.e. the selected household roster. Only on the serialnos path; the `serialnos is None` route returns at acs_pums.py:190 before the check.", + "count_at_tenth": "153,161 selected ACS households (1,531,614/10)", + "count_at_full_source": "1,531,614 selected ACS households \u2014 the non-vacant ACS catalogue records that survey_catalogue_selection can draw from (catalogue counts 1,631,969 total minus 100,355 vacancies; 1,348,408 occupied HU + 84,422 institutional GQ + 98,784 noninstitutional GQ = 1,531,614). Measured from the pilot's preparation.json catalogue block.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "At HEAD it no longer binds: 1,531,614 of 7,000,000 is 22% used. At the pre-lift 1,000,000 it bound at full source (1.53x over) and, because snapshot_serialnos is called at binding.py:561 before anything else, it would have been the FIRST refusal in program order for a full-source selection. It is moot either way: two tighter gates fire earlier as the fraction rises \u2014 MAX_BODY_BYTES at ~5,730 households and the 1 MiB canonical-JSON cap on the serialnos list at binding.py:564 at 65,535 households.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_pums.py:56 MAX_EXACT_PERSON_ROWS", + "value": "14_000_000 at HEAD 14defbfc0. WAS 1_000_000 at 1d5cfafb1; raised mid-census by commit 14defbfc0.", + "enforced_at": "acs_pums.py:270-273 inside load_acs_pums_tables, on the serialnos branch only.", + "refusal_code": "bare `ValueError(\"ACS selected complete roster exceeds native person budget.\")` \u2014 acs_pums.py:272. Converted downstream to `ACSHousingSourceError(\"FRAME_PREPARATION_REFUSED\")` (acs_housing_universe_source.py:918-928) and then to `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")` (acs_native_coverage_binding.py:735-736).", + "protects": "memory-or-time", + "protects_detail": "Gates `household.NP.sum()` after the exact selection is applied and immediately before _read_archive materializes the person table \u2014 the module docstring (acs_pums.py:17-21) states the selected tables 'necessarily materialize: the returned Frame itself is the dense base-pool artifact'. A real RAM ceiling on what is about to be built. Movable.", + "counted_quantity": "Sum of NP over the SELECTED occupied households = the exact selected person-row count. The selected roster, not the source file, not the stacked or cloned roster.", + "count_at_tenth": "342,289 selected ACS person rows", + "count_at_full_source": "3,422,888 selected ACS person rows", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "At HEAD it no longer binds: 24% used. At the pre-lift 1,000,000 it bound at full source (3.42x over). It could never have been the refusal a user saw, though: acs_native_coverage_binding.py:336 tests the identical quantity (`expected_rows = sum(households.values())`) against the identical pre-lift value inside _preflight (binding.py:573), which runs before prepare_acs_housing_population (binding.py:578) \u2014 so \"NATIVE_ROW_BUDGET\" always fired first. Both are now outranked by MAX_BODY_BYTES.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_person_coverage_authentication.py:36 MAX_HEADER_BYTES", + "value": "1024**2 = 1_048_576", + "enforced_at": "Default `cap` of _json at acs_person_coverage_authentication.py:70; the refusals live inside _json at :79 and :120. Call sites that pass it explicitly: :585 (the UngrantedACSNativeBinding receipt) and :695 (`raw_header = _json(header, MAX_HEADER_BYTES)`, the coverage envelope header).", + "refusal_code": "`ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` \u2014 acs_person_coverage_authentication.py:79 (pre-charge pass) and :120 (encoder pass)", + "protects": "memory-or-time", + "protects_detail": "A canonical-encoding size ceiling on the coverage envelope's JSON header, charged byte-by-byte before the encoder can allocate a whole token (see the comment at :72-74). Movable.", + "counted_quantity": "ASCII-escaped canonical JSON bytes of the coverage receipt header. Its contents are O(ZIP members) and O(columns), never O(rows): a per-member inventory (name/bytes/sha256/compressed_bytes/crc32/applicable/header_sha256/rows) for 3+3 members, the field contract, the archives list, and the original_literal_receipt whose row-dependent parts are integers and sha256 digests. It does NOT embed the roster.", + "count_at_tenth": "order 10 kB (unchanged by fraction; not measured directly)", + "count_at_full_source": "order 10 kB. Not measured directly \u2014 I did not import the module. Computed from the encoder and the schema against the measured archive structure: 6 member entries of roughly 250 bytes each plus a ~1.5 kB field contract. Row counts enter only as integers.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and cannot scale: no per-row datum is charged against it. The one sibling cap that IS row-scaling is ACS_HU_RECEIPT_MAX_BYTES, which has the same 1 MiB value and is applied to the serialnos list at acs_native_coverage_binding.py:564.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_person_coverage_authentication.py:37 MAX_BODY_BYTES", + "value": "64 * 1024**2 = 67_108_864", + "enforced_at": "acs_person_coverage_authentication.py:379-382 (\"SELECTED_BODY_BUDGET\"); acs_person_coverage_authentication.py:661 (\"BODY_SIZE\"); acs_native_coverage_binding.py:640-647 as the _json cap on the sorted raw person-key list (refuses inside _json at auth:79/:120).", + "refusal_code": "`ACSCoverageAuthenticationError(\"SELECTED_BODY_BUDGET\")` at auth:381; `ACSCoverageAuthenticationError(\"BODY_SIZE\")` at auth:661; `ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` at auth:79/:120 for the binding:646 site. All three surface out of issue_acs_native_coverage as `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")` (binding.py:735-736), because ACSCoverageAuthenticationError is not a subclass of ACSNativeCoverageBindingError \u2014 both subclass ValueError independently.", + "protects": "memory-or-time", + "protects_detail": "auth:380's comment states it exactly: 'Upper bound on escaped row + statuses + lineage before the low-level reader allocates its selected DataFrame.' It is a deliberate, arbitrary RAM budget for the selected NDJSON payload, charged at `selected_budget += 6 * len(raw) + 1024` per selected row. Nothing downstream reads a fixed-width field or a HEADER + ROW_BYTES * N offset, so it is not a wire format. Movable.", + "counted_quantity": "auth:381 \u2014 a pessimistic byte estimate of the selected roster, 6x the raw CSV record bytes plus 1024, summed over the SELECTED rows of one role (a fresh `selected_budget` per _inventory call). auth:661 \u2014 the actual NDJSON body, one canonical-JSON line per selected person row. Both are the selected survey roster, not the source file and not the stacked or cloned roster.", + "count_at_tenth": "auth:381 person role 153,161 hh -> 342,289 selected persons x (6 x 702.76 + 1024 = 5,240.6 B) = 1.794e9 B, 26.7x over. Household role 153,161 x (6 x 598.96 + 1024 = 4,617.8 B) = 7.073e8 B, 10.5x over. auth:661: 342,289 x ~92 B = 3.15e7 B, passes.", + "count_at_full_source": "auth:381 person role 3,422,888 x 5,240.6 = 1.794e10 B, 267x over; household role 1,531,614 x 4,617.8 = 7.073e9 B, 105x over. auth:661: 3,422,888 x ~92 B = 3.15e8 B, 4.7x over. The 702.76 and 598.96 B/record averages are MEASURED this session (csv_pus 2,405,468,113 B over 3,422,888 rows; csv_hus 977,484,520 B over 1,631,969 rows). The ~92 B/line NDJSON figure is computed from the encoder and _COLUMNS, not measured.", + "binds_at_tenth": true, + "binds_at_full_source": true, + "binding_note": "THIS IS THE FIRST CEILING THIS FAMILY HITS, AND THE LIFT COMMIT DID NOT TOUCH IT. auth:381 exhausts at 67,108,864/5,240.6 = 12,805 selected person rows on the person role, i.e. 12,805/2.23485 = 5,730 selected ACS households, i.e. selection fraction 1/267 \u2014 only 3.7x above the 1/1000 pilot, which used 17,419,588 B (26% of budget) on the person role and 7,060,538 B (10.5%) on the household role. The household role exhausts at 14,533 households (1/105) and is scanned first inside _preflight, so above 1/105 it is the household _inventory that refuses; between 1/267 and 1/105 the household inventory passes and the person inventory refuses. Either way the code is \"SELECTED_BODY_BUDGET\". It is 78x tighter than the pre-lift 1,000,000 row ceilings and 1,278x tighter than the post-lift 14,000,000 ones, so every row-count ceiling in this family is now unreachable behind it. auth:661 is a distant second at ~721,600 persons (1/4.7). This bound is memory-or-time and movable, but note the 6x multiplier is an estimate of escaped JSON against raw CSV, so raising it should be argued against the measured ~92 B/line the body actually costs, not against the 5,240 B/row the estimator charges.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": true + } + }, + { + "constant": "acs_person_coverage_authentication.py:38 MAX_RECORD_BYTES", + "value": "400_000", + "enforced_at": "acs_person_coverage_authentication.py:198 (selects `cap` for non-first records) with the refusals at :206 and :210; also the _json cap at :499 and :657 (both `MAX_RECORD_BYTES - 1`), refusing inside _json at :79/:120.", + "refusal_code": "`ACSCoverageAuthenticationError(\"CSV_RECORD_BYTES\")` at auth:206 and :210; `ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` at auth:79/:120 for the :499 and :657 sites.", + "protects": "memory-or-time", + "protects_detail": "The byte fence's per-record ceiling, charged before UTF-8 decoding or any CSV allocation (docstring at :222-238). Resets at every record. Also caps one roster line and one NDJSON line. Movable.", + "counted_quantity": "Bytes in ONE raw logical CSV record (including its terminator), and separately the canonical JSON bytes of ONE roster/NDJSON row. Per-record; does not scale with the survey.", + "count_at_tenth": "1,882 worst case (unchanged by fraction)", + "count_at_full_source": "1,882 \u2014 MEASURED; the longest logical record in either archive is the 1,882-byte psam_pus header, so every data record is shorter. The JSON sites carry ~92 B/row (computed).", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 0.5% used, 212x headroom. It also underwrites the _records scanner's structural invariant that a record spans at most two blocks, since MAX_RECORD_BYTES (400,000) < _SCAN_BLOCK (1,048,576) \u2014 see the docstring at auth:232-238. Lowering it below _SCAN_BLOCK is safe; raising it above 1,048,576 would break that invariant.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_person_coverage_authentication.py:39 MAX_CSV_HEADER_BYTES", + "value": "64 * 1024 = 65_536", + "enforced_at": "acs_person_coverage_authentication.py:198 (`cap = MAX_CSV_HEADER_BYTES if first else MAX_RECORD_BYTES`) with the refusals at :206 and :210; fast path at :199.", + "refusal_code": "`ACSCoverageAuthenticationError(\"CSV_RECORD_BYTES\")` \u2014 acs_person_coverage_authentication.py:206 and :210", + "protects": "memory-or-time", + "protects_detail": "A looser per-record ceiling for the FIRST record of a member, because a CSV header is wider than a data row. Per-record allocation guard. Movable.", + "counted_quantity": "Bytes in the CSV header line of one archive member. Per-file; does not scale with rows.", + "count_at_tenth": "1,882 (unchanged by fraction)", + "count_at_full_source": "1,882 for psam_pus* (286 columns), 1,570 for psam_hus* (241 columns) \u2014 both MEASURED this session by reading the first line of each member.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 2.9% used, 34.8x headroom. It would only bind if the Census widened the PUMS dictionary by 35x.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_person_coverage_authentication.py:40 MAX_TOKEN_BYTES", + "value": "64 * 1024 = 65_536", + "enforced_at": "acs_person_coverage_authentication.py:199 (fast-path `min(cap, MAX_TOKEN_BYTES)`) and :215 (the refusal).", + "refusal_code": "`ACSCoverageAuthenticationError(\"CSV_TOKEN_BYTES\")` \u2014 acs_person_coverage_authentication.py:215", + "protects": "memory-or-time", + "protects_detail": "Per-field ceiling, reset at each unquoted comma (auth:212-213); the docstring at :224-226 notes it is conservative because it includes raw CSV quoting. Movable.", + "counted_quantity": "Bytes in ONE CSV field. Per-field; does not scale.", + "count_at_tenth": "13 (unchanged by fraction)", + "count_at_full_source": "13 \u2014 MEASURED; the longest comma-delimited field anywhere in either archive is 13 bytes (that is the SERIALNO). Both archives contain zero quote characters, so the quoted-field path never engages.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 0.02% used, 5,041x headroom.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_native_coverage_binding.py:28 MAX_ARCHIVE_BYTES", + "value": "8 * 1024**3 = 8_589_934_592 (the source comment reads 'combined compressed bytes, before capture')", + "enforced_at": "acs_native_coverage_binding.py:568", + "refusal_code": "`ACSNativeCoverageBindingError(\"ARCHIVE_BUDGET\")` \u2014 acs_native_coverage_binding.py:568 (the class is declared at binding.py:52-53 and subclasses ValueError)", + "protects": "memory-or-time", + "protects_detail": "It gates the disk capture: housing._capture copies both archives into a private snapshot and separately demands free space of sum(pins) + 2*ACS_HU_SOURCE_MAX_BYTES + 2*ACS_HU_RECEIPT_MAX_BYTES + 1 GiB at acs_housing_universe_source.py:200-207 (\"INSUFFICIENT_DISK\"). The sizes it sums are already exact per-file pins from the manifest, so it is an operational ceiling on capture cost, not the size assertion \u2014 that assertion is the manifest's own exact `size_bytes` equality. Movable.", + "counted_quantity": "Sum of the two pinned archives' declared size_bytes. Upstream compressed bytes; identical at every selection fraction.", + "count_at_tenth": "854,347,733 bytes (unchanged by fraction)", + "count_at_full_source": "854,347,733 bytes = 251,500,587 (csv_hus.zip) + 602,847,146 (csv_pus.zip), read from acs_2024_1yr_sources.json and confirmed against the pilot's captured files on disk.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 9.9% used, 10.05x headroom. acs_population_catalogue.py:481 (\"ARCHIVE_LIMIT\") is the same value on the same quantity.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_native_coverage_binding.py:29 MAX_EXPANDED_BYTES", + "value": "16 * 1024**3 = 17_179_869_184 (comment: 'combined, before opening any member')", + "enforced_at": "acs_native_coverage_binding.py:303-307", + "refusal_code": "`ACSNativeCoverageBindingError(\"EXPANDED_BUDGET\")` \u2014 acs_native_coverage_binding.py:307", + "protects": "memory-or-time", + "protects_detail": "A decompression-bomb gate: `expanded` accumulates member.file_size across BOTH archives from the central directory and is checked before any member is opened. Movable.", + "counted_quantity": "Total uncompressed bytes declared by every ZIP member of csv_hus.zip plus csv_pus.zip. Upstream bytes; identical at every fraction.", + "count_at_tenth": "3,383,159,949 bytes (3.151 GiB, unchanged by fraction)", + "count_at_full_source": "3,383,159,949 bytes = 977,588,178 (csv_hus: psam_husa 495,752,055 + psam_husb 481,732,465 + README.pdf 103,658) + 2,405,571,771 (csv_pus: psam_pusa 1,226,543,489 + psam_pusb 1,178,924,624 + README.pdf 103,658). MEASURED this session from the pilot's captured archives.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 19.7% used, 5.08x headroom. The per-archive twin _EXPANDED_MAX (acs_person_coverage_authentication.py:325, acs_housing_universe_source.py:295) has the same value against a 2.24 GiB worst case, and acs_population_catalogue.py:149 (\"EXPANDED_LIMIT\") is the same cross-archive form.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_native_coverage_binding.py:30 MAX_SOURCE_ROWS", + "value": "6_000_000 (comment: 'per role, before full-source construction')", + "enforced_at": "acs_native_coverage_binding.py:313-315, :328-330, :342-344 \u2014 three sites, all on `sum(m[\"rows\"] for m in inventories[role])`.", + "refusal_code": "`ACSNativeCoverageBindingError(\"SOURCE_ROW_BUDGET\")` \u2014 acs_native_coverage_binding.py:315, :330, :344", + "protects": "upstream-file-size", + "protects_detail": "The counted rows come back from coverage._inventory, which streams the upstream archive and retains nothing when the selection set is empty. It is the same quantity and the same number as acs_person_coverage_columns.MAX_ROWS, which the lift commit explicitly classified as a source-file assertion. Its comment frames it as a pre-construction guard, but the thing it counts is the Census file's own rows. Treat as not movable downward.", + "counted_quantity": "Upstream archive data rows per role: household 1,631,969, person 3,422,888. Not the selected, stacked or cloned roster; identical at every fraction.", + "count_at_tenth": "household 1,631,969; person 3,422,888 (unchanged by fraction)", + "count_at_full_source": "household 1,631,969; person 3,422,888 \u2014 MEASURED, and matching the pilot's catalogue counts exactly.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds, and in fact CANNOT fire: coverage._inventory already refuses at acs_person_coverage_authentication.py:373 with `rows <= literal.MAX_ROWS` at the identical value 6,000,000 while scanning, so every inventory it is handed already satisfies `sum <= 6,000,000`. All three sites are dead code. Same for acs_population_catalogue.py:156.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": false, + "protects": "upstream-file-size", + "safe_to_move": false, + "agrees_with_census": true + } + }, + { + "constant": "acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES", + "value": "2 * 1024**2 = 2_097_152", + "enforced_at": "acs_native_coverage_binding.py:465 (_frame_sha256 details); :564 (`min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES)` on the serialnos tuple); :711 (`payload = coverage._json(receipt, MAX_EVIDENCE_BYTES)`). All three refuse inside coverage._json at acs_person_coverage_authentication.py:79/:120.", + "refusal_code": "`ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")`, surfaced by issue_acs_native_coverage as `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")` (binding.py:735-736) and by verify_acs_native_coverage as `ACSNativeCoverageBindingError(\"NATIVE_VERIFICATION_REFUSED\")` (binding.py:542-543).", + "protects": "memory-or-time", + "protects_detail": "A canonical-encoding ceiling on the issuance receipt. The :465 site is O(entities x columns) and cannot scale. The :711 and :564 sites ARE row-scaling, because the receipt embeds `\"requested_serialnos\": serialnos` (binding.py:684) and `\"vacant_serialnos\"` (binding.py:697-699) \u2014 the literal selected household key list, at 16 canonical bytes per 13-character key. Movable.", + "counted_quantity": "Canonical ASCII-JSON bytes of the native-coverage issuance receipt, dominated by the selected ACS household key list. The selected roster, not the source file.", + "count_at_tenth": "153,161 keys -> 16 x 153,161 + 1 = 2,450,577 bytes for the key list alone, before the ~20-40 kB fixed producer closure: 1.17x over the 2 MiB cap at :711, and 2.34x over the tighter 1 MiB cap at :564.", + "count_at_full_source": "On the production path (serialnos given): 1,531,614 keys -> 24,505,825 bytes, 11.7x over. On the `serialnos is None` route the field is `null` and the receipt stays at order 40 kB, so it does not bind there \u2014 but survey_population_preparation.py:1959 never takes that route.", + "binds_at_tenth": true, + "binds_at_full_source": true, + "binding_note": "Binds at 1/10 and above, at both the :564 and :711 sites; the :564 site is tighter because of the min() with the 1 MiB ACS_HU_RECEIPT_MAX_BYTES, exhausting at n = 65,535 selected households (fraction 1/23.4). Does not bind at 1/100 (15,316 keys -> 245,057 bytes). It is not the first thing to fire: MAX_BODY_BYTES refuses at 5,730 households, 11x earlier. Note the asymmetry the lift did not address \u2014 these two caps make a full-source run IMPOSSIBLE on the selection path regardless of any row-count ceiling, because 1.5M keys cannot fit a 1 MiB JSON list.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": true + } + }, + { + "constant": "acs_native_coverage_binding.py:45 _COMPILE_CACHE_MAX_ENTRIES", + "value": "128", + "enforced_at": "acs_native_coverage_binding.py:96 (eligibility: `and _COMPILE_CACHE_MAX_ENTRIES > 0`) and :141 (FIFO eviction condition).", + "refusal_code": "NONE. There is no refusal and no exception. At :96 an ineligible input simply bypasses the cache and compiles normally; at :139-145 the oldest entry is deleted silently.", + "protects": "memory-or-time", + "protects_detail": "The comment at binding.py:141-142 says it plainly: 'FIFO eviction bounds retained source keys; it is not a hard RSS cap.' It bounds a bytecode cache for the producer drift check.", + "counted_quantity": "Number of (source, filename, mode, flags, dont_inherit, optimize, compiler-id) entries retained in _COMPILE_CACHE. Python module compilations, not data rows \u2014 it is completely independent of survey scale.", + "count_at_tenth": "order 30-60 module compilations (unchanged by fraction)", + "count_at_full_source": "order 30-60 \u2014 the modules named by implementation_manifest plus the three added at binding.py:250-253. Not measured exactly; I did not import the module.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds on data scale and cannot: it counts compiled Python modules, not rows. Included only because the name contains MAX, per the census rule.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_native_coverage_binding.py:46 _COMPILE_CACHE_MAX_SOURCE_BYTES", + "value": "16 * 1024**2 = 16_777_216", + "enforced_at": "acs_native_coverage_binding.py:95 (eligibility) and :142-144 (eviction while the retained source total plus this source would exceed it).", + "refusal_code": "NONE \u2014 ineligible sources bypass the cache; over-budget entries are evicted silently.", + "protects": "memory-or-time", + "protects_detail": "Bounds the total bytes of module source text retained as cache keys. Not data.", + "counted_quantity": "Sum of len(source) over retained cache keys \u2014 Python source bytes, not survey rows.", + "count_at_tenth": "order 1-3 MB of module source (unchanged by fraction)", + "count_at_full_source": "order 1-3 MB. Not measured; the modules in this family alone total about 200 kB.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds on data scale and cannot. Included per the name-contains-MAX rule.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_native_coverage_binding.py:47 _COMPILE_CACHE_MAX_ENTRY_BYTES", + "value": "1024**2 = 1_048_576", + "enforced_at": "acs_native_coverage_binding.py:94 (eligibility only)", + "refusal_code": "NONE \u2014 a source larger than this simply is not cached.", + "protects": "memory-or-time", + "protects_detail": "Per-entry ceiling on cached module source bytes.", + "counted_quantity": "len(source) for one Python module. Not survey data.", + "count_at_tenth": "largest module in scope is under 60 kB (unchanged by fraction)", + "count_at_full_source": "under 60 kB \u2014 the largest file in this family is acs_housing_universe_source.py at 933 lines / roughly 35 kB.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds on data scale and cannot. Included per the name-contains-MAX rule.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:42 ACS_HU_SOURCE_MAX_BYTES", + "value": "8 * 1024**3 = 8_589_934_592", + "enforced_at": "acs_housing_universe_source.py:360-364 (\"PROJECTION_TOO_LARGE\", the per-role streaming accumulator); :574 -> :180 (\"OUTPUT_SIZE\", writing full-projection.json); :577 (\"PROJECTION_TOO_LARGE\", the selected projection); :645 -> :137 (\"FILE_TOO_LARGE\", renamed to \"RECONSTRUCTION_SIZE\" at :659-661); :673 -> :180 (\"OUTPUT_SIZE\", publishing projection.json); acs_population_catalogue.py:494-497 (\"PROJECTION_SIZE\"). Also doubled into the disk-space demand at :203.", + "refusal_code": "`ACSHousingSourceError(\"PROJECTION_TOO_LARGE\")` at housing:363 and :577; `ACSHousingSourceError(\"OUTPUT_SIZE\")` at housing:180; `ACSHousingSourceError(\"FILE_TOO_LARGE\")` at housing:137, renamed to `ACSHousingSourceError(\"RECONSTRUCTION_SIZE\")` at housing:661; `ACSSourceCatalogueError(\"PROJECTION_SIZE\")` at catalogue:497. ACSHousingSourceError is declared at housing:65-66 and subclasses ValueError.", + "protects": "memory-or-time", + "protects_detail": "The lexical projection is a Python list of lists held entirely in memory and then serialized to one JSON document on disk; the module docstring for the catalogue calls this out ('The full housing lexical projection and raw catalogue remain in memory'). This is a byte budget for that document, not a wire format \u2014 nothing reads it at a fixed offset. Movable.", + "counted_quantity": "Canonical JSON bytes of the ACS lexical projection. At housing:361 the accumulator runs over EVERY applicable row of one archive (despite being named selected_bytes), so it is the full upstream projection per role; at housing:577 it is the selected projection. Full-source quantity at the accumulator sites regardless of the selection fraction.", + "count_at_tenth": "Full projection 292,484,053 bytes (unchanged by fraction \u2014 _reconstruct always builds the whole source before selecting). The selected projection at 1/10 is roughly 29 MB.", + "count_at_full_source": "292,484,053 bytes \u2014 MEASURED directly: that is the byte size of full-projection.json as the pilot wrote it, at /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/snapshots/acs-hu-capture-e5pg8w1z/full-projection.json (and byte-identical in the native-preparation capture).", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 3.4% used, 29.4x headroom. Demonstrated rather than merely computed \u2014 the 1/1000 pilot ran the full-source _reconstruct and wrote that exact 292 MB file, and the catalogue lane re-read it under catalogue:496.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:43 ACS_HU_RECEIPT_MAX_BYTES", + "value": "1024**2 = 1_048_576", + "enforced_at": "acs_housing_universe_source.py:616 (\"RECEIPT_TOO_LARGE\"); :232 -> :180 (\"OUTPUT_SIZE\", the best-effort failure.json); :647 -> :137 (\"FILE_TOO_LARGE\"/\"RECONSTRUCTION_SIZE\"); :678 -> :180 (\"OUTPUT_SIZE\"); acs_native_coverage_binding.py:562-565 (`min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES)` applied to the serialnos tuple). Also doubled into the disk demand at :204.", + "refusal_code": "`ACSHousingSourceError(\"RECEIPT_TOO_LARGE\")` at housing:616; `ACSHousingSourceError(\"OUTPUT_SIZE\")` at housing:180; `ACSHousingSourceError(\"FILE_TOO_LARGE\")` at housing:137; and, at the binding:564 site, `ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` surfaced as `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")`.", + "protects": "memory-or-time", + "protects_detail": "Two distinct roles. As a receipt cap (housing:616) it bounds a document whose contents are O(ZIP members) \u2014 per-member name/bytes/sha256/crc32/header_hex/header/rows for 6 members, plus counts as integers and selections as digests. As the min() term at binding:564 it becomes the tightest row-scaling cap in the family's evidence path, because it is applied to the literal selected key tuple. Movable.", + "counted_quantity": "housing:616 \u2014 canonical JSON bytes of the AuthenticatedACSHousingSource receipt (no per-row data). binding:564 \u2014 canonical JSON bytes of the selected ACS household SERIALNO tuple, at 16 bytes per key: 2 + (n-1) + 15n = 16n + 1.", + "count_at_tenth": "Receipt: order 25 kB (unchanged by fraction). Serialnos list at 1/10: 16 x 153,161 + 1 = 2,450,577 bytes, 2.34x over.", + "count_at_full_source": "Receipt: order 25 kB, computed from the schema against the MEASURED header widths (1,570 B household / 1,882 B person -> 3,140 and 3,764 hex characters each, x2 data members per archive) plus the column-name lists (241 and 286 names). Not measured directly; I did not import the module. Serialnos list at full source: 16 x 1,531,614 + 1 = 24,505,825 bytes, 23.4x over.", + "binds_at_tenth": true, + "binds_at_full_source": true, + "binding_note": "The receipt use never binds (about 2.4% used and not row-scaling). The binding:564 use binds at n > 65,535 selected households, i.e. selection fraction above 1/23.4 \u2014 so it binds at 1/10 and at full source, and not at 1/100. This is the hard stop that makes a full-source exact-key ACS run impossible no matter how far the row ceilings are lifted, and the lift commit did not address it. Still not the first refusal: MAX_BODY_BYTES fires 11x earlier at 5,730 households.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": true, + "protects": "memory-or-time", + "safe_to_move": true, + "agrees_with_census": false + } + }, + { + "constant": "acs_housing_universe_source.py:45 _MEMBER_MAX", + "value": "8 * 1024**3 = 8_589_934_592", + "enforced_at": "acs_housing_universe_source.py:293; acs_person_coverage_authentication.py:323. Recorded at acs_person_coverage_authentication.py:182.", + "refusal_code": "`ACSHousingSourceError(\"ZIP_MEMBER_SIZE\")` at housing:293; `ACSCoverageAuthenticationError(\"ZIP_MEMBER_SIZE\")` at auth:323", + "protects": "memory-or-time", + "protects_detail": "Per-member decompression-bomb gate, read from the ZIP central directory before the member is opened. Movable.", + "counted_quantity": "One ZIP member's declared uncompressed file_size. Upstream bytes; identical at every fraction.", + "count_at_tenth": "1,226,543,489 (largest member, unchanged by fraction)", + "count_at_full_source": "1,226,543,489 bytes \u2014 psam_pusa.csv, MEASURED from the captured archive's central directory. The other members are psam_pusb 1,178,924,624, psam_husa 495,752,055, psam_husb 481,732,465, and two 103,658-byte README PDFs.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 14.3% used, 7.0x headroom.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:46 _EXPANDED_MAX", + "value": "16 * 1024**3 = 17_179_869_184", + "enforced_at": "acs_housing_universe_source.py:295; acs_person_coverage_authentication.py:325. Recorded at acs_person_coverage_authentication.py:183.", + "refusal_code": "`ACSHousingSourceError(\"ZIP_EXPANDED_SIZE\")` at housing:295; `ACSCoverageAuthenticationError(\"ZIP_EXPANDED_SIZE\")` at auth:325", + "protects": "memory-or-time", + "protects_detail": "Per-archive cumulative decompression-bomb gate, accumulated across the central directory before any member is opened. Movable.", + "counted_quantity": "Sum of member.file_size within ONE archive (the accumulator is local to each _archive/_members call). Upstream bytes; identical at every fraction.", + "count_at_tenth": "2,405,571,771 (csv_pus, the larger archive; unchanged by fraction)", + "count_at_full_source": "csv_pus 2,405,571,771 bytes and csv_hus 977,588,178 bytes \u2014 both MEASURED.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 14.0% used at the worst archive, 7.1x headroom. Its cross-archive siblings are acs_native_coverage_binding.py:307 and acs_population_catalogue.py:149, at the same value against 3.15 GiB combined.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:47 _MEMBER_COUNT_MAX", + "value": "64", + "enforced_at": "acs_housing_universe_source.py:269; acs_person_coverage_authentication.py:303. Recorded at acs_person_coverage_authentication.py:184.", + "refusal_code": "`ACSHousingSourceError(\"ZIP_MEMBER_COUNT\")` at housing:269; `ACSCoverageAuthenticationError(\"ZIP_MEMBER_COUNT\")` at auth:303", + "protects": "memory-or-time", + "protects_detail": "Bounds the central-directory walk and the number of streams that will be opened and hashed, before any member is read. Both sites also require `0 < len(members)`. It doubles as a structural claim about the Census packaging, but its placement among ZIP_MEMBER_PATH/TYPE/ENCODING/SIZE checks makes it an archive-hygiene resource gate. Movable.", + "counted_quantity": "Number of entries in one archive's ZIP central directory \u2014 a count of groups, not rows. Identical at every fraction.", + "count_at_tenth": "3 per archive (unchanged by fraction)", + "count_at_full_source": "3 per archive \u2014 MEASURED: csv_hus.zip holds psam_husa.csv, psam_husb.csv and ACS2024_PUMS_README.pdf; csv_pus.zip holds psam_pusa.csv, psam_pusb.csv and the same README.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 3 of 64, 21x headroom. It also makes acs_population_catalogue.py:147 (\"MEMBER_LIMIT\", same value 64) dead code \u2014 catalogue:146 obtains `members` from native.coverage._members, which already enforced the identical bound at auth:303.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:48 _RECORD_MAX", + "value": "1024**2 = 1_048_576", + "enforced_at": "acs_housing_universe_source.py:239 (the refusal) and :313-316 (bounds the readline request to `min(_RECORD_MAX + 1, member.file_size - count + 1)`). Recorded in the receipt's parser_profile at :596.", + "refusal_code": "`ACSHousingSourceError(\"CSV_RECORD_SIZE\")` \u2014 acs_housing_universe_source.py:239", + "protects": "memory-or-time", + "protects_detail": "Per-record allocation ceiling for the housing owner's own physical-record parser. Movable.", + "counted_quantity": "Bytes in ONE physical CSV record including its terminator. Per-record; does not scale with the survey.", + "count_at_tenth": "1,882 worst case (unchanged by fraction)", + "count_at_full_source": "1,882 \u2014 MEASURED across every record of both archives.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 0.18% used, 557x headroom.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_housing_universe_source.py:49 _FIELD_MAX", + "value": "64 * 1024 = 65_536", + "enforced_at": "acs_housing_universe_source.py:255-258. Recorded in the receipt's parser_profile at :595.", + "refusal_code": "`ACSHousingSourceError(\"CSV_FIELD_SIZE\")` \u2014 acs_housing_universe_source.py:257", + "protects": "memory-or-time", + "protects_detail": "Per-field ceiling on the UTF-8 encoded length of each parsed CSV value. Movable.", + "counted_quantity": "Bytes in ONE CSV field after parsing. Per-field; does not scale.", + "count_at_tenth": "13 (unchanged by fraction)", + "count_at_full_source": "13 \u2014 MEASURED; longest comma-delimited field in either archive.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 5,041x headroom. Note the runtime's own csv.field_size_limit is 131,072 (recorded in the pilot producer), which is looser than this, so _FIELD_MAX is the operative field cap on this path.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:25 _ARCHIVE_BYTES", + "value": "8 * 1024**3 = 8_589_934_592", + "enforced_at": "acs_population_catalogue.py:481. Checked for identity against the literal 8*1024**3 in the producer record at catalogue:110/:118-124 (\"LIMITS\"). Also consumed for disk planning at survey_population_preparation.py:510-511 and :548.", + "refusal_code": "`ACSSourceCatalogueError(\"ARCHIVE_LIMIT\")` \u2014 acs_population_catalogue.py:481 (class declared at catalogue:35-36, subclasses ValueError)", + "protects": "memory-or-time", + "protects_detail": "Same role as MAX_ARCHIVE_BYTES: an operational ceiling on the compressed bytes about to be copied into a private capture. Movable.", + "counted_quantity": "Sum of the two pinned archives' declared size_bytes. Upstream compressed bytes; identical at every fraction.", + "count_at_tenth": "854,347,733 (unchanged by fraction)", + "count_at_full_source": "854,347,733 \u2014 from the manifest and confirmed on disk.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 9.9% used, 10.05x headroom. Demonstrated: the 1/1000 pilot ran issue_acs_source_catalogue to completion past this gate.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:26 _EXPANDED_BYTES", + "value": "16 * 1024**3 = 17_179_869_184", + "enforced_at": "acs_population_catalogue.py:145-149. Identity-checked in the producer record at catalogue:111.", + "refusal_code": "`ACSSourceCatalogueError(\"EXPANDED_LIMIT\")` \u2014 acs_population_catalogue.py:149", + "protects": "memory-or-time", + "protects_detail": "Cross-archive decompression gate in the catalogue's own _preflight, whose docstring at catalogue:141 says it 'Keep[s] full-source gates separate from native's 1M selected-row gate'. Movable.", + "counted_quantity": "Sum of member.file_size across BOTH archives (the accumulator spans the loop at catalogue:143-149). Upstream bytes; identical at every fraction.", + "count_at_tenth": "3,383,159,949 (unchanged by fraction)", + "count_at_full_source": "3,383,159,949 bytes (3.151 GiB) \u2014 MEASURED.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 19.7% used, 5.08x headroom. Demonstrated by the completed pilot catalogue.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:27 _MEMBERS", + "value": "64", + "enforced_at": "acs_population_catalogue.py:147. Identity-checked in the producer record at catalogue:112.", + "refusal_code": "`ACSSourceCatalogueError(\"MEMBER_LIMIT\")` \u2014 acs_population_catalogue.py:147", + "protects": "memory-or-time", + "protects_detail": "Bounds the number of archive members the catalogue will inventory. Movable.", + "counted_quantity": "Number of ZIP central-directory entries in one archive \u2014 a group count, not rows. Identical at every fraction.", + "count_at_tenth": "3 per archive (unchanged by fraction)", + "count_at_full_source": "3 per archive \u2014 MEASURED.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds AND cannot fire: `members` at catalogue:147 is the return value of native.coverage._members(archive, role) called at :146, which already enforced `0 < len(members) <= custody._MEMBER_COUNT_MAX` at the identical value 64 (acs_person_coverage_authentication.py:303). Dead code \u2014 the same upstream-gate-first pattern the transport lane found in survey_population_domains.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:28 _SOURCE_ROWS", + "value": "6_000_000", + "enforced_at": "acs_population_catalogue.py:153-157. Identity-checked in the producer record at catalogue:113.", + "refusal_code": "`ACSSourceCatalogueError(\"SOURCE_ROW_LIMIT\")` \u2014 acs_population_catalogue.py:156", + "protects": "upstream-file-size", + "protects_detail": "Same quantity and same number as acs_person_coverage_columns.MAX_ROWS: the upstream archive's own rows, counted by a scan (catalogue:151 passes `frozenset()`, so nothing is retained \u2014 and catalogue:152 asserts `not empty`). Treat as not movable downward.", + "counted_quantity": "Upstream archive data rows per role. Identical at every fraction.", + "count_at_tenth": "household 1,631,969; person 3,422,888 (unchanged by fraction)", + "count_at_full_source": "household 1,631,969; person 3,422,888 \u2014 MEASURED, and equal to the pilot receipt's `counts` for the same run.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds (57% used) AND cannot fire: coverage._inventory already refuses above 6,000,000 at acs_person_coverage_authentication.py:373 while scanning, so the sum it returns is always within this bound. Dead code, same as acs_native_coverage_binding.py:315/:330/:344.", + "family": "ACS source and coverage", + "verified": { + "binds_at_full_source": false, + "protects": "upstream-file-size", + "safe_to_move": false, + "agrees_with_census": false + } + }, + { + "constant": "acs_population_catalogue.py:29 _BATCH_PEOPLE", + "value": "100_000", + "enforced_at": "acs_population_catalogue.py:252 (\"HOUSEHOLD_EXCEEDS_BATCH\"); acs_population_catalogue.py:311-315 (\"BATCH_PERSON_LIMIT\", as `min(_BATCH_PEOPLE, literal.MAX_SELECTED_ROWS)`); and as the batching threshold in the assembly loop at acs_population_catalogue.py:327-333 (no refusal there). Identity-checked at catalogue:114.", + "refusal_code": "`ACSSourceCatalogueError(\"HOUSEHOLD_EXCEEDS_BATCH\")` at catalogue:252; `ACSSourceCatalogueError(\"BATCH_PERSON_LIMIT\")` at catalogue:314", + "protects": "memory-or-time", + "protects_detail": "The comment at catalogue:308-309 says it preserves 'the prior selected-reader ceiling on each assembly batch, even though no selected DataFrame is allocated here'. It is the streaming batch size that keeps whole-household record assembly bounded. Movable.", + "counted_quantity": "catalogue:252 \u2014 NP for ONE household. catalogue:314 \u2014 the number of person rows in the current assembly batch. Both are full-source quantities (the catalogue always runs at full source), never the selected roster.", + "count_at_tenth": "catalogue:252 worst case 20; catalogue:314 at most 100,000 by construction (unchanged by fraction \u2014 the catalogue does not vary with the selection fraction at all)", + "count_at_full_source": "catalogue:252 worst case 20 persons; catalogue:314 at most 100,000. The pilot recorded literal_batches = 35 for 3,422,888 people across 1,631,969 households \u2014 97,797 people per batch, i.e. the loop rides right at the 100,000 threshold as designed.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "NEITHER SITE CAN FIRE. catalogue:252 tests one household's NP against 100,000, but NP is already constrained to 0..20 by housing._tables at acs_housing_universe_source.py:441 (`(\"NP\", 0, 20)`), which ran at catalogue:222 \u2014 a structural invariant makes it unreachable by a factor of 5,000. catalogue:314 tests the batch's person count, but the loop at catalogue:329 consumes and resets the batch whenever `batch_people + count > _BATCH_PEOPLE`, so the only way to exceed it is a single household above 100,000, which catalogue:252 already refused. This is the exact pattern the task flagged in survey_population_domains: the tighter batching limit applies first, and the nominal ceiling is dead. The `literal.MAX_SELECTED_ROWS` term in the min() is dead twice over.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:30 _CATALOGUE_BYTES", + "value": "2 * 1024**3 = 2_147_483_648", + "enforced_at": "acs_population_catalogue.py:202 (_json cap on one record, shrinking as `max(0, _CATALOGUE_BYTES - size - 1)`) and :204 (\"CATALOGUE_BYTE_LIMIT\"); :232 (_json cap, `max(0, _CATALOGUE_BYTES - prospective_size - punctuation)`) and :235 (\"CATALOGUE_BYTE_LIMIT\"). The two accountings must agree at :335 (\"CATALOGUE_BYTE_ACCOUNTING\"). Identity-checked at catalogue:115.", + "refusal_code": "`ACSSourceCatalogueError(\"CATALOGUE_BYTE_LIMIT\")` at catalogue:204 and :235; `ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` at the :202/:232 _json caps, surfaced as `ACSSourceCatalogueError(\"CATALOGUE_ISSUANCE_REFUSED\")` (catalogue:603-604).", + "protects": "memory-or-time", + "protects_detail": "A byte budget for the in-memory raw catalogue, reserved before literal fields are retained (the `reserve` comment at catalogue:224-226) and charged again as records are emitted. The module docstring is explicit that 'The full housing lexical projection and raw catalogue remain in memory. ... this does not make national preparation memory-bounded.' Movable.", + "counted_quantity": "Canonical JSON bytes of every catalogue record: one household tuple (SERIALNO, TYPEHUGQ, NP, WGTP, source_member, source_row_ordinal, members) plus one nested person tuple per member. Full upstream source, at every selection fraction \u2014 the catalogue is always built with serialnos=None (catalogue:489).", + "count_at_tenth": "374,174,650 bytes (unchanged by fraction)", + "count_at_full_source": "374,174,650 bytes \u2014 MEASURED, read directly from the pilot receipt at /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json, key /catalogues/acs/counts/canonical_record_bytes, for the run that inventoried 1,631,969 households and 3,422,888 people.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds: 17.4% used, 5.74x headroom. This is the family's largest genuinely-consumed byte budget and it is comfortably clear. Demonstrated, not estimated \u2014 the 1/1000 pilot ran this exact full-source path to completion.", + "family": "ACS source and coverage" + }, + { + "constant": "acs_population_catalogue.py:31 _RECEIPT_BYTES", + "value": "1024**2 = 1_048_576", + "enforced_at": "acs_population_catalogue.py:472-475 (\"CANDIDATE_TYPE_SIZE\"); and as the _json cap at :440, :449, :483, :561 and :587, refusing inside coverage._json at acs_person_coverage_authentication.py:79/:120. Identity-checked at catalogue:116.", + "refusal_code": "`ACSSourceCatalogueError(\"CANDIDATE_TYPE_SIZE\")` at catalogue:474; `ACSCoverageAuthenticationError(\"CANONICAL_SIZE\")` at the _json sites, surfaced as `ACSSourceCatalogueError(\"CATALOGUE_ISSUANCE_REFUSED\")` (catalogue:603-604) or `ACSSourceCatalogueError(\"CATALOGUE_VERIFICATION_REFUSED\")` (catalogue:457-458).", + "protects": "memory-or-time", + "protects_detail": "Canonical-encoding ceiling on the compact catalogue receipt and on the producer record it embeds. Movable.", + "counted_quantity": "Canonical JSON bytes of the catalogue receipt (protocol, producer closure, archives, the 6-member inventory, counts as integers, the field contract, the record schema) and of the producer record alone. O(members) and O(manifest entries), never O(rows) \u2014 the 1,631,969 and 3,422,888 totals enter only as integers.", + "count_at_tenth": "order 60-150 kB (unchanged by fraction)", + "count_at_full_source": "order 60-150 kB, dominated by the embedded implementation manifest (native._producer -> implementation_manifest module and resource digests). Not measured directly \u2014 I did not import the module. It is demonstrably under 1 MiB, because the 1/1000 pilot issued this receipt at full source and recorded its sha256 at /catalogues/acs/receipt_sha256.", + "binds_at_tenth": false, + "binds_at_full_source": false, + "binding_note": "Never binds and cannot scale with the survey: no per-row datum is charged against it. Note this is the same 1 MiB value that, at acs_native_coverage_binding.py:564 via ACS_HU_RECEIPT_MAX_BYTES, DOES bind hard \u2014 the difference is entirely whether the document embeds the roster.", + "family": "ACS source and coverage" + } + ], + "verdicts": [ + { + "constant": "asec_housing_status.py:54 ROW_BYTES", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": true, + "correction": "No substantive correction \u2014 value, classification, and binds-at-full verdict all hold. Four precision addenda the census row should absorb: (1) the per-column width check it cites spans asec_housing_status.py:299-304, not 299-302, and its refusal literal is \"BUFFER_SIZE\" (line 303), which the row omits. (2) The derived budget's refusal SITE is asec_housing_status.py:436-440 (literal \"PAYLOAD_SIZE\" at line 440); the row cited only the exception definition (86-92). Exception type HousingStatusRefusalError (line 86, ValueError subclass) raised by _require (90-92) is correct. (3) \"the literal serialized width of one household row\" is loose: the wire layout is column-major (11 int64 buffers of n*8 bytes then 14 uint8 buffers of n bytes, _parts 406-410, _issue 390-394), so no contiguous 102-byte row exists; 102 is the per-household byte TOTAL across those buffers. The arithmetic and the conclusion are unaffected. (4) For the record, the N that ROW_BYTES multiplies is the ASEC verified-source household roster (source.scope.household_ids, asec_housing_status_source.py:70-120), measured at 55,762 by this lane's own roster census (test_us_native_row_ceilings.py:35) \u2014 NOT the 1,587,376 stacked survey. Actual payload ~5.69 MB (55,762*102) against a 40,865,581-byte (~39 MB) budget. The sibling MAX_HOUSEHOLDS = 400_000 (line 53, enforced 174 and 294 with \"ROW_COUNT\") is thus an upstream-roster bound and must not be lifted for full-source scale; the row under test correctly does not claim otherwise.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and not necessary \u2014 it never binds. Two independent code-read reasons. (a) Fixed-width encoding: ROW_BYTES is not an independent knob but the arithmetic restatement of the actual serialized column widths, which are enforced per column at asec_housing_status.py:300-304 (\"BUFFER_SIZE\", len(buf) == n * (8 if RAW else 1)) and fixed by the dtypes at _issue 361-363 and array() 278-281. It is consumed only by PAYLOAD_MAX_BYTES (line 55), on which a reader depends: decode_housing_status refuses len(payload) > PAYLOAD_MAX_BYTES with \"PAYLOAD_SIZE\" (436-440) and read_housing_status truncates its disk read at PAYLOAD_MAX_BYTES + 1 (line 473). Shrink ROW_BYTES and legitimate full-size artifacts are truncated and refused; grow it and the budget goes slack. The budget is currently exactly tight: header <= HEADER_MAX_BYTES (102/445) plus buffers <= ROW_BYTES*MAX_HOUSEHOLDS (ROW_COUNT at 174/294) plus MAGIC+4+32, so raising the separate MAX_HOUSEHOLDS would already scale the budget automatically without touching ROW_BYTES. (b) Module-identity coupling the census row missed: _definition_identity() (151-168) hashes this module's own source bytes, stored as header[\"definition_sha256\"] (382) and required by validate() with \"DEFINITION_CHANGED\" (290-292); authenticated statuses additionally require data[\"implementation\"] == _implementation() with \"IMPLEMENTATION_CHANGED\" (323-328), and asec_housing_status_source._implementation (44-58) also hashes asec_housing_status.py. Any edit to this file, ROW_BYTES included, invalidates every previously issued artifact. No tests or committed fixtures pin PAYLOAD_MAX_BYTES/ROW_BYTES (grep over packages/*/tests found none), so the breakage would surface at build time, outside PR CI.", + "evidence": "packages/microcosm-build/src/microcosm/build/us_runtime/asec_housing_status.py:19-31 (RAW_COLUMNS, 11 entries); :34-49 (DERIVED_COLUMNS, 14 entries); :50 (COLUMNS); :51-52 (MAGIC 9 bytes, HEADER_MAX_BYTES 65536); :53 (MAX_HOUSEHOLDS = 400_000); :54 (ROW_BYTES = 8*len(RAW_COLUMNS)+len(DERIVED_COLUMNS) = 102); :55 (PAYLOAD_MAX_BYTES = 9+4+65536+102*400000+32 = 40,865,581); :86-87 (HousingStatusRefusalError(ValueError)); :90-92 (_require raises it); :151-168 (_definition_identity hashes asec_housing_status.py bytes); :174 (_require 0 < n <= MAX_HOUSEHOLDS, \"ROW_COUNT\"); :278-281 (array(): ' two package records); :319-327 (_aggregates: rows, nonzero_rows, unweighted_total, minimum, maximum \u2014 five scalars per root); :344-345 (rows = frame.n(\"person\"); _require(0 < rows <= EVALUATION_MAX_ROWS, \"ENGINE_ROW_BOUND\")); :356-384 (header construction, incl. \"person_rows\" :366, \"aggregates\" :382, \"bindings\" :383); :388-403 (_evaluation_parts) with :391 (_require(0 < len(header) <= EVALUATION_HEADER_MAX_BYTES, \"HEADER_SIZE\")) and :395 (parts = [EVALUATION_MAGIC, struct.pack(\" MAX_ROWS: raise ValueError(...)`). Catalogue twin: acs_population_catalogue.py:28 (`_SOURCE_ROWS = 6_000_000`), :154-157 (`\"SOURCE_ROW_LIMIT\"`), :35-41 (ACSSourceCatalogueError), :112-125 (limits/ceilings tuple pinning it at 6_000_000, refusal `\"LIMITS\"`). Producer publication of the constant: acs_native_coverage_binding.py:261-268. Pin of the upstream gate's file: acs_native_coverage_binding.py:32-42 (_ACCEPTED) enforced at :233-237 (\"UNREVIEWED_PREPARATION\"); digests verified against the working tree with shasum -a 256 \u2014 all four match. Counts (not re-derived, cross-checked against in-repo receipts): experiments/native-row-ceilings/roster-census.json:11 (`\"households\": 1631969`), :16 (`\"people\": 3422888`), :27-28 (acs_households 1531614 = occupied_hu 1348408 + institutional_gq 84422 + noninstitutional_gq 98784; 1631969 = that plus vacancies 100355). The person figure is tied to the source file's own row count in code by acs_population_catalogue.py:296-301 (`_require(rows == len(positions), \"GLOBAL_PERSON_KEYS\")` against literal._scan_acs_person_coverage). Not established: I did not independently confirm that the ACS housing archive's raw CSV data-row count equals the catalogue's 1,631,969 household keys (no equivalent in-code identity read); I relied on the catalogue's own whole-source housing count, which is ~1.6M either way and far below 6,000,000. Also noted: acs_person_coverage_columns.py:38 cites docs/us-native-row-ceilings.md, which does not exist in this worktree (checked docs/ and a repo-wide find)." + }, + { + "constant": "acs_population_catalogue.py:28 _SOURCE_ROWS = 6_000_000", + "binds_at_full_source_verdict": false, + "protects_verdict": "upstream-file-size", + "agrees_with_census": false, + "correction": "Every load-bearing element of the row is CONFIRMED (value, quantity, counts, refusal code + exception type, \"binds at full: false\", the unreachability argument, and the upstream-file-size classification). Three defects:\n\n(1) CITATION OFF BY ONE, twice. The row says \"catalogue:151 passes `frozenset()`\" and \"catalogue:152 asserts `not empty`\". Actual: line 152 is `inventory, empty = native.coverage._inventory(path, role, frozenset())`; line 153 is `_require(not empty, \"UNEXPECTED_SELECTED_ROWS\")`. Line 151 is the `for role, path in paths.items():` header. Likewise \"enforced at: 153-157\" over-includes line 153, which carries a different refusal code (\"UNEXPECTED_SELECTED_ROWS\"); the SOURCE_ROW_LIMIT require is exactly acs_population_catalogue.py:154-157, literal at :156.\n\n(2) MOVABILITY DIRECTION IS BACKWARDS. The row says \"Treat as not movable downward.\" For a category-(c) bound the danger is UPWARD: raising 6,000,000 weakens the structural assertion that a genuine 2024 ACS PUMS archive is near its real size. The repo says so in its own words at acs_person_coverage_columns.py:33-38 (\"MAX_ROWS bounds the source file's own person records and stays where it is ... the bound is the structural assertion that a genuine file is near that size. Only the requested-roster ceiling moves\") and this lane's own test asserts it at packages/microcosm-build/tests/test_us_native_row_ceilings.py:88-100. The row should read \"must not be moved in either direction\".\n\n(3) TWO GUARDS THE ROW OMITS, both of which strengthen \"do not touch this\".\n (a) A hardcoded second copy of the number: acs_population_catalogue.py:118 `ceilings = (8*1024**3, 16*1024**3, 64, 6_000_000, 100_000, 2*1024**3, 1024**2)` with `_require(all(type(v) is int and 0 < v <= cap ...), \"LIMITS\")` at :119-125. Raising `_SOURCE_ROWS` alone makes `_producer()` refuse with ACSSourceCatalogueError(\"LIMITS\") on every issuance AND every re-validation. Any change also moves `limits` (:137) and `catalogue_sha256` (:135) inside the producer record, which the pilot receipt pins (`/producer/acs/limits` 4th element = 6000000, `/producer/acs/catalogue_sha256` = 4e3118be...).\n (b) Because `_inventory` raises ACSCoverageAuthenticationError (a different type, acs_person_coverage_authentication.py:57, raised at :63) and `issue_acs_source_catalogue` converts every non-ACSSourceCatalogueError to ACSSourceCatalogueError(\"CATALOGUE_ISSUANCE_REFUSED\") at :601-604, an operator feeding an oversized real archive would see CATALOGUE_ISSUANCE_REFUSED, never \"SOURCE_ROW_LIMIT\". The row's \"dead code\" verdict is right, and this is why.\n\nAlso, for the record, the row's \"Dead code\" is a reachability fact, not license to delete: deleting it would not weaken the real check (authentication:373 does the work) but would change the module bytes and the producer record.", + "safe_to_move": false, + "safe_to_move_reason": "Category (c), an upstream file's true size. `_preflight` passes `frozenset()` (acs_population_catalogue.py:152), so `_inventory` retains nothing and `member[\"rows\"]` (acs_person_coverage_authentication.py:410) is the raw CSV data-row count of each member of the real ACS PUMS zip for that role, summed per role at :155. That is the archive's own size, not the selected/stacked/cloned roster: issue_acs_source_catalogue:508-512 requires `counts[\"households\"] == sum(m[\"rows\"] for m in inventories[\"household\"])` and the same for persons under \"GLOBAL_SOURCE_COMPLETENESS\", and the recovered 1/1000 pilot recorded those counts as 1,631,969 / 3,422,888 \u2014 i.e. unchanged by the selection fraction (the same run's selected native/acs rows are 1,529 households / 3,324 persons). It is the same number and the same quantity as acs_person_coverage_columns.MAX_ROWS (=6_000_000, :39), whose comment at :33-38 declares it a structural assertion about the real file, and as acs_native_coverage_binding.MAX_SOURCE_ROWS (:30, \"per role, before full-source construction\"). Raising it weakens a real check on a file this build does not produce. Independently, it cannot be raised without also editing the hardcoded 6_000_000 in the ceilings tuple at :118 or `_producer()` refuses with \"LIMITS\", and any edit moves `limits`/`catalogue_sha256` in the producer record that every receipt re-verifies (:440, :449). No fixed-width encoding or byte budget is derived from it (`_CATALOGUE_BYTES`/`_RECEIPT_BYTES` are enforced by actual accumulation in `_charge` :201-206 and `reserve` :227-235), so it is not category (b) \u2014 but it is firmly category (c).", + "evidence": "VALUE + PRODUCER PIN: packages/microcosm-build/src/microcosm/build/us_runtime/acs_population_catalogue.py:28 (`_SOURCE_ROWS = 6_000_000`); :113 (inside the `limits` tuple, :109-117); :118 (`ceilings` hardcodes 6_000_000 as the cap); :119-125 (`_require(..., \"LIMITS\")`); :135-137 (`catalogue_sha256`, `\"limits\": list(limits)` in the producer record).\nENFORCEMENT: acs_population_catalogue.py:154-157 \u2014 `_require(sum(member[\"rows\"] for member in inventory) <= _SOURCE_ROWS, \"SOURCE_ROW_LIMIT\")`, literal at :156. Exception type ACSSourceCatalogueError (subclass of ValueError) defined at :36-37, raised by `_require` at :40-42. Preceding lines: :152 `_inventory(path, role, frozenset())`, :153 `_require(not empty, \"UNEXPECTED_SELECTED_ROWS\")`. Wrapper conversion at :601-604.\nQUANTITY: acs_person_coverage_authentication.py:334-335 (`rows` initialized once per `_inventory` call, i.e. per role); :371-373 (`rows += 1; row_count += 1; _require(rows <= literal.MAX_ROWS, \"SOURCE_ROWS\")` \u2014 counted only for applicable `psam_hus`/`psam_pus` members, after the header `continue` at :366); :410 (`\"rows\": row_count` in the returned inventory member), so `sum(member[\"rows\"]) == rows` exactly. Exception type there: ACSCoverageAuthenticationError (ValueError) at :57, raised at :61-63.\nTIGHTER UPSTREAM LIMIT (makes it unreachable): acs_person_coverage_authentication.py:373 refuses at the 6,000,001st row of the same per-role total before `_inventory` returns \u2014 so acs_population_catalogue.py:155 can never see a sum above 6,000,000. Identical redundancy at acs_native_coverage_binding.py:314-315, :329-330, :343-344 (refusal code \"SOURCE_ROW_BUDGET\"), against MAX_SOURCE_ROWS at :30.\nTWIN CONSTANT + ITS DOCUMENTED INTENT: acs_person_coverage_columns.py:33-38 (comment), :39 (`MAX_ROWS = 6_000_000`), :40 (`MAX_SELECTED_ROWS = 14_000_000`), :198 (second use, person scan).\nCOUNTS (measured, not derived): /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json, sha256 34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0 (I recomputed it): `/catalogues/acs/counts/households = 1631969`, `/catalogues/acs/counts/people = 3422888`, and `/producer/acs/limits = [8589934592, 17179869184, 64, 6000000, 100000, 2147483648, 1048576]`. Cross-check: occupied_hu 1,348,408 + institutional_gq 84,422 + noninstitutional_gq 98,784 + vacancies 100,355 = 1,631,969. Committed copy at experiments/native-row-ceilings/roster-census.json:11,16. Utilisation: 3,422,888/6,000,000 = 57.0% (person role), 1,631,969/6,000,000 = 27.2% (household role).\nLANE POSITION: packages/microcosm-build/tests/test_us_native_row_ceilings.py:88-100 (`test_the_acs_source_file_bound_does_not_move`) asserts MAX_ROWS == 6_000_000 and MAX_SOURCE_ROWS == 6_000_000; it does NOT cover acs_population_catalogue._SOURCE_ROWS, the third copy.\nNOT ESTABLISHED / SIDE NOTE: `docs/us-native-row-ceilings.md`, cited by acs_person_coverage_columns.py:38 and by the new test's docstring, does not exist on disk or in `git ls-files docs` at HEAD 2ebd246f1." + }, + { + "constant": "asec_housing_universe.py:40 HEADER_MAX_BYTES = 65536", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "Three corrections. The bottom-line \"binds at full: false\" is correct, but the classification and the enforcement inventory are both wrong.\n\n(1) CLASSIFICATION IS WRONG: not encoding-width, it is memory-or-time. The census says \"Header length is a ` MAX_ROWS: raise ValueError(\"ACS coverage source row bound exceeded\")` \u2014 bare ValueError; grep of the module shows only bare `raise ValueError(...)` (lines 97,100,107,110,113,115,133,141,145,157,169,179,192,199,201,235,237,240,250,260,267), no _require and no typed error class.\n- acs_person_coverage_authentication.py:373 `_require(rows <= literal.MAX_ROWS, \"SOURCE_ROWS\")`; :61-63 `def _require(condition, code): raise ACSCoverageAuthenticationError(code)`; :57 `class ACSCoverageAuthenticationError(ValueError)`.\n- acs_person_coverage_authentication.py:173 `\"source_rows\": literal.MAX_ROWS` inside `_producer()` (def :125); compared to persisted receipts at :595 `_require(coverage.receipt[\"producer\"] == _producer(), \"PRODUCER_CHANGED\")`.\n- git show 14defbfc0 -- acs_person_coverage_columns.py: only `MAX_SELECTED_ROWS 1_000_000 -> 14_000_000` plus the 6-line comment; MAX_ROWS untouched.\n\nCOUNTED QUANTITY IS FRACTION-INDEPENDENT:\n- acs_person_coverage_columns.py:185-206 \u2014 members = sorted psam_pus*.csv of `source.person_zip`; `rows` counted for every record before `consume_row`, whose retention filter is at :244-245.\n- acs_person_coverage_authentication.py:334-411 `_inventory(path, role, serialnos)`; :300 `prefix = \"psam_hus\" if role == \"household\" else \"psam_pus\"`; :370-373 rows += 1 and the bound, then :375 `if cells[0] not in serialnos: continue`; :404-409 per-member `\"rows\": row_count`.\n\nINDEPENDENT MEASUREMENT (this session, read-only, streaming the pinned archives at /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/snapshots/acs-hu-capture-e5pg8w1z/):\n- sha256 csv_hus.zip = 8281008e53de98f0ef81e7a2ee5a8725991dda1ecfd2713ead73246425e515d0, csv_pus.zip = afdc6d90c6e2f0bab365ed32d95ba4c4d8ac651162f46ac7861295b2dc469894 \u2014 both match the pins in acs_2024_1yr_sources.json exactly (sizes 251500587 / 602847146 match too).\n- Streamed each member: psam_pusa.csv 1,226,543,489 B, 1,743,752 newlines, 0 `\"` bytes, ends with newline -> 1,743,751 data rows; psam_pusb.csv 1,178,924,624 B -> 1,679,137; total person 3,422,888. psam_husa.csv 495,752,055 B -> 827,133; psam_husb.csv 481,732,465 B -> 804,836; total household 1,631,969. Zero quote bytes anywhere, so line count equals record count exactly.\n- grep of the pilot receipt financial-artifacts/preparation.json returns exactly the strings 3422888 and 1631969 \u2014 my counts match the receipt.\n\nNO TIGHTER UPSTREAM GATE:\n- acs_housing_universe_source.py:45-47 `_MEMBER_MAX = 8*1024**3`, `_EXPANDED_MAX = 16*1024**3`, `_MEMBER_COUNT_MAX = 64`; enforced acs_person_coverage_authentication.py:293-295 (\"ZIP_MEMBER_SIZE\"/\"ZIP_EXPANDED_SIZE\") before any row is read.\n- acs_native_coverage_binding.py:28-30 MAX_ARCHIVE_BYTES 8 GiB / MAX_EXPANDED_BYTES 16 GiB / MAX_SOURCE_ROWS 6_000_000; acs_population_catalogue.py:24-28 _ARCHIVE_BYTES/_EXPANDED_BYTES/_MEMBERS/_SOURCE_ROWS, checked at :143-157 after `_inventory` returns.\n- Measured zip expanded totals: csv_pus.zip 2.24 GiB across 3 members, csv_hus.zip 0.91 GiB \u2014 ~703 B per person row, so 6,000,000 rows (~4.2 GB) is reached long before any byte ceiling.\n\nCORROBORATION:\n- acs_pums.py:50-56 \"Neither is a statement about the source file, which MAX_ROWS in acs_person_coverage_columns still makes at 6,000,000\"; MAX_EXACT_HOUSEHOLDS = 7_000_000, MAX_EXACT_PERSON_ROWS = 14_000_000.\n- packages/microcosm-build/tests/test_us_native_row_ceilings.py:92-99 asserts MAX_ROWS == 6_000_000 and > ACS_PERSONS." + }, + { + "constant": "asec_housing_status.py:52 HEADER_MAX_BYTES = 65536", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "I confirm the census's headline verdict (row-independent, cannot bind at full source) but three of its supporting statements are wrong, and its classification rests on a non-sequitur.\n\n1. WRONG CLASSIFICATION REASONING (the ` raise HousingStatusRefusalError(reason)\n :101-108 _parse; :102 _require(type(value) is bytes and 0 < len(value) <= HEADER_MAX_BYTES, \"HEADER_SIZE\") \u2014 OUTSIDE the try at :103; :105 \"HEADER_CANONICAL\" inside the try; :107-108 except (ValueError, ...) -> HousingStatusRefusalError(\"HEADER_FORMAT\")\n :115-148 _codebook(); :67-83 DICTIONARIES; :57-65 ZERO_ORIGIN_EVIDENCE\n :151-168 _definition_identity() hashes this module's own bytes plus codebook/dictionaries/zero_origin_evidence/dependency versions\n :174 _require(0 < n <= MAX_HOUSEHOLDS, \"ROW_COUNT\")\n :269-272 header_data -> _parse(self.header)\n :283-328 validate(); :291 \"DEFINITION_CHANGED\"; :294 \"ROW_COUNT\"; :299-304 \"BUFFER_SIZE\"; :323-328 \"IMPLEMENTATION_CHANGED\"\n :348-396 _issue(); :369-388 header dict, incl. :378 \"household_rows\": len(raw), :383 \"source\": source, :384 \"implementation\": implementation; :391 _json(header); :395 result.validate()\n :406-410 _parts(); :408 struct.pack(\" AuthenticatedAsecSource(_json(evidence))\n\npackages/microcosm-build/src/microcosm/build/us_runtime/graph_implementation_inventory.json: asec_housing_status.py appears in stages asec_prepared_v3 (modules[17]), composed_asec_binding_v1 (modules[19]), composed_population_v1 (modules[19]).\n\nMeasurements run with the worktree's existing .venv/bin/python (numpy 2.4.6, pandas 3.0.3), read-only, no installs, no repository file modified; scripts written to the session scratchpad only:\n synthetic header bytes: n=1 -> 2443, n=10 -> 2444, n=1000 -> 2446, n=100000 -> 2448, n=400000 -> 2448\n payload bytes: n=400000 -> 40,802,493 (vs PAYLOAD_MAX_BYTES 40,865,581)\n authenticated header estimate (real FIELDS roster + placeholder digests): 5,887 bytes; source 2,769 / codebook 688 / implementation 673 / zero_origin_evidence 578 / columns 361 / dictionaries 264\n refusal checks: _parse(b'{\"a\":\"'+b'x'*70000+b'\"}') -> HousingStatusRefusalError('HEADER_SIZE'); _parse(b'') -> HousingStatusRefusalError('HEADER_SIZE'); _parse(b'{\"b\":1, \"a\":2}') -> HousingStatusRefusalError('HEADER_FORMAT'); decode_housing_status(MAGIC + pack(\" HousingStatusRefusalError('HEADER_SIZE'); issubclass(HousingStatusRefusalError, ValueError) -> True" + }, + { + "constant": "asec_income_observations.py:43 _HEADER_MAX = 65536", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "Agree on the load-bearing verdict (row-independent, does not bind at full source), but the row is wrong on two points and thin on a third.\n\n(1) WRONG CLASSIFICATION. The row says \"encoding-width\". The code does not support it. The on-wire header-length field is uint32: `MAGIC + struct.pack(\" n = 432,523.\nImplementation digest coupling: asec_income_observations.py:150-168 (_implementation includes _sha of asec_income_observations.py), :333 and :469 (re-verified), graph_asec_income.py:201.\nMeasurements I ran in the worktree's existing .venv (read-only, no installs): _implementation() JSON = 1344 bytes; contract price = 234 bytes, monetary_basis = 715 bytes, whole resource = 3274 bytes; reconstructed money evidence JSON = 2267 bytes; full reconstructed header = 9424 bytes vs _HEADER_MAX 65536.\nRepo-wide grep for 65536/65_536 across *.py/*.json/*.md: no other module, test, manifest or checked-in receipt pins this bound." + }, + { + "constant": "asec_current_money_selection.py:83 SELECTED_PAYLOAD_MAX_BYTES", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The operative verdict fields hold (value, sites, refusal code, counted quantity, binds-at-full=false, encoding-width, not safe to move), but four statements in the row are wrong or unsupported and one material mechanism is missing.\n\n(1) MISCITED MECHANISM. The row cites \"`_positions`, asec_current_money_selection.py:311-325\" for the strictly-increasing duplicate-free slice. `_positions` is defined at asec_current_money_selection.py:122-132 (`0 < len(values) <= rows` at :129, strict increase at :130, `values[-1] < rows` at :131, refusal codes \"PERSON_POSITIONS\"/\"HOUSEHOLD_POSITIONS\" supplied at the call sites :347-353). Lines 311-325 are the BoundCurrentMoneyBody constructor call and the `select_current_money` signature \u2014 no ordering check lives there. The cited claim is true; the citation is not.\n\n(2) MISCITED COMPOSED-PATH CHECK. graph_composed_asec_binding.py:313 is a docstring line inside `_resolved_positions` (:306-332). The enforcement is at :322-325 (\"BINDING_AMBIGUOUS_ROWS\") and :328-331 (\"BINDING_NON_MONOTONE\"), with \"BINDING_UNRESOLVED\" at :327. Those are the requires that make a cloned/duplicated roster unresolvable against the body, so the conclusion survives, but the row points at prose rather than at the check.\n\n(3) THE \"~154.1 MB BODY\" AND \"~2.3x HEADROOM AT THE THEORETICAL MAXIMUM\" ARE NOT ESTABLISHED. I could not reproduce 154.1 MB from any roster in the lane's own artifact. Under the codec's own formula 11*(32*P + H): ASEC catalogue (142,125 p / 55,762 h) = 50.6 MB; ASEC including the 33,170 unrepresented households (est. ~226,700 p) = 80.8 MB; stacked = 1,272 MB; clone = 2,545 MB; at the row ceilings = 356.4 MB. The true theoretical maximum is NOT 2.3x under the cap \u2014 with the canonical 32-person/1-household roster and a maximal 65,536-byte header the payload equals 356,465,581 bytes EXACTLY, i.e. the cap is tight to the byte at MAX_PERSONS/MAX_HOUSEHOLDS by construction. The real headroom statement is the full-source one: 50,641,427 bytes + header (<= 65,536) vs 356,465,581 = 7.0x.\n\n(4) \"~52 MB at full source\" is ~3% high. Exact: 9 + 4 + len(header) + 11*(32*142125 + 55762) + 32 = 50,641,427 + len(header) bytes, i.e. ~50.6-50.7 MB (48.3 MiB). The roster itself is right and comes from the lane's own artifact (experiments/native-row-ceilings/roster-census.json:29-30) and from the pilot preparation.json (/catalogues/asec/counts/households = 55762, persons = 142125).\n\n(5) MISSING: A TIGHTER CHECK UPSTREAM OF THE ENCODE SITE. `_selected_parts` calls `_validate_selected` at :446 BEFORE the :461 PAYLOAD_SIZE require. `_validate_selected` already enforces rows <= MAX_PERSONS/MAX_HOUSEHOLDS at :430-435 (\"SELECTED_ROW_BOUND\"), `_field_shapes` fixes each field's buffers at 11*count bytes (:149-168, \"FIELD_BUFFER_SHAPE\"), and `_parse(selected.header, HEADER_MAX_BYTES)` at :399 caps the header (asec_current_money.py:82-83, \"JSON_SIZE_OR_TYPE\"). So for the canonical roster the encode-side PAYLOAD_SIZE is arithmetically implied and unreachable. It is reachable only through a producer-declared `field_entities` roster that assigns MORE than 32 of the 33 FIELDS to \"person\": `_entities` (:89-105) validates only the names and that each entity is in {\"person\",\"household\"}, never the 32/1 split the budget's comment at :82 assumes. That makes :461 a real roster/budget mismatch backstop, not dead code \u2014 a point the row misses. On the decode side, :483-487 is genuinely the first ceiling applied to untrusted bytes, before `struct.unpack_from` (:501) and before `_split` slices buffers (:527), so it is not redundant there.\n\n(6) SCALE QUESTION, ANSWERED IN THE ROW'S FAVOUR. The counted roster is the ASEC arm, not the stacked/cloned survey. body rows are pinned to the prepared ASEC receipt's entity_rows (graph_asec_prepared.py:422-432 \"BODY_ROW_ALIGNMENT\"; graph_composed_asec_binding.py:593-604 same code), selected rows <= body rows via `_positions`, and the composed arm's positions must be unique and strictly increasing in the source's own identity (graph_composed_asec_binding.py:322-331). A 1,587,376-household stacked roster or the doubled clone therefore cannot enter this payload, and would in any case be refused earlier by \"BODY_ROW_BOUND\" (:296) / \"SELECTED_ROW_BOUND\" (:434) against MAX_PERSONS=1,000,000, not by \"PAYLOAD_SIZE\".", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and it does not need to move. Three independent reasons, each read at HEAD.\n\n(a) It is not an independent constant. asec_current_money_selection.py:83-89 defines it purely as len(SELECTED_MAGIC) + 4 + HEADER_MAX_BYTES + 11*(32*MAX_PERSONS + MAX_HOUSEHOLDS) + 32, from the symbols imported at :30-33. You cannot raise it without raising MAX_PERSONS (1,000,000) or MAX_HOUSEHOLDS (400,000) at asec_current_money.py:48-49 \u2014 and those are the row assertions applied to the ASEC source roster itself: AsecMoneyScope.__post_init__ refuses `0 < p <= MAX_PERSONS and 0 < h <= MAX_HOUSEHOLDS` with \"SOURCE_SIZE\" at asec_current_money.py:521, the money header is bounded by \"HEADER_ROW_BOUND\" at :800-806, and the body/selected headers by \"BODY_ROW_BOUND\" (:296) and \"SELECTED_ROW_BOUND\" (:434, :521) here. Moving the byte budget therefore means weakening a size assertion on an upstream CPS ASEC roster whose real catalogued size is 55,762 households / 142,125 persons (88,932 households counting the 33,170 unrepresented) \u2014 7x below the person ceiling already.\n\n(b) A reader depends on the exact number. decode_selected_current_money refuses `len(payload) > SELECTED_PAYLOAD_MAX_BYTES` at :483-487 as its first act, before the header length is even unpacked (:501) and before `_split` allocates buffers (:527). The artifact is a versioned, digest-pinned graph type (US_ASEC_SELECTED_MONEY_TYPE = ArtifactType(\"microcosm.us.asec_selected_current_money\", 1), :48-50; the encoded bytes are hashed into receipts at graph_asec_prepared.py:569 and graph_asec_income.py:380). Changing the cap changes the set of payloads an existing decoder accepts, i.e. a producer/consumer compatibility break, even though it would not re-lay-out any byte \u2014 the only fixed-width field in the format is the ` 2.37x headroom. Conclusion unchanged; the number is just not the one the row states.\n Per the task's instruction about ASEC-side rosters: this bound is on an ASEC-source roster, and the real upstream size is the three-cohort pool at 168,852 households / 432,523 persons (measured, above). The 2024-only ASEC catalogue in the same pilot is 55,762 represented + 33,170 unrepresented = 88,932 households / 142,125 persons (preparation.json catalogues.asec) \u2014 that is one cohort's selectable slice, NOT what this artifact encodes.\n Side finding with the same evidence: the lane's own untracked test packages/microcosm-build/tests/test_us_native_row_ceilings.py:38-39 and :106-116 asserts \"person_rows and household_rows are the ASEC source scope's own sizes, and the catalogue measures them at 142,125 persons in 55,762 households\". That is the wrong roster by a factor of ~3. The real scope sizes are 168,852/432,523, so headroom on asec_current_money.MAX_HOUSEHOLDS (400,000) is 2.37x, not 7.2x, and on MAX_PERSONS (1,000,000) 2.31x, not 7.0x. Still non-binding, but the stated rationale should be corrected before it is relied on.\n\n3) WRONG CLASSIFICATION. The row says \"encoding-width ... the decoder relies on before parsing\". At the site it cites, the decoder explicitly does NOT rely on it: decode_housing_universe compares the whole payload to a re-encode of the independently issued `expected` (asec_housing_universe.py:386 \"EXPECTED_CONTENT\") with the comment at :385 \"Compare before parsing any candidate header or buffer; no alternate encoding\", and read_housing_universe never mentions PAYLOAD_MAX_BYTES at all \u2014 it requires `info.st_size == size` where size = len(encode(expected)) (:409, :414 \"ARTIFACT_FILE\") and reads size+1 bytes. (Contrast asec_housing_status.py:473, which really does read PAYLOAD_MAX_BYTES+1.) Only the site the row missed (graph_housing_universe.py:80) is a genuine pre-parse budget, and even there the format is self-describing: magic, a uint32 header length bounded independently by HEADER_MAX_BYTES (:100-104 \"BOUND_HU_HEADER_SIZE\"), a JSON header carrying `household_rows`, then buffers, then a 32-byte digest, with the exact length re-derived from rows at :150-155 \"BOUND_HU_LENGTH\". No field width depends on PAYLOAD_MAX_BYTES; changing it changes no wire format. It is a resource pre-check whose magnitude is a pure derivation of MAX_HOUSEHOLDS x 63 + framing \u2014 hence memory-or-time in character, driven by an upstream-roster bound.", + "safe_to_move": false, + "safe_to_move_reason": "No, and it also needs no move. (a) It is unreachable-by-construction for valid artifacts: rows are capped at MAX_HOUSEHOLDS=400,000 in three places (asec_housing_universe.py:147, :241, :307, all \"ROW_COUNT\"; plus graph_housing_universe.py:145-149 \"BOUND_HU_ROWS\") and the header at HEADER_MAX_BYTES=65,536 (:72 \"HEADER_SIZE\"), so PAYLOAD_MAX_BYTES is exactly the supremum those two already imply \u2014 9+4+65,536+400,000*63+32 = 25,265,581. Any artifact that passes ROW_COUNT necessarily passes PAYLOAD_SIZE/BOUND_HU_SIZE. (b) Editing it independently breaks that equality: lowering it would refuse artifacts the row ceilings accept; raising it alone accomplishes nothing, because rows stay capped at 400,000 and graph_housing_universe.py:155 re-requires the exact length from rows. (c) The only way it legitimately moves is by moving MAX_HOUSEHOLDS, and MAX_HOUSEHOLDS bounds the ASEC source pool's own household count \u2014 an upstream source roster (168,852 measured, 2.37x headroom), not a selected/stacked/cloned roster that grows with the selection fraction. Nothing in a full-source native build pushes it: at 1/1000, 1/10 and full source the artifact is the same ~10.64 MB. This lane's ceiling rule correctly leaves it alone.", + "evidence": "packages/microcosm-build/src/microcosm/build/us_runtime/asec_housing_universe.py:19-27 (RAW_COLUMNS, 7 int64), :29-37 (DERIVED_COLUMNS, 7 uint8), :39-41 (MAGIC=b\"MCAHUNIV\\x01\" len 9, HEADER_MAX_BYTES=65536, MAX_HOUSEHOLDS=400_000), :42-48 (PAYLOAD_MAX_BYTES = 9+4+65536+400000*63+32 = 25,265,581), :52-53 (class HousingUniverseRefusalError(ValueError)), :56-58 (_require raises it), :72 (\"HEADER_SIZE\" bounds header to HEADER_MAX_BYTES), :147/:241/:307 (\"ROW_COUNT\" caps rows at MAX_HOUSEHOLDS), :353-357 (_parts: MAGIC, \" HousingUniverseRefusalError), :100-105 (\"BOUND_HU_HEADER_SIZE\" then _parse), :145-149 (\"BOUND_HU_ROWS\": rows <= hu.MAX_HOUSEHOLDS and rows == prepared_receipt[\"entity_rows\"][\"household\"]), :150-155 (\"BOUND_HU_LENGTH\": len(payload) == cursor + rows*63 + 32).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_housing_universe_source.py:74-87 (raw built from source.scope.household_ids / household_years / household_native_keys), :283-331 (verify_housing_universe_rows maps clone/subset lineage_ids into universe positions \u2014 clones do not enlarge the universe).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_current_money_source.py:251-277 (_scope; requires source years == {2022,2023,2024}), :534-535 (\"the two reviewed full-source files and exact scope\"), :544-560 (byte-pinned checkpoint staging, three cohorts), :584 (scope = _scope(frame)).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_current_money.py:48-49 (MAX_PERSONS=1_000_000, MAX_HOUSEHOLDS=400_000), :521 (\"SOURCE_SIZE\"), :913 (p,h from source_scope), :1021-1022 (money_header person_rows/household_rows).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_prepared_source.py:351-352 and :529 (entity_rows = frame.n(entity) of the prepared ASEC source frame), :449 (load_authenticated_housing_universe).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_housing_status.py:473 (contrast: that module's reader really does read PAYLOAD_MAX_BYTES+1).\n/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/manifest.json \u2014 money_header at /nodes/survey_predictors.asec_current_columns/receipt/money_header (and .attach, .source_projection, plus the /content_addressed/ copies): household_rows 168852, person_rows 432523, source_authentication \"checkpoint_bytes_verified\".\n/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json \u2014 catalogues.asec: households 55762, persons 142125, unrepresented_households 33170 (2024 cohort only, a different roster from the pooled scope above).\npackages/microcosm-build/tests/test_us_native_row_ceilings.py:38-39, :106-116 (the lane's own ASEC_HOUSEHOLDS=55_762 / ASEC_PERSONS=142_125 rationale, contradicted by the measured money_header scope)." + }, + { + "constant": "_asec_current_money_codec.py:26 PAYLOAD_MAX_BYTES = len(MAGIC) + 4 + HEADER_MAX_BYTES + 11 * (32 * MAX_PERSONS + MAX_HOUSEHOLDS) + 32 = 9 + 4 + 65536 + 11*(32*1_000_000 + 400_000) + 32 = 356,465,581", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The census is right on the constant, its value, every enforcement site, the refusal code, the exception type, the derived-knob analysis and the non-binding verdict. It is wrong on one count it asserted, and therefore on its byte figure.\n\n1) household_rows is 168,852, not 169,699. Evidence: the recovered 1/1000 pilot artifact /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/projection.json (same object in manifest.json) carries the real money_header: {\"recipe\":\"us_asec_current_money_ccpiu_2024_v1\", ..., \"person_rows\":432523, \"household_rows\":168852, \"schema_version\":2, \"source_authentication\":\"checkpoint_bytes_verified\"}. I re-canonicalized that header with the module's own rule (json.dumps sort_keys, separators (\",\",\":\"), ensure_ascii \u2014 asec_current_money.py:76-81) and its sha256 reproduces the recorded money_header_sha256 8cec113a39db5ac1c26cf65e1ab4d8c194e8b58784f2468ee555b5eb6a642868 exactly, so these two row counts are authenticated, not inferred. person_rows=432,523 matches the census; household_rows does not. The string \"169699\" appears nowhere in any of the four pilot artifacts and nowhere in the repo except unrelated PUF projection JSON.\n\n2) Consequently the payload size is 154,106,694 bytes (= 9 + 4 + 1181-byte header + 11*(32*432523 + 168852) + 32), not the census's 154,114,785 (it is over by 847 households * 11 bytes = 9,317 bytes). Headroom is 356,465,581 / 154,106,694 = 2.313x, so the census's \"~2.3x headroom\" stands.\n\n3) One framing refinement, not an error: the census calls this encoding-width, and under the task taxonomy that is the right bucket (a HEADER + ROW_BYTES*MAX_N byte budget a reader relies on before parsing, _asec_current_money_codec.py:97-101). But moving it would NOT change the wire format: the layout is validated by the exact-length equality at :122-125 (and the same equality in asec_current_money_selection.py:298-301), which is computed from the header's own person_rows/household_rows. PAYLOAD_MAX_BYTES is strictly a pre-parse bounded-input guard on an arbitrary byte string, derived from the row ceilings; the census says exactly this itself, so its prose is accurate.\n\nEverything else I confirmed independently: MAGIC is 9 bytes (b\"MCAM2024\\x01\", :23); HEADER_MAX_BYTES=64*1024, MAX_PERSONS=1_000_000, MAX_HOUSEHOLDS=400_000 (asec_current_money.py:47-49); 11 bytes/field-row = amount \" encode_current_money(ready)); the scope builder requires set(source_year) == {2022, 2023, 2024} (asec_current_money_source.py:261-262, \"SOURCE_YEAR_MAPPING\"), i.e. exactly three ASEC cohort years; graph_asec_prepared.py:427-431 binds body.person_rows/household_rows to the prepared ASEC frame's entity rows (\"BODY_ROW_ALIGNMENT\"). And the observed 432,523 persons / 168,852 households in a 1/1000 pilot is ~3x the pilot's own recorded single-year ASEC catalogue (asec counts households 55,762, persons 142,125 in preparation.json), while the ACS spine in that same run is 1,631,969 catalogue households / 1,587,376 supplied. So the payload is fixed at ~154.1 MB whether the run is 1/1000, 1/10 or full source. It never approaches 356.5 MB.\n\nNo cloning path can inflate it: the only subsetting API, select_current_money, routes positions through _positions (asec_current_money_selection.py:122-131), which requires strictly increasing int64 positions with 0 < len(values) <= rows and values[-1] < rows. A selection is a strict subset, never a repeat, so SELECTED_PAYLOAD_MAX_BYTES (identical formula, :83) is bounded by the body and also never binds.\n\nTighter upstream limits: yes, and they make this one unreachable for well-formed data. The same rows are bounded earlier by MAX_PERSONS/MAX_HOUSEHOLDS at asec_current_money.py:520-521 (\"SOURCE_SIZE\"), asec_current_money.py:800-805 (\"HEADER_ROW_BOUND\") and asec_current_money_selection.py:291-293 (\"BODY_ROW_BOUND\"). Since PAYLOAD_MAX_BYTES is literally those two ceilings times the row width, any conforming payload that exceeded it must already have failed a row bound. The graph's raw-bytes-v1 cap RAW_BYTES_MAX_BYTES = 64 MiB (packages/microcosm-graph/src/microcosm/graph/codecs.py:78, enforced :571-580) does NOT apply here \u2014 the money body travels as an in-graph ArtifactOutput (\"current_money\", US_ASEC_CURRENT_MONEY_BODY_TYPE; graph_asec_prepared.py:302, :324, :389) rather than as a declared raw-bytes source. Worth flagging: at 154 MB the body already exceeds that 64 MiB codec cap by 2.3x, so it must never be routed through raw-bytes-v1. I found no size limit on graph artifact payloads (ArtifactValue only type-checks bytes, kernel.py:301-309).", + "safe_to_move": false, + "safe_to_move_reason": "No move is needed and a move in isolation would be wrong. (1) It never binds: the encoded body is a fixed ~154.1 MB against a 356.5 MB ceiling (2.31x headroom) at every sample fraction, because it measures the three-year pooled ASEC source, which does not scale with the ACS spine. (2) It is not an independent knob \u2014 it is a derived expression over MAX_PERSONS/MAX_HOUSEHOLDS (asec_current_money.py:47-49). Raising it alone accomplishes nothing because HEADER_ROW_BOUND/BODY_ROW_BOUND/SOURCE_SIZE still cap the same rows; lowering it would be actively unsafe, since anything below ~154.2 MB refuses today's real production body with \"PAYLOAD_SIZE\". (3) The underlying ceilings are themselves a structural statement about an upstream file: 1,000,000 persons / 400,000 households against a real pooled ASEC of 432,523 / 168,852, and the same two numbers are declared in the sha256-pinned resource asec_current_money_domains_v1.json (\"max_persons\":1000000, \"max_households\":400000), pinned by RESOURCE_PINS and checked as \"RESOURCE_FINGERPRINT\" (asec_current_money.py:37-41, 404-411). Note the code never reads those two JSON keys (only domains[\"fields\"], :467-469), so changing the Python constants would silently contradict a digest-pinned declaration. (4) Any edit to _asec_current_money_codec.py at all \u2014 including a comment \u2014 changes the module's sha256, which is hashed live into execution.modules (asec_current_money_resources.py:31-33, 44-52), validated as \"EXECUTION_MODULES\" (asec_current_money.py:421-429) and folded into the spec identity (:479-490) and thus spec_sha256, the money header, money_header_sha256, money_content_sha256 and every downstream receipt pin (the pilot header's spec_sha256 is 4daec9d262c916be8df59032712e5845eaa80c6944fd6cd27d0e0afd04b85440). So touching it forces an attested-identity re-pin for zero benefit. It is not a fixed-width field and moving it would not corrupt a reader \u2014 the exact-length check does that work \u2014 but there is no reason to move it, and it must not be lowered.", + "evidence": "Code at HEAD (2ebd246f1), all under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/packages/microcosm-build/src/microcosm/build/us_runtime/:\n- _asec_current_money_codec.py:23 MAGIC = b\"MCAM2024\\x01\" (len 9); :25-28 PAYLOAD_MAX_BYTES definition (value computes to 356,465,581); :46-55 the 11-byte lanes (amount/status/validity/zero_origin appended per field); :56 _require(sum(map(len, parts)) + 32 <= PAYLOAD_MAX_BYTES, \"PAYLOAD_SIZE\"); :97-101 _require(... len(payload) <= PAYLOAD_MAX_BYTES, \"PAYLOAD_SIZE\") on decode; :122-125 expected_size = start + size + 11*(32*person_rows + household_rows) + 32 with _require(len(payload) == expected_size, \"PAYLOAD_LENGTH\"); :126 \"PAYLOAD_CHECKSUM\".\n- asec_current_money.py:47-49 HEADER_MAX_BYTES = 64*1024, MAX_PERSONS = 1_000_000, MAX_HOUSEHOLDS = 400_000; :59-64 class MoneyRefusalError(ValueError); :68-70 _require raises MoneyRefusalError; :520-521 \"SOURCE_SIZE\"; :636-659 MoneyField (\" subsets only); :461 and :485 the selected-artifact \"PAYLOAD_SIZE\" sites.\n- asec_current_money_source.py:255-277 _scope(...) building AsecMoneyScope, with _require(set(years) == {2022, 2023, 2024}, \"SOURCE_YEAR_MAPPING\") at :261-262.\n- asec_prepared_source.py:407-471 prepare_asec_current_money_population -> money_payload = encode_current_money(ready) over the full authenticated three-cohort ASEC population.\n- graph_asec_prepared.py:302, :324, :389 the money body as an in-graph ArtifactOutput; :422-431 parse_current_money_body + \"BODY_ROW_ALIGNMENT\" tying body rows to the prepared ASEC frame's entity rows.\n- asec_current_money_domains_v1.json: 33 fields (32 person + HTOTVAL household), max_persons 1000000, max_households 400000 (keys not read by code).\n- asec_current_money_resources.py:31-33, 44-52: live sha256 of _asec_current_money_codec.py into the execution identity.\n- packages/microcosm-graph/src/microcosm/graph/codecs.py:78 RAW_BYTES_MAX_BYTES = 64*1024*1024, enforced :571-580 (not on this path); kernel.py:301-309 ArtifactValue has no size cap.\n\nArtifact evidence (read-only): /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/projection.json and manifest.json, money_header with person_rows 432523, household_rows 168852, canonical header length 1181 bytes, sha256 reproducing the recorded money_header_sha256 8cec113a39db5ac1c26cf65e1ab4d8c194e8b58784f2468ee555b5eb6a642868; preparation.json catalogues/asec/counts households 55762, persons 142125 and selection/supplied_households 1587376. Exact payload = 154,106,694 bytes; ceiling/payload = 2.313.\n\nNo test in packages/microcosm-build/tests pins the literal 356465581; grep across the repo finds that number nowhere." + }, + { + "constant": "acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES = 2 * 1024**2 = 2,097,152", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": true, + "correction": "I agree with every substantive verdict (value, three enforcement sites, refusal string AND exception type, counted quantity, binds-at-full, memory-or-time classification, movable). Three corrections, none of which flips anything:\n\n(1) WRONG LINE CITES for the two receipt fields that make :711 row-scale. `\"requested_serialnos\": serialnos` is at binding.py:638, NOT :684. `\"vacant_serialnos\"` is at binding.py:649-651, NOT :697-699. Line 684 is `prepared.source.projection_json` inside `design_anchors`; 697-699 is inside the `\"capture\"` block. The substance is correct -- both fields do exist and do carry the literal key list -- only the cites are off (consistent ~46-line offset, suggesting the census was written against an earlier revision). The three ENFORCEMENT cites (:465, :564, :711) and the two wrapper cites (:542-543, :735-736) are all exactly right at HEAD.\n\n(2) ARITHMETIC, minor. The \"~20-40 kB fixed producer closure\" is measurably 17,610 canonical bytes: I canonicalized `/producer/acs/native_source_closure` from the pilot artifact (preparation 15,889 + coverage 1,469 + issuer_sha256 66 + limits 56 + accepted_age_transform 49). So :711 exhausts at n ~= 129,800 keys (fraction ~1/11.8), not quite the census's implied figure. Also 1/10 is 153,160 ACS keys, not 153,161 (per-cell flooring: 8,442+9,878+166+134,674). Both immaterial: still 1.17x over at :711 and 2.34x over at :564.\n\n(3) THE \"Movable\" VERDICT NEEDS SCOPING, and the census's own note that \"these two caps make a full-source run IMPOSSIBLE regardless of any row-count ceiling\" understates it. At :564 the cap is `min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES)`. ACS_HU_RECEIPT_MAX_BYTES = 1024**2 (acs_housing_universe_source.py:43) is already the smaller of the two, so raising MAX_EVIDENCE_BYTES alone accomplishes NOTHING at :564 -- the 1 MiB cap simply becomes the binding one and the 65,535-household wall stands. And ACS_HU_RECEIPT_MAX_BYTES is not a free-floating ceiling: it is a receipt-file write/read budget with real readers (acs_housing_universe_source.py:204, :232, :616 \"RECEIPT_TOO_LARGE\", :647, :678; graph_acs_housing_universe.py:202, :218, :294, :308, :310), so it must be argued on its own terms and is NOT covered by this row's \"movable\".\n\nAlso worth recording, omitted by the census: MAX_EVIDENCE_BYTES appears a FOURTH time, at binding.py:265, as entry 4 of the `_producer()[\"limits\"]` tuple. That is not an enforcement site but it is an attestation site -- the value is embedded in the receipt at :633 (`\"producer\": producer`), hashed into the payload, and written into build artifacts (pilot `/producer/acs/native_source_closure/limits` = [8589934592, 17179869184, 6000000, 2097152, 1000000, 1000000]). Any edit to this file also moves `issuer_sha256` (:262). So moving the constant carries a downstream digest re-pin cost -- exactly the cost this lane already paid for entries 5 and 6 of that same tuple (1,000,000 -> 7,000,000 / 14,000,000 at acs_pums.py:55-56).\n\nNOT ESTABLISHED: the census's \"MAX_BODY_BYTES refuses at 5,730 households, 11x earlier\". I confirmed the site (SELECTED_BODY_BUDGET, acs_person_coverage_authentication.py:380-382, `selected_budget += 6*len(raw) + 1024` per selected row against MAX_BODY_BYTES = 64 MiB at :37) and that it refuses at far fewer than 65,535 households, but the exact household number depends on the raw ACS CSV row length, which I did not measure. Directionally correct; the precise figure is unverified and belongs to that constant's own row.", + "safe_to_move": true, + "safe_to_move_reason": "Safe to move as a category (a) memory/time ceiling, with a re-pin cost and one scoping caveat.\n\nNOT (b) a fixed-width encoding or a reader-dependent byte budget. I read the encode/decode path. `_json` (acs_person_coverage_authentication.py:69-124) takes `cap` as a plain argument, charges escaped bytes in a pre-pass and again during `iterencode`, and refuses; nothing sizes a buffer or computes HEADER + ROW_BYTES*MAX_N from it. The only consumer of the produced payload is `AuthenticatedACSNativeCoverage.receipt` -> `json.loads(_owned(self).payload)` (binding.py:374-376), which is cap-free. The candidate-bytes path at binding.py:715-720 uses `housing._copy(..., len(payload), exact_size=len(payload))` -- the actual length, never the ceiling. Grep confirms MAX_EVIDENCE_BYTES is referenced only inside acs_native_coverage_binding.py (lines 31, 265, 465, 564, 711); no other module reads it.\n\nNOT (c) an upstream file's true size. It counts `requested_serialnos`, a list the CALLER constructs from `plan.selected` (survey_population_preparation.py:1947-1948, fraction-parameterized at survey_catalogue_selection.py:125 `0 < fraction <= 1`). It scales with the selection fraction. Contrast the genuine (c) bound in the same file: MAX_SOURCE_ROWS = 6,000,000 (binding.py:29), which the lane's own test declares must not move because it asserts the ACS person file's real 3,422,888 records (test_us_native_row_ceilings.py:91-102). The coincidence that at fraction 1 the selected list equals the ACS non-vacant household count does not make it a file-size assertion.\n\nNOT (d) structural/domain.\n\nThe :465 site is correctly described as non-row-scaling: I read `_frame_sha256` (binding.py:439-465) and `details` holds only per-entity index/column dtype reprs, a `dtypes` list of one repr per column, `attrs`, flags, and fixed-size weight descriptors plus a `weight_sha256` digest -- O(entities x columns). `grep attrs acs_housing_universe_source.py` returns nothing, so the prepared frame carries no row-scaling attrs payload.\n\nNo tighter upstream limit masks it. `AcsPumsSource.snapshot_serialnos` (acs_pums.py:188-209) runs at binding.py:557, BEFORE :564, and bounds `len(serialnos) <= MAX_EXACT_HOUSEHOLDS = 7,000,000` (acs_pums.py:55) -- far looser. `_preflight` (binding.py:299, called at :574) and its SELECTED_BODY_BUDGET run AFTER :564, so for n > 65,535 the CANONICAL_SIZE refusal is what the operator actually sees. `survey_population_domains.MAX_HOUSEHOLDS = 100_000` (:19) is enforced at :551 as \"HOUSEHOLD_BATCH\" against a batch whose size is capped at `_BATCH_HOUSEHOLDS <= 10,000` (survey_catalogue_selection.py:129) -- a batch bound, not a total, so it does not mask either.\n\nCAVEATS on \"movable\": (i) moving MAX_EVIDENCE_BYTES alone does not unblock :564, where `min(MAX_EVIDENCE_BYTES, ACS_HU_RECEIPT_MAX_BYTES)` already resolves to the 1 MiB ACS_HU_RECEIPT_MAX_BYTES; that constant is a receipt-file budget with real readers and needs its own argument. (ii) The value is attested in `_producer()[\"limits\"]` (binding.py:265) and hashed into the receipt (:633) and build artifacts, and any edit to the file moves `issuer_sha256` (:262), so downstream producer/receipt digests must be re-pinned. It does NOT move the inventory contract: graph_implementation_inventory.json:136 pins imports plus resource_accesses_sha256 and unbound_uses_sha256, none of which a numeric literal change touches.", + "evidence": "CONSTANT AND SITES (all read at HEAD 2ebd246f1):\n- packages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py:31 -- `MAX_EVIDENCE_BYTES = 2 * 1024**2`\n- acs_native_coverage_binding.py:265 -- 4th entry of `_producer()[\"limits\"]` (attestation, not enforcement; census omitted it)\n- acs_native_coverage_binding.py:439 `def _frame_sha256(frame):` ... :465 `return coverage._sha(coverage._json(details, MAX_EVIDENCE_BYTES))`\n- acs_native_coverage_binding.py:563-565 -- `coverage._json(serialnos, min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES))`\n- acs_native_coverage_binding.py:711 -- `payload = coverage._json(receipt, MAX_EVIDENCE_BYTES)`\n- acs_housing_universe_source.py:43 -- `ACS_HU_RECEIPT_MAX_BYTES = 1024**2`\n\nREFUSAL:\n- acs_person_coverage_authentication.py:57-58 -- `class ACSCoverageAuthenticationError(ValueError)`\n- acs_person_coverage_authentication.py:61-63 -- `def _require(condition, code): raise ACSCoverageAuthenticationError(code)`\n- acs_person_coverage_authentication.py:79 and :120 -- `_require(count <= cap, \"CANONICAL_SIZE\")` (pre-count pass and iterencode pass)\n- acs_native_coverage_binding.py:54-59 -- `class ACSNativeCoverageBindingError(ValueError)` / its `_require`\n- acs_native_coverage_binding.py:540-543 -- `except ACSNativeCoverageBindingError: raise` / `except Exception: raise ACSNativeCoverageBindingError(\"NATIVE_VERIFICATION_REFUSED\") from None`\n- acs_native_coverage_binding.py:733-736 -- same pattern raising `ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\")`\n (ACSCoverageAuthenticationError is not a subclass of ACSNativeCoverageBindingError, so it falls through to `except Exception` -- the census's chain is exactly right.)\n\nCOUNTED QUANTITY:\n- acs_native_coverage_binding.py:638 -- `\"requested_serialnos\": serialnos` (census cited :684 -- WRONG)\n- acs_native_coverage_binding.py:649-651 -- `\"vacant_serialnos\": None if selected_roster is None else sorted(...)` (census cited :697-699 -- WRONG)\n- survey_population_preparation.py:1947-1948 -- `acs_keys = tuple(r.key.native_id for r in plan.selected if r.key.source is domains.Source.ACS)`\n- survey_population_preparation.py:1959-1961 -- `acs_native.issue_acs_native_coverage(root/\"acs\", snapshot_root=snapshots, serialnos=acs_keys)`; serialnos is never None on this path, confirming the census\n- acs_pums.py:188-209 `snapshot_serialnos` -- keys are `len(key) == 13`, prefix 2024HU/2024GQ, so 15 canonical bytes each + separator = 16/key; list total = 16n + 1 per the charge rules at acs_person_coverage_authentication.py:88-101\n- packages/microcosm-build/tests/test_us_native_row_ceilings.py:44-49 -- the repo's own measured constants: ACS_HOUSEHOLDS = 1_531_614, ACS_PERSONS = 3_422_888, STACKED_HOUSEHOLDS = 1_587_376\n- /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json -- selection.cells: acs institutional_gq 84,422 / noninstitutional_gq 98,784 / residual_housing 1,665 / shared_housing 1,346,743 = 1,531,614 ACS eligible; asec 55,697; sum 1,587,311. Selected at 1/1000: 84+98+1+1,346 = 1,529 ACS (matches /native/acs/households = 1,529) + 55 asec = 1,584. catalogues.acs.counts.households 1,631,969 minus vacancies 100,355 = 1,531,614 exactly.\n- same artifact, /producer/acs/native_source_closure = 17,610 canonical bytes; its `limits` = [8589934592, 17179869184, 6000000, 2097152, 1000000, 1000000] (4th entry is this constant)\n\nUPSTREAM / MASKING CHECKS:\n- acs_pums.py:55-56 -- MAX_EXACT_HOUSEHOLDS = 7_000_000, MAX_EXACT_PERSON_ROWS = 14_000_000; :194 `0 < len(serialnos) <= MAX_EXACT_HOUSEHOLDS`, called at binding.py:557 before :564 -- far looser\n- acs_native_coverage_binding.py:299 `def _preflight`, called at :574 (after :564)\n- acs_person_coverage_authentication.py:37 `MAX_BODY_BYTES = 64 * 1024**2`; :380-382 `_require(selected_budget <= MAX_BODY_BYTES, \"SELECTED_BODY_BUDGET\")`; :661 `_require(size <= MAX_BODY_BYTES, \"BODY_SIZE\")`\n- survey_population_domains.py:19 MAX_HOUSEHOLDS = 100_000; :551 `_require(type(rows) is tuple and 0 < len(rows) <= MAX_HOUSEHOLDS, \"HOUSEHOLD_BATCH\")`; survey_catalogue_selection.py:129 `0 < _BATCH_HOUSEHOLDS <= min(10_000, domains.MAX_HOUSEHOLDS)` -- batch bound, not total\n\nMOVABILITY:\n- acs_native_coverage_binding.py:374-376 -- `receipt` property is a bare `json.loads(payload)`; no cap-dependent reader\n- acs_native_coverage_binding.py:715-720 -- candidate copy uses `len(payload)` and `exact_size=len(payload)`\n- acs_native_coverage_binding.py:262 -- `\"issuer_sha256\": coverage._sha(Path(__file__).read_bytes())`\n- graph_implementation_inventory.json:136-143 -- the entry for this module pins only `imports`, `resource_accesses_sha256`, `unbound_uses_sha256`\n- test_us_native_row_ceilings.py:91-102 -- the lane's stated contrast case: MAX_SOURCE_ROWS / MAX_ROWS = 6,000,000 must NOT move because they assert the ACS person file's real 3,422,888 records" + }, + { + "constant": "asec_income_observations.py:45 INCOME_PAYLOAD_MAX_BYTES", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "I confirm the value, the enforcement site, the refusal code and type, and \"binds at full: false\" \u2014 but the census is wrong on the classification, and its \"protects\" text hides the one fact that decides whether this may be touched.\n\n(1) NOT encoding-width. I read the encoder and decoder. Encoder (asec_income_observations.py:354-360): payload = MAGIC + struct.pack(\" 64 bytes/row); :354-360 (encode_income_observations: MAGIC + struct.pack(\" 25,398,017 B = 24.22x over (not 24,505,825 / 23.4x)\n - 1/10 = 158,738 keys -> 2,539,809 B = 2.42x over (not 2,450,577 / 2.34x)\n - 1/100 = 15,874 keys -> 253,985 B, under. Conclusion unchanged.\n - threshold is exactly n <= 65,535 (65,535 -> 1,048,561 B ok; 65,536 -> 1,048,577 B over), i.e. selection fraction 65,535/1,587,376 = 4.13% = 1/24.2, NOT \"above 1/23.4\".\n\n(2) \"STILL NOT THE FIRST REFUSAL\" IS BACKWARDS IN EXECUTION ORDER. binding:563-565 runs at the top of issue_acs_native_coverage, immediately after snapshot_serialnos at :556 and BEFORE housing._pins(), _producer(), housing._capture() and _preflight() at :572. MAX_BODY_BYTES/SELECTED_BODY_BUDGET (acs_person_coverage_authentication.py:379-381) is only reached inside _preflight -> coverage._inventory, i.e. later. So at n >= 65,536 the refusal an operator actually sees IS this one (CANONICAL_SIZE -> NATIVE_ISSUANCE_REFUSED), before a single archive byte is opened. The 5,730-household figure is fine as a *threshold* (I measure ~5,950: mean raw person record 689.3 B over 20k records of psam_pusa.csv, budget 6*689.3+1024 = 5,160 B/person row, 64 MiB / 5,160 = 13,005 person rows / 2.187 p-hh = 5,947 households), but it is not \"first\" in any run that supplies more than 65,535 keys.\n\n(3) BINDS-AT-FULL NEEDS AN EXACT-KEY QUALIFIER. serialnos=None (the whole-source route) charges 4 bytes through coverage._json, so a full-source selection=\"all\" run is NOT blocked by this constant at all. It blocks only a full-source *exact-key* run. The pilot lane is exact-key (preparation.json /selection/selected is a literal list of 1,584 SERIALNOs), so the row's operational conclusion holds for this lane, but \"binds at full source: true\" is unconditional as written and should not be.\n\n(4) ENFORCEMENT LIST IS INCOMPLETE \u2014 it omits the whole reader side. graph_acs_housing_universe.py:202, :218, :308, :310 parse receipts with this cap -> ACSHousingSourceError(\"JSON_SIZE\") (housing:82). graph:293-297 is a DECODER acceptance bound on the declared source_receipt_bytes section of the microcosm.acs_housing_graph_evidence.v2 wire payload -> ACSHousingSourceError(\"GRAPH_ACS_SECTION_SIZE\") (graph:123-124 prefixes every code with \"GRAPH_ACS_\" and delegates to source._require). That is the one site where the constant touches a transport format, so it had to be checked before answering \"movable\".\n\n(5) MOVING IT ALONE BUYS EXACTLY 2x. The same serialnos tuple is embedded verbatim as receipt[\"selection\"][\"requested_serialnos\"] (binding:645-ish, inside the receipt dict) and re-capped at binding:711 by coverage._json(receipt, MAX_EVIDENCE_BYTES) with MAX_EVIDENCE_BYTES = 2*1024**2 (binding:31). So lifting ACS_HU_RECEIPT_MAX_BYTES stops the exact-key path at ~131,000 households instead of 65,535. The row calls :564 \"the tightest row-scaling cap in the family's evidence path\" \u2014 true, but it does not say the next one is only 2x away and carries the same payload.\n\nEverything else checks out and I confirmed it independently: the constant's value (also corroborated by the pilot artifact's /producer/acs/limits list, whose last element is 1048576); RECEIPT_TOO_LARGE at :616, OUTPUT_SIZE at :180 reached from :232 (genuinely best-effort \u2014 the :232 call is inside `with suppress(Exception)` at :230) and from :678; FILE_TOO_LARGE at :137 reached from :647 and renamed to RECONSTRUCTION_SIZE at :659-661; the 2x term in the INSUFFICIENT_DISK demand at :204; the exception types (ACSHousingSourceError at :63-65; ACSCoverageAuthenticationError(\"CANONICAL_SIZE\") at coverage:57-64/:79/:120 surfaced as ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\") at binding:736); and that the receipt carries no per-row data (_select returns five ints as full_counts at :529-535, selected keys only as a digest at :605). I also verified there is NO tighter upstream cap on the serialnos tuple: the only earlier check is snapshot_serialnos (acs_pums.py:188-209) whose MAX_EXACT_HOUSEHOLDS = 7,000,000 (acs_pums.py:55) is 107x looser.\n\nFinally I measured the receipt instead of estimating it. Opening the real pinned archives read-only: csv_hus.zip and csv_pus.zip each hold 3 members (psam_husa/husb.csv + README.pdf; psam_pusa/pusb.csv + README.pdf) = 6 total, header lines 1,570 B / 1,882 B, 241 / 286 columns \u2014 every one of the row's \"MEASURED header widths\" is correct. Reconstructing the receipt dict exactly as :597-615 builds it and canonicalizing with the same json.dumps settings as microcosm.graph.canonical.canonical_json gives 26,056 bytes = 2.48% of the cap, fixed for the sha-pinned archives. The row's \"order 25 kB / about 2.4% used / never binds\" is correct.", + "safe_to_move": true, + "safe_to_move_reason": "Safe to move as a resource ceiling, but it is not a free edit and the census row's bare \"Movable\" understates the cost.\n\nNot category (c): nothing here asserts an upstream file's true size. The receipt is a document this module generates; the serialnos tuple is caller-supplied selection. The real upstream-size assertions live elsewhere (the sha256+size pins in acs_2024_1yr_sources.json enforced by _copy(exact_size=...) -> \"SOURCE_SIZE\" at housing:139, and ZIP_MEMBER_SIZE/_MEMBER_MAX at housing:293) and are untouched by this constant.\n\nNot category (b) in the layout sense, but I had to read the codec to say so. graph_acs_housing_universe.py:257 frames the payload as MAGIC + struct.pack(\" ACSNativeCoverageBindingError(\"UNREVIEWED_PREPARATION\").\n2. The module's bytes are digested into implementation_hash(ACS_HU_STAGE) (graph_implementation.py:431-435 modules[name] = _digest(payload); :528-531), which is recorded as receipt[\"implementation_sha256\"] (housing:596-597) and header[\"implementation_sha256\"] (graph:239). The recovered pilot artifact records exactly that digest under /producer/acs/native_source_closure/preparation/modules, so every downstream pinned receipt/payload digest shifts too. This is the same re-pin chain the lane's last two commits already walked.\n\nOne design note: raising the cap is the wrong lever for the binding:564 problem. The housing receipt already solves it correctly by storing selected_serialnos_sha256 (housing:605) instead of the key list; the issuance receipt stores the list verbatim. Carrying a digest there would remove both the 1 MiB and the 2 MiB row-scaling caps at once rather than trading 65,535 households for 131,000.", + "evidence": "Constant and value: packages/microcosm-build/src/microcosm/build/us_runtime/acs_housing_universe_source.py:43. Corroborated at runtime by /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json -> /producer/acs/limits = [8589934592, 17179869184, 64, 6000000, 100000, 2147483648, 1048576].\n\nRefusal machinery (module acs_housing_universe_source.py, exception ACSHousingSourceError): class at :63-64, _require at :67-69. Sites: :82 \"JSON_SIZE\"; :137 \"FILE_TOO_LARGE\"; :180 \"OUTPUT_SIZE\"; :204 (2 * ACS_HU_RECEIPT_MAX_BYTES inside the INSUFFICIENT_DISK demand at :200-206); :230-232 best-effort failure.json write under `with suppress(Exception)`; :616 \"RECEIPT_TOO_LARGE\"; :647 + :659-661 \"RECONSTRUCTION_SIZE\"; :676-680 receipt.json write -> :180.\n\nReceipt content (no row scaling): :592-615 receipt dict; :529-535 _select returns full_counts as five ints; :605 selected_serialnos_sha256 is a digest; member inventory fields at :366-381 (header_hex :376, header :378, rows :379) are O(ZIP members), bounded by _MEMBER_COUNT_MAX = 64 at :47 / enforced :269.\n\nBinding site: acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES = 2*1024**2; :556 snapshot_serialnos; :563-565 coverage._json(serialnos, min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES)); :572 _preflight (proving the cap runs first); :711 coverage._json(receipt, MAX_EVIDENCE_BYTES) re-capping the same requested_serialnos; :735-736 except Exception -> ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\"); :54-58 class + _require; :37 the _ACCEPTED source pin; :234-238 \"UNREVIEWED_PREPARATION\".\n\nByte accounting: acs_person_coverage_authentication.py:57-59 class ACSCoverageAuthenticationError; :61-64 _require; :70-122 _json, \"CANONICAL_SIZE\" at :79 and :120; list charge 2+(n-1) at :121; str charge 2 + 1/printable at :76-83. MAX_BODY_BYTES = 64*1024**2 at :37, SELECTED_BODY_BUDGET at :379-381 inside _inventory (:334-335, selected_budget is per-role-call).\n\nNo tighter upstream cap on serialnos: acs_pums.py:55 MAX_EXACT_HOUSEHOLDS = 7_000_000, enforced acs_pums.py:188-209 (:194).\n\nReader/wire side: graph_acs_housing_universe.py:123-124 _require prefixes \"GRAPH_ACS_\"; :202, :218, :308, :310 parse with the cap; :257 struct.pack(\" MAX_HOUSEHOLDS at issue and at every validate(), before a payload can be encoded; and upstream AsecMoneyScope.__post_init__ refuses \"SOURCE_SIZE\" (MoneyRefusalError, asec_current_money.py:59-70, raised at :521 with 0 < h <= MAX_HOUSEHOLDS = 400_000 defined at :49), re-raised as HousingStatusRefusalError with the same reason at asec_housing_status_source.py:85-86. asec_current_money.py:800-806 adds \"HEADER_ROW_BOUND\" on the same ceiling. Note asec_housing_status.MAX_HOUSEHOLDS (:53) and asec_current_money.MAX_HOUSEHOLDS (:49) are two independent definitions that happen to share the value 400_000 \u2014 raising one alone is inert.\n\nMinor: the \"PAYLOAD_SIZE\" string literal is at :439, not :438; the _require call spans :436-440 with the PAYLOAD_MAX_BYTES comparison on :438. Exception type and definition site in the census are exactly right.", + "safe_to_move": false, + "safe_to_move_reason": "Do not move it \u2014 not because moving it would break a reader, but because there is no reason to and a real check would be weakened.\n\nNarrow wire-format safety: raising it IS harmless to the format. No field width, offset, or checksum depends on it (only the header length is fixed-width, and it is guarded separately by HEADER_MAX_BYTES at :445); payload length is self-describing and re-checked exactly at :446-449 \"PAYLOAD_LENGTH\" against expected.buffers.\n\nBut: (i) It never binds. The quantity is the pinned pooled ASEC source roster, measured 168,852 households (432,523 persons) in the exact parent.h5 the build consumes, fixed by sha256 pins and independent of the spine fraction. 2.364x headroom minimum at full source, and the code path that enforces it (decode/read) has no caller anywhere in the repo.\n\n(ii) It is derived, so \"moving it\" means moving asec_housing_status.MAX_HOUSEHOLDS = 400_000 (:53). That constant is a structural assertion about the real upstream ASEC cohort files \u2014 three ASEC years at ~56k interviewed households each \u2014 and it is mirrored by asec_current_money.MAX_HOUSEHOLDS (:49) enforced as \"SOURCE_SIZE\" at :521 and \"HEADER_ROW_BOUND\" at :804. Raising it in asec_housing_status alone is inert (the money scope refuses first), and raising both weakens a genuine upstream-size check for zero capability gain. A full-source native build reaches nothing near it.\n\nIf SOURCE_YEARS (asec_prepared_source.py:72) ever grew past roughly seven ASEC cohort years, this ceiling would become the binding one and should be revisited then, deliberately, as an upstream-roster assertion \u2014 not as part of a scale lift.", + "evidence": "All paths relative to /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/packages/microcosm-build/src/microcosm/build/us_runtime/ unless absolute.\n\nConstant and its inputs (verified by importing the module with the worktree .venv: MAGIC len 9, RAW 11, DERIVED 14, ROW_BYTES 102, HEADER_MAX_BYTES 65536, MAX_HOUSEHOLDS 400000, PAYLOAD_MAX_BYTES 40865581):\n- asec_housing_status.py:51 MAGIC = b\"MCAHSTAT\\x01\"\n- asec_housing_status.py:52 HEADER_MAX_BYTES = 65536\n- asec_housing_status.py:53 MAX_HOUSEHOLDS = 400_000\n- asec_housing_status.py:54 ROW_BYTES = 8 * len(RAW_COLUMNS) + len(DERIVED_COLUMNS)\n- asec_housing_status.py:55 PAYLOAD_MAX_BYTES = ...\n\nEnforcement and refusal:\n- asec_housing_status.py:436-440 _require(type(payload) is bytes and len(MAGIC) + 4 + 32 < len(payload) <= PAYLOAD_MAX_BYTES, \"PAYLOAD_SIZE\") [condition :438, literal :439]\n- asec_housing_status.py:473 payload = handle.read(PAYLOAD_MAX_BYTES + 1)\n- asec_housing_status.py:86-87 class HousingStatusRefusalError(ValueError); :90-92 _require raises it\n- asec_housing_status.py:445 \"HEADER_SIZE\"; :446-449 \"PAYLOAD_LENGTH\"; :450 \"PAYLOAD_CHECKSUM\"; :102 \"HEADER_SIZE\" in _parse\n- asec_housing_status.py:174 and :294 \"ROW_COUNT\" (n <= MAX_HOUSEHOLDS, before any encode)\n- asec_housing_status.py:300-304 \"BUFFER_SIZE\" exact buffer lengths; :372 header \"household_rows\": len(raw)\n- asec_housing_status.py:406-410 _parts; :413-425 encode; :427-460 decode; :462-465 write; :468-474 read\n\nUpstream gate on the same roster:\n- asec_current_money.py:49 MAX_HOUSEHOLDS = 400_000; :59-70 MoneyRefusalError + _require; :521 _require(0 < p <= MAX_PERSONS and 0 < h <= MAX_HOUSEHOLDS, \"SOURCE_SIZE\"); :800-806 \"HEADER_ROW_BOUND\"\n- asec_housing_status_source.py:85-86 raise HousingStatusRefusalError(error.reason) from a MoneyRefusalError\n\nWhat the rows are:\n- asec_housing_status_source.py:104-117 output arrays sized from scope.household_ids / household_native_keys; :122-127 \"HOUSEHOLD_ORDER\" ties them to source.frame.table(\"household\"); :133 _stage_verified(cohort_paths[year], pin, staged); :164-168 joins record source_rows / joined_rows / unreferenced_source_rows\n- asec_current_money_source.py:251-277 _scope(frame) \u2192 AsecMoneyScope over the full household table, _require(set(years) == {2022, 2023, 2024}) at :263\n- asec_prepared_source.py:72 SOURCE_YEARS = (2022, 2023, 2024); :73 PARENT_FILENAME = ASEC_RAW_STAGE_CHECKPOINT_FILENAME; :77-79 COHORT_FILENAMES; :411-420 load_authenticated_restored_current_money_source(paths[PARENT_FILENAME], ...); :442-443 load_authenticated_housing_status(source, cohort_paths=...); :445 attach_housing_status\n\nMeasured upstream file (read-only, h5py):\n- /Users/maxghenis/PolicyEngine/_recovered/microcosm-native-sources-20260909/sources/asec/parent.h5 \u2014 sha256 e2f2b7495bfcf1448dfb0acb0a17e93f86a6ab8bef8a70ec0a2981983028cbb5, 653,419,922 bytes; artifact_kind \"populace_us_asec_raw_stage\"; identity.row_counts person 432,523 / household 168,852 / tax_unit 231,007 / spm_unit 176,039 / family 189,401 / marital_unit 345,873; ED_VAL audit rows 146,133 (2022), 144,265 (2023), 142,125 (2024)\n- /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json (sha256 34b362d8...) source_files pins that same \"asec/parent.h5\" hash and size; selection.supplied_households = 1,587,376 at fraction [1, 1000]; catalogues.asec.counts households 55,762 / persons 142,125 / unrepresented 33,170 (one ASEC year, the survey-preparation catalogue \u2014 a different stage from the housing-status roster)\n\nCaller census: grep -rn --include='*.py' \"encode_housing_status|decode_housing_status|read_housing_status|write_housing_status\" over the whole worktree returns only the six definition/self-reference lines in asec_housing_status.py. No 169,699 anywhere in packages/, docs/, tools/, or experiments/native-row-ceilings/." + }, + { + "constant": "graph_asec_income.py:81 ACCOUNTING_HEADER_MAX = 65536", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "I confirm the identifiers, the value, the enforcement site, the refusal string, the exception type, the 28-column roster, and the bottom-line \"does not bind at full source.\" I disagree with three things.\n\n(1) WRONG MECHANISM: \"row-independent\" is false. The header embeds \"rows\" (graph_asec_income.py:449-450) and, per column, an \"offset\" and a \"bytes\" field computed from the row count (graph_asec_income.py:431-443). Its length IS a function of rows \u2014 logarithmically, not linearly. Measured by replicating the module's exact canonicalizer (`_json` = json.dumps sort_keys, separators=(\",\",\":\"), ensure_ascii, asec_current_money.py:77) over the exact ACCOUNTING_COLUMNS/ACCOUNTING_DTYPES tuples (graph_asec_income.py:77-78): the 28-descriptor block is 4,250 bytes at 3,464 rows, 4,305 at 34,714, 4,357 at 200,000, 4,361 at 600,000 (the source._MAX_PERSONS ceiling), and 4,415 even at a hypothetical 3,471,383. Total growth across 1000x in rows: 165 bytes. So the verdict \"cannot bind on row count\" survives, but the stated reason (\"row-independent\") is not the mechanism; the real reason is that row dependence is logarithmic and worth ~0.25% of the 65,536 budget.\n\n(2) UNDERCOUNTED SIZE: \"roughly 3-4 kB\" is too low and omits two terms. The descriptor block alone is 4.25 kB, and the header additionally carries the 715-byte `monetary_basis` document (graph_asec_income.py:477, copied from asec_income_observations.py:500 / asec_income_observations_v1.json:69) and the full ASEC catalogue `source_identity` re-embedded as an escaped JSON string (graph_asec_income.py:472-475, sourced from asec_income_observations.py:484, built at asec_population_catalogue.py:644-678). With source_identity stubbed to empty, the header measures 6,571 bytes at 3,464 rows and 6,684 at 600,000. So the honest floor is ~6.6 kB PLUS the escaped source_identity, whose real size I could not establish (no accounting artifact exists in the recovered pilot run \u2014 financial-artifacts/ holds only graph.json, manifest.json, preparation.json, projection.json, recipient-matrix.bin). The only path by which this header could approach 65,536 is that identity document, which is an ASEC-source-catalogue artifact and does not scale with the survey sample fraction either. It is itself already constrained: it was accepted inside the income header, which is capped at source._HEADER_MAX = 65,536 (asec_income_observations.py:43, enforced at :315 and again at graph_asec_income.py:159-161).\n\n(3) WRONG CLASSIFICATION: \"encoding-width\" is not supported by the code. The `67.1 MB, binds); at 1/100 ~11 MB (does not bind); at full source ~1.1-1.3 GB against the module's own 1.02 GiB figure at survey_population_preparation.py:51-53. Binds at full source.", + "safe_to_move": true, + "safe_to_move_reason": "Safe, but NOT by editing line 61 alone \u2014 and the census's reasoning would have produced exactly that broken edit.\n\nNo reader depends on a byte budget: microcosm/graph/store.py:984 put_bytes and :1002 load_bytes apply no size bound, and graph_survey_population.py:495-... `_artifact` compares payload bytes for equality, never length. `microcosm.graph.codecs.RAW_BYTES_MAX_BYTES = 64 * 1024 * 1024` (codecs.py:78, enforced :571-579) bounds raw *source-file* reads, not artifact payloads, so it is not a coupled reader budget. There is no fixed-width field and no HEADER + ROW_BYTES * MAX_N computation anywhere on this path. Not (b). And the payload measures only selected/stacked rows, never an upstream file's true row count, so not (c).\n\nThe blocking coupling is internal: _bounded_json hard-codes the same 64 MiB as the maximum permissible `limit` (graph_survey_population.py:264), and line 470 feeds PREPARATION_MAX_BYTES straight into it. Two safe shapes:\n (1) Preferred, and it matches precedent already in these two modules: leave PREPARATION_MAX_BYTES at 64 MiB as the per-encode ceiling (so 264/470 and the CONTEXT_BYTES sites 309/582 are untouched \u2014 the context is a few kB and can never bind) and introduce a separate explicit total ceiling for the payload sites 305 and 577. This is exactly what ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES does at graph_survey_population.py:66-67 and what MAX_ROSTER_BYTES = 64 * MAX_SEGMENT_BYTES does at survey_population_preparation.py:58, each with a comment saying the old ceiling stays put as the transport shape.\n (2) If the constant itself is raised, line 264's `64 * 1024**2` must be raised in the same edit or every allocation run dies with TRANSPORT_LIMIT, plus test_us_graph_survey_population.py:682.\n\nEither way the payload ceiling must stay at or under the producer's MAX_ROSTER_BYTES = 4 GiB (survey_population_preparation.py:58), which is where the receipt is actually materialized as one bytes object; ~1.02-1.3 GiB at full source fits.\n\nTwo mandatory follow-ons: editing this module moves the stage implementation hash (graph_implementation.py:431-435 digests whole module bytes into the manifest, :528 implementation_hash), so the authenticated_survey_population_v1 pin must be regenerated. And survey_population_preparation.py:678-699 embeds MAX_PAYLOAD_BYTES/MAX_SEGMENT_BYTES/MAX_ROSTER_BYTES in the producer \"contract\" tuple that lands inside the receipt bytes, so touching those producer constants (option 2's neighbors) changes every preparation digest; option 1 confined to graph_survey_population.py avoids that.", + "evidence": "graph_survey_population.py:61 (PREPARATION_MAX_BYTES = 64 * 1024**2); :66-67 (ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES precedent); :108 (class SurveyPopulationGraphError(ValueError)); :111-113 (_require raises it); :262-280 (_bounded_json, with :264 `_require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\")`, :273 and :275 the per-chunk TRANSPORT_LIMIT); :296-315 (_checked_preparation, :305-306 \"PREPARATION_BYTES\", :309-310 \"CONTEXT_BYTES\"); :378 and :383 (_bounded_json at 4096 for allocation instructions); :470 (new_context = _bounded_json(context, PREPARATION_MAX_BYTES)); :575-584 (allocation kernel ctor, :578 \"PREPARATION_BYTES\", :583 \"CONTEXT_BYTES\"); :855 (def run_authenticated_survey_population); :921 (owner, view = _checked_preparation(preparation)); :926 (instructions = allocation_instructions); :1190-1196 (final _checked_preparation, \"FINAL_PREPARATION_SEAL\").\nsurvey_population_preparation.py:50 (MAX_PAYLOAD_BYTES = 64 * 1024**2); :51-53 (the 1.02 GiB / 6.1% comment); :57-58 (MAX_SEGMENT_BYTES, MAX_ROSTER_BYTES = 64 * MAX_SEGMENT_BYTES); :129-168 (_roster_segments, :162 \"ROSTER_LIMIT\"); :232-254 (_roster_payload, :250 payload = b\"\".join(segments), :252 \"ROSTER_LIMIT\"); :678-699 (producer contract tuple embedding the byte constants); :1143-1151 (_bounded_append, :1149 \"ORIGIN_LIMIT\" against MAX_ROSTER_BYTES); :1154-1320 (_origins); :1257-1285 (per-household dict append); :1291-1307 (per-person row append); :1896 (\"CANDIDATE_LIMIT\"); :1948-1955 (acs_keys/asec_keys from plan.selected); :1960 (issue_acs_native_coverage(serialnos=acs_keys)); :1963 (load_authenticated_asec_2024_native_population(selected_households=asec_keys)); :1970-1974 (source_copies, stack_survey_spines); :1983 (origins = _origins(frame, source_copies, plan.selected, native_receipts)); :2010-2035 (_roster_payload call building the receipt); :2022-2023 (native household/person counts are scalars); :384-396 (_plan_document: selected/excluded/cells); :1795-1840 (AuthenticatedSurveyPopulationPreparation.checked_view returning the same payload bytes).\nsurvey_population_domains.py:18-20 (MAX_MEMBERS 20, MAX_HOUSEHOLDS 100_000, MAX_TOTAL_MEMBERS 1_000_000); :547-563 (classify_households, :551 \"HOUSEHOLD_BATCH\", :563 \"BATCH_MEMBER_BOUND\").\nsurvey_catalogue_selection.py:19-20 (_BATCH_HOUSEHOLDS 10_000, _BATCH_PEOPLE 100_000); :196-203 (batch cut and consume).\nsurvey_observed_age.py:19-20 (MAX_ROWS 14_000_000); :52 (\"ROW_BOUND\").\ngraph_context.py:27 (_ENTITY_FIELDS); :104-120 (_row_identity); :123-160 (encode_us_frame_context).\ngraph_implementation.py:431-435 (module bytes digested into the manifest); :528-532 (implementation_hash).\npackages/microcosm-graph/src/microcosm/graph/store.py:984-1009 (put_bytes / load_bytes, no size bound); packages/microcosm-graph/src/microcosm/graph/codecs.py:78, :571-579 (RAW_BYTES_MAX_BYTES, source-file reads only).\npackages/microcosm-build/tests/test_us_graph_survey_population.py:682 (test passes PREPARATION_MAX_BYTES as the _bounded_json limit).\nIndependent re-measurement of /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json: 1,136,063 B, sha256 34b362d8\u2026, supplied 1,587,376 / selected 1,584 / excluded 65 / cells 5; origins households 1,584, persons 3,464; entities family 1,594, household 1,584, marital_unit 2,774, person 3,464, spm_unit 1,585, tax_unit 2,128; top-level origins 718,800 B, selection 393,395 B, producer 21,135 B, source_files 1,004 B, catalogues 726 B, native 413 B." + }, + { + "constant": "graph_asec_income.py:83 ACCOUNTING_PAYLOAD_MAX_BYTES", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The verdict is confirmed (never binds; must not be moved), but two cells of the row are wrong and one is incomplete.\n\n(1) WRONG \u2014 \"at 1/10: ~39.4 MB (same); fixed\". The ASEC prepared arm is itself household-sampled at the run fraction: graph_asec_prepared.py:454 `fraction=context.params[\"fraction\"]` passed to `sample_frame_households`, whose realized count is floor(fraction * eligible) per stratum (frame_sampling.py:108-126 docstring + validate_sample_fraction at :60-66, fraction in (0,1]); keep/person_positions are derived from the sampled household ids at graph_asec_prepared.py:466-470. The accounting rows are exactly those selected persons (graph_asec_income.py:629-631 person_ids = person.person_id, rows = len(person_ids) at :445). So at 1/10 the payload is ~4.0 MB (43,252 rows -> 13 + 65,536 + 91*43,252 + 32 = 4,001,513 bytes), not ~39.4 MB. The ~39.4 MB figure is the fraction=1 (full-source) value, which the row also states correctly. \"fixed\" is true only in the sense that it does not scale with the 1,587,376-household ACS source; it does scale with the build fraction, capped by the ASEC roster.\n\n(2) INCOMPLETE \u2014 \"inherited from asec_income_observations._MAX_PERSONS\" understates why it can never bind. The hard cap is not _MAX_PERSONS (600,000) but the real, content-pinned upstream ASEC roster. derive_reported_income indexes person_ids into the income-observations roster and requires every position >= 0, unique, and strictly increasing (graph_asec_income.py:385-394, \"ACCOUNTING_SOURCE_POSITIONS\"), so rows <= income rows. Income rows are pinned to the three real Census person archives: asec_income_observations.py:409 `_require(len(positions) == rows, \"INCOME_COHORT_ROWS\")` against `_MEMBER_PINS` (asec_income_observations.py:92-95) built from ASEC_EDUCATION_ASSISTANCE_ARCHIVES rows=146_133 / 144_265 / 142_125 (education_assistance_source.py:117, 139, 161) for pppub23/24/25, each with pinned member_sha256 and size checked at \"INCOME_SOURCE_BYTES\". Total = 432,523. bind_income_observations re-pins it: graph_asec_income.py:173-180 requires `rows == binding[\"rows\"] == prepared_receipt[\"entity_rows\"][\"person\"]` (\"BOUND_INCOME_ROWS\"). Full-source payload = 13 + header(<=65,536) + 91*432,523 + 32 = 39,425,174 bytes vs ceiling 54,665,581 -> 1.387x headroom (row says 1.39x, correct). The strictly-increasing-positions rule also means the combined clone can never double this roster, so the cloned ~6.9M persons are irrelevant here.\n\n(3) Not stated by the row: the encode-side check at :495 is unreachable by construction. ReportedIncomeAccounting is token-gated (:322 \"ACCOUNTING_CONSTRUCTOR\"), and derive already enforces rows <= source._MAX_PERSONS and cursor == rows*91 (:446-448 \"ACCOUNTING_LENGTH\") plus header <= ACCOUNTING_HEADER_MAX (:483 \"ACCOUNTING_HEADER_SIZE\"), which together imply len(payload)+32 <= ACCOUNTING_PAYLOAD_MAX_BYTES exactly. Only the read-side guard (:501-503) is live, and only as a pre-reconstruction size check on untrusted candidate bytes.\n\nEverything else in the row checks out: value 9+4+65536+600000*91+32 = 54,665,581 (verified: len(b\"MCAINACC\\x03\")=9, EVIDENCE_COLUMNS=19, ACCOUNTING_ROW_BYTES=8*9+19=91); enforcement at :495 and :501-503; refusal literal \"ACCOUNTING_SIZE\"; exception type MoneyRefusalError(ValueError) defined at asec_current_money.py:59-65 and raised by _require at :68-70; quantity = bytes of the encoded accounting payload = 13 + header + 91*rows + 32; ASEC-source-bounded, not stacked-survey-bounded; binds at full = false.", + "safe_to_move": false, + "safe_to_move_reason": "Do not move it, on three independent grounds. (a) It is not an independently tunable number: it is a derived expression over len(ACCOUNTING_MAGIC), ACCOUNTING_HEADER_MAX, source._MAX_PERSONS and ACCOUNTING_ROW_BYTES (graph_asec_income.py:83-89). Raising the byte budget alone buys nothing, because rows are separately capped at source._MAX_PERSONS by \"ACCOUNTING_LENGTH\" (:446-448); the only way to move it is to raise _MAX_PERSONS or ACCOUNTING_HEADER_MAX. (b) source._MAX_PERSONS (asec_income_observations.py:44) is a guard over an upstream roster whose true size is fixed and content-pinned: the pooled CPS ASEC person files pppub23/24/25 at 146,133 + 144,265 + 142,125 = 432,523 rows (education_assistance_source.py:117/139/161), enforced per cohort at asec_income_observations.py:409 and re-pinned at graph_asec_income.py:173-180. Raising it loosens a guard on real upstream file sizes for no build reason. (c) It is a reader-facing byte budget: read_reported_income rejects candidate payloads on size before reconstructing (graph_asec_income.py:500-503, \"ACCOUNTING_SIZE\"), so changing it changes what a decoder will admit. And it simply never binds: 39,425,174 bytes at full source against a 54,665,581 ceiling (1.387x headroom), with the row count structurally incapable of exceeding 432,523.", + "evidence": "Constant and derivation: packages/microcosm-build/src/microcosm/build/us_runtime/graph_asec_income.py:62-66 (CONTRIBUTING_FIELDS, 5 fields), :69-76 (EVIDENCE_COLUMNS = 5*3 + 4 = 19), :79 (ACCOUNTING_MAGIC = b\"MCAINACC\\x03\", 9 bytes), :81 (ACCOUNTING_HEADER_MAX = 65536), :82 (ACCOUNTING_ROW_BYTES = 8*9 + 19 = 91), :83-89 (ACCOUNTING_PAYLOAD_MAX_BYTES = 9+4+65536+600000*91+32 = 54,665,581).\nEnforcement sites and refusal code: graph_asec_income.py:495 `_require(len(payload) + 32 <= ACCOUNTING_PAYLOAD_MAX_BYTES, \"ACCOUNTING_SIZE\")` inside encode_reported_income (:487-496); :500-503 `_require(type(payload) is bytes and len(payload) <= ACCOUNTING_PAYLOAD_MAX_BYTES, \"ACCOUNTING_SIZE\")` inside read_reported_income (:498-507).\nException type: asec_current_money.py:59-65 `class MoneyRefusalError(ValueError)`; :68-70 `def _require(...): raise MoneyRefusalError(reason, field)`; imported into graph_asec_income.py:31.\nWire format actually written: graph_asec_income.py:489-493 (payload = MAGIC + struct.pack(\"= 0, unique and strictly increasing); :444-448 rows = len(person_ids), \"ACCOUNTING_LENGTH\"; kernel call site :629-631 (person_ids = context.tables[\"person\"].person_id, income_years = person.source_year); composed path graph_composed_asec_measures.py:455-469 uses the arm's ORIGINAL ASEC person ids, so the combined clone cannot double this roster either.\nUpstream roster size (fixed, real): asec_income_observations.py:44 (_MAX_PERSONS = 600_000), :92-95 (_MEMBER_PINS), :377 \"INCOME_ROWS\", :404-409 (`_capture(...) == pin`, \"INCOME_SOURCE_BYTES\"; `_require(len(positions) == rows, \"INCOME_COHORT_ROWS\")`); education_assistance_source.py:97-163 archives with rows=146_133 (pppub23, income_year 2022), rows=144_265 (pppub24, 2023), rows=142_125 (pppub25, 2024) = 432,523 persons; graph_asec_income.py:173-180 \"BOUND_INCOME_ROWS\" ties the graph artifact to prepared_receipt[\"entity_rows\"][\"person\"]; graph_asec_prepared.py:417-421 \"PREPARED_RECEIPT_ROWS\" and :427-431 \"BODY_ROW_ALIGNMENT\".\nFraction scaling of the ASEC arm: graph_asec_prepared.py:451-470 (sample_frame_households with fraction=context.params[\"fraction\"], keep/person_positions), :1016 (select node params carry fraction/seed); frame_sampling.py:108-126 (floor(fraction * eligible) per stratum, fraction in (0,1]).\nArithmetic check (scratch computation, no files touched): 91*432,523 = 39,359,593; full-source payload 13 + 65,536 + 39,359,593 + 32 = 39,425,174 bytes; 54,665,581 / 39,425,174 = 1.387; 600,000 / 432,523 = 1.387; at 1/10, 91*43,252 -> 4,001,513 bytes.\nPilot artifact consulted for scale context: /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/execution-config.proposed.json:159-168 (single run-level \"fraction\": [1, 1000]). Not established: no recovered pilot artifact records the realized ASEC accounting person_rows, so the 1/10 figure above is computed from the sampling rule, not observed." + }, + { + "constant": "graph_survey_calibration.py:44 MAX_BYTES = 64 * 1024**2 = 67,108,864", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The row is right on value, sites, counted quantity, protects-class and binds-at-full. Two defects.\n\n(1) EXCEPTION TYPE AT :155 IS WRONG (or ambiguous enough to mislead). The row writes: '\"NUMERIC_LIMIT\" at graph_survey_budget.py:155 and \"TRANSPORT_LIMIT\" from graph._bounded_json for the :153 path (SurveyPopulationGraphError)'. The parenthetical reads as covering both. It does not. graph_survey_budget.py:39-41 is `def _require(condition, reason): if not condition: raise ValueError(\"SURVEY_BUDGET_TRANSPORT_\" + reason)` \u2014 a PLAIN ValueError, message \"SURVEY_BUDGET_TRANSPORT_NUMERIC_LIMIT\". SurveyPopulationGraphError (graph_survey_population.py:108, `class SurveyPopulationGraphError(ValueError)`, raised by _require at :112-114) applies ONLY to the _bounded_json TRANSPORT_LIMIT path. The row should read: \"NUMERIC_LIMIT\" raised as plain ValueError(\"SURVEY_BUDGET_TRANSPORT_NUMERIC_LIMIT\") at graph_survey_budget.py:155 via :39-41; \"TRANSPORT_LIMIT\" raised as SurveyPopulationGraphError from graph_survey_population.py:264/272-276 for the :153 path.\n\n(2) THE ROW OMITS THE DECISIVE MOVE-BLOCKER. It says \"No reader depends on the byte figure\", which is true of the decoder, and leaves the impression the constant is a free-standing memory dial. It is not. graph_survey_population.py:264 validates _bounded_json's LIMIT ARGUMENT itself:\n _require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\")\nnumeric.MAX_BYTES is passed as exactly that argument at graph_survey_budget.py:153, and it already sits exactly at that hardcoded cap. Raising MAX_BYTES by any amount therefore makes every call to numeric_survey_budget_payload refuse unconditionally with SurveyPopulationGraphError(\"TRANSPORT_LIMIT\") \u2014 an invalid-limit refusal on a zero-row document, not a payload-size refusal. MAX_BYTES is pinned to a second, separately-hardcoded 64 MiB in another module and cannot be moved in isolation. The row must say this.\n\n(2a) Secondary, same area: MAX_ROWS = MAX_BYTES // 128 (:45) is DERIVED. The row's \"the two are enforced independently\" is true of the enforcement sites (:105 vs :119) but the values are not independent \u2014 moving MAX_BYTES moves MAX_ROWS by construction, so the pair cannot be tuned separately without restructuring the definition.\n\n(3) Numbers, checked rather than accepted. I re-encoded the document with the module's exact encoder settings (json sort_keys=True, separators=(\",\",\":\"), ensure_ascii=False, allow_nan=False \u2014 identical to graph_survey_population.py:265-267) at full-source-width household ids and full-mantissa float64 hex tokens (float.hex() = 20 chars, e.g. '0x1.37bdd2f1a9fbfp+9'): 317,476 cloned rows -> 23,271,279 bytes = 73.30 B/row. So \"~22-24 MB\" is right, and the headline \"~70 B/row\" is a little low but inside the row's own stated 64-75 B/row band. Full source 3,174,752 cloned rows -> ~233 MB, inside \"~220-240 MB\". 233 MB > 67.1 MB, so binds-at-full = true stands.\n\n(4) The upstream-ordering reasoning stands, but the row understates how far it is from binding first. Tightest on this path is survey_origin_budget.MAX_PAYLOAD_BYTES (survey_origin_budget.py:52, 64 MiB), enforced on the INPUT budget payload at graph_survey_budget.py:49-52 (_document, code \"PAYLOAD\", SurveyOriginBudgetError is not the raiser here \u2014 graph_survey_budget's plain ValueError is) and at survey_origin_budget.py:86. This lane measured that document at 764 B/group = 87,838 selected households = 175,676 cloned rows (docs/us-native-row-ceilings.md section 5, experiments/native-row-ceilings/origin_budget_size.py). Then graph_survey_age_artifact.py:33+56-59 at 441,074 cloned rows ((67,108,864 - 9 - 4 - 65,536) // 152, with _WIDTH = 19 from the 18 bands asserted at national_age_activation.py:309). Then MAX_ROWS at 524,288 (~38.4 MB at the measured 73.3 B/row). Then this byte ceiling, first reachable at ~915,000 cloned rows. The row's \"two ~90k-household ceilings upstream\" is underspecified \u2014 it names neither, and one of the pair (the 96,860-household preparation-receipt ceiling) the lane doc describes as already lifted by the transport lane while graph_survey_population.py:61 PREPARATION_MAX_BYTES is still 64 MiB at this head; I did not establish which artifact that 96,860 figure belongs to. The conclusion is unaffected: MAX_PAYLOAD_BYTES alone, at 87,838 households, refuses long before anything in graph_survey_calibration.\n\n(5) Not corrections, confirmed as written: the counted quantity is the CLONED roster (graph_survey_budget.py:82 `len(ids) == 2 * len(records)`; records are the `origins` = one allocation instruction per selected household), plus one group_upper_hex per selected household (:140 appends once per record; decode bounds it at :127 \"GROUP_COUNT\" with `0 < len(raw_upper) <= len(ids)`). This is the selected/stacked-then-cloned survey roster, NOT an upstream ASEC (~55,762 hh per the lane catalogue) or ACS file roster. And it is not a fixed-width encoding: decode does json.loads (:107) then `canonical_json(document) == payload` (:111) and derives nothing from MAX_BYTES.", + "safe_to_move": false, + "safe_to_move_reason": "NOT safe to move upward, but for a reason outside the task's (a)/(b) test \u2014 I want that distinction on the record.\n\nIt is NOT (b): this bounds the build's own cloned roster document, not a real CPS/ACS/PUF file's row count, so moving it weakens no structural assertion about upstream data.\n\nIt is NOT (a) in the classic sense: decode_numeric_survey_bounds (graph_survey_calibration.py:104-141) re-parses ordinary canonical JSON \u2014 json.loads at :107, canonical_json(document) == payload at :111 \u2014 and no reader computes a HEADER + ROW_BYTES * MAX_N budget from MAX_BYTES. Contrast graph_survey_age_artifact.py:56-59, which genuinely does derive its byte budget from a row count.\n\nIt is blocked by a third thing the row missed: graph_survey_population.py:264 validates _bounded_json's LIMIT ARGUMENT against a separately hardcoded `64 * 1024**2`, and graph_survey_budget.py:153 passes numeric.MAX_BYTES as precisely that argument. MAX_BYTES already equals that cap. So MAX_BYTES = 64 MiB + 1 does not buy a larger document; it makes numeric_survey_budget_payload raise SurveyPopulationGraphError(\"TRANSPORT_LIMIT\") on every call, including a one-row one, because the limit is rejected before any encoding happens. Any move requires editing graph_survey_population._bounded_json's own cap in the same change, which is a change to the shared transport primitive used by survey_origin_budget.py:86, graph_survey_population.py:378/383/470, graph_puf_diagnostic_consumer.py:403, current_survey_amounts.py:362 and current_survey_predictors.py:508 \u2014 a much wider blast radius than one module's dial.\n\nTwo further reasons not to move it here. (i) MAX_ROWS = MAX_BYTES // 128 (:45) is derived, so the move is not a byte-only move. (ii) It is dead weight even if moved: survey_origin_budget.MAX_PAYLOAD_BYTES refuses at 87,838 selected households (175,676 cloned rows), graph_survey_age_artifact at 441,074 cloned rows and MAX_ROWS at 524,288, all before this ceiling's first reachable ~915,000 cloned rows. And this lane's own written rule (docs/us-native-row-ceilings.md section 1a and section 5) assigns byte transports to the segmented-stream argument rather than to the 4x-rounded row rule, so raising it here would contradict the neighbouring lane.", + "evidence": "VALUE AND DERIVATION\npackages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_calibration.py:44 \u2014 MAX_BYTES = 64 * 1024**2\ngraph_survey_calibration.py:45 \u2014 MAX_ROWS = MAX_BYTES // 128 (= 524,288; derived, not independent)\n\nENFORCEMENT AND REFUSAL CODES\ngraph_survey_calibration.py:61-63 \u2014 def _require(condition, reason): raise ValueError(\"SURVEY_CALIBRATION_\" + reason) [plain ValueError]\ngraph_survey_calibration.py:105 \u2014 _require(type(payload) is bytes and 0 < len(payload) <= MAX_BYTES, \"PAYLOAD\")\ngraph_survey_calibration.py:119 \u2014 _require(type(ids) is list and 0 < len(ids) <= MAX_ROWS, \"ROW_COUNT\")\ngraph_survey_calibration.py:127 \u2014 _require(type(raw_upper) is list and 0 < len(raw_upper) <= len(ids), \"GROUP_COUNT\")\ngraph_survey_budget.py:39-41 \u2014 def _require(condition, reason): raise ValueError(\"SURVEY_BUDGET_TRANSPORT_\" + reason) [plain ValueError \u2014 NOT SurveyPopulationGraphError]\ngraph_survey_budget.py:142-154 \u2014 output = graph._bounded_json({...}, numeric.MAX_BYTES)\ngraph_survey_budget.py:155 \u2014 _require(len(output) <= numeric.MAX_BYTES, \"NUMERIC_LIMIT\")\ngraph_survey_population.py:108 \u2014 class SurveyPopulationGraphError(ValueError)\ngraph_survey_population.py:112-114 \u2014 def _require(condition, code): raise SurveyPopulationGraphError(code)\ngraph_survey_population.py:264 \u2014 _require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\") <-- the move-blocker\ngraph_survey_population.py:272-276 \u2014 running per-piece \"TRANSPORT_LIMIT\" checks\n\nWHAT IS COUNTED (cloned roster, not an upstream file)\ngraph_survey_budget.py:72-74 \u2014 ids = document[\"household_ids\"], groups = ..., records = document[\"origins\"]\ngraph_survey_budget.py:77 \u2014 0 < len(ids) <= numeric.MAX_ROWS (code \"COMPLETE_CLONE_SHAPE\")\ngraph_survey_budget.py:80-82 \u2014 0 < len(records) <= budgets.MAX_GROUPS ... and len(ids) == 2 * len(records)\ngraph_survey_budget.py:131-136 \u2014 incoming[position] / row_upper[position] written once per cloned row\ngraph_survey_budget.py:140 \u2014 upper.append(record[\"upper_float64_hex\"]) (one per origin = per selected household)\ngraph_survey_budget.py:146 \u2014 \"population\": clone.COMBINED_CLONE_NODE\ngraph_survey_calibration.py:351-355 \u2014 bounds.population == counts.population == context.node.base and ordered-id identity\n\nNOT A FIXED-WIDTH ENCODING\ngraph_survey_calibration.py:107 \u2014 json.loads(payload)\ngraph_survey_calibration.py:111 \u2014 _require(canonical_json(document) == payload, \"CANONICAL\")\n(contrast, a real byte-from-rows budget: graph_survey_age_artifact.py:56-59 _shape ... len(MAGIC)+4+MAX_HEADER_BYTES+rows*_WIDTH*8 <= MAX_BYTES, \"LIMIT\")\n\nTIGHTER UPSTREAM CEILINGS\nsurvey_origin_budget.py:52 \u2014 MAX_PAYLOAD_BYTES = 64 * 1024**2\nsurvey_origin_budget.py:62 \u2014 MAX_GROUPS = 7_000_000\nsurvey_origin_budget.py:86 \u2014 def _json(value): return graph._bounded_json(value, MAX_PAYLOAD_BYTES)\ngraph_survey_budget.py:49-52 \u2014 _require(type(payload) is bytes and 0 < len(payload) <= budgets.MAX_PAYLOAD_BYTES, \"PAYLOAD\")\ndocs/us-native-row-ceilings.md section 5 \u2014 measured 764 bytes/group => 87,838 households = 175,676 cloned rows (receipt: experiments/native-row-ceilings/origin_budget_size.py)\ngraph_survey_age_artifact.py:33-37 \u2014 MAX_BYTES = 64*1024**2, MAX_HEADER_BYTES = 65_536, _WIDTH = 1 + len(_COLUMNS)\nnational_age_activation.py:309 \u2014 _require(len(self.bands) == 18, ...) => _WIDTH = 19 => (67,108,864-9-4-65,536)//152 = 441,074 cloned rows\n\nBYTES/ROW, MEASURED NOT EXTRAPOLATED\nRe-encoded a faithful bounds document with graph_survey_population.py:265-267's exact encoder settings (sort_keys=True, separators=(\",\",\":\"), ensure_ascii=False, allow_nan=False), full-source-width ids and full-mantissa float64 hex tokens: 317,476 cloned rows -> 23,271,279 bytes = 73.30 B/row; scaling to 3,174,752 cloned rows -> ~233 MB vs the 67,108,864-byte cap.\n\nHEAD: cb5a62cf05f92040a2e0ee03cfcb2406ea5c8d76 (branch native-row-ceilings). No file was edited." + }, + { + "constant": "survey_atomic_geography.py:63 MAX_SUPPORT_BYTES (= RAW_BYTES_MAX_BYTES = 64 * 1024 * 1024 = 67,108,864)", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The census's facts are right; its CLASSIFICATION is wrong, plus one line-number slip.\n\n1) WRONG CATEGORY (substantive). Census says \"protects: upstream-file-size \u2014 a structural assertion about a real on-disk publisher-derived input\". The code says the opposite. The size ceiling asserts nothing about the real file; the sha256 does, exactly, at survey_atomic_geography.py:158-163 (`_sha(payload) == config.support_sha256`, refusal \"SUPPORT_CHANGED\"). The 64 MiB number is a bounded-read resource ceiling, documented as such at its definition: codecs.py:79 \"\"\"The most ``raw-bytes-v1`` will read from one source file (64 MiB).\"\"\" and codecs.py:541-547 \"\"\"The read is bounded rather than trusted to ``st_size``, because a file may grow after it is measured. Detecting that a source moved is the executor's own content check ... this bound only keeps the codec from reading an unbounded amount first.\"\"\" Its siblings in the same module are unmistakably anti-zip-bomb memory bounds: MAX_SUPPORT_EXPANDED_BYTES = 2 GiB and MAX_SUPPORT_MEMBERS = 32 (lines 64-65), enforced at :172-180 under the comment \":169-170 Bound declared decompressed bytes before numpy opens any array\", and the per-member NPY header check at :191-197 (\"SUPPORT_ARRAY_SIZE\"). Correct category: (a) memory-or-time. This matters: filed as (c), the row reads as \"must not be moved, ever, it is a real-file check\", when the true reason not to touch it is blast radius plus zero benefit.\n\n2) LINE-NUMBER SLIP. The `_require` helper is at survey_atomic_geography.py:68-70 (`def _require` :68, `raise ValueError(\"ATOMIC_SURVEY_RECONSTRUCTION_\" + reason)` :70), not :67-69. The exception type (plain ValueError) and the literal \"SUPPORT_SIZE\" are exactly right, and the enforcement span :151-154 is exactly right.\n\n3) INCOMPLETE ON QUESTION 3 (tighter bound). The census does not mention that the same 64 MiB is enforced on this very artifact a second time, on the producer side: atomic_block_sources.py:365 `_require(len(payload) <= RAW_BYTES_MAX_BYTES, \"SUPPORT_BYTES\")` guards the output of `normalized.assemble_atomic_block_support(...)` \u2014 i.e. the national-atomic-support payload itself, not some unrelated raw-bytes source (atomic_block_api_sources.py:264 is the API variant). It also misses that the sibling expanded ceiling is effectively tighter for this artifact: the real file's members expand to 1,061,673,840 bytes (49.4% of MAX_SUPPORT_EXPANDED_BYTES), a 36.78x ratio, so at that ratio the 2 GiB expanded ceiling would refuse at ~55.7 MiB compressed, before the 64 MiB ceiling under test is ever reached.\n\nEverything else in the row checks out, verified independently: value 67,108,864 (codecs.py:78); counted quantity is the byte length of the file at AtomicSurveyReconstruction.support_path, which is 2020-census-block national geography (atomic_block_support.py:21-22 SYSTEM=\"us_census_block_2020\", SOURCE=\"us_atomic_block_support\"; source_ids forced to (\"district\",\"population\",\"puma\") at :112-118) and carries no survey roster; measured file 28,862,508 bytes, sha256 5edc0e77471ba31d550a1eed416d5b46ada0a35425718eb87cfabe4d66fe4960 (matches the pin in experiments/us-atomic-native-national-1-control-20260909.json:3356 and the three harnesses), = 43.0% of the ceiling; its members are area/county/district/population/puma/state/tract with ~5.77M block rows (area.npy 346,196,648 bytes at U15 = 60 B/row), invariant to the survey fraction. So binds_at_full = false is correct and robustly so.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe to move, but for different reasons than the census gave. It is NOT a fixed-width encoding and NOT a structural claim about an upstream file's true size (the sha256 at :158-163 is that claim), so rules (a)/(b) of the \"not safe\" test do not apply \u2014 it is a movable resource ceiling in principle. It is nonetheless wrong to move here, on three read-from-code grounds. (i) No benefit: nothing in this lane is near it (43.0% at 28,862,508 bytes) and the quantity is national block geography, which does not scale with the survey fraction, so a row-ceiling lift gains nothing. (ii) A shared-constant move has real collateral: RAW_BYTES_MAX_BYTES also caps every raw-bytes source read (codecs.py:571-580) and, crucially, backs a genuinely survey-scaled bound at puf55_survey_recipients.py:422 \u2014 `len(ids) * (1 + len(profile.predictors)) * 8 <= RAW_BYTES_MAX_BYTES`, \"MATRIX_SIZE\", where `ids` are selected tax_unit ids \u2014 so relaxing the shared constant would silently relax a real row ceiling that must be argued on its own. It also reds the explicit pin at packages/microcosm-build/tests/test_us_atomic_block_sources.py:243 (`source.RAW_BYTES_MAX_BYTES == RAW_BYTES_MAX_BYTES == 64 * 1024**2`). (iii) A local-only move (redefining MAX_SUPPORT_BYTES away from the alias) buys nothing, because a larger support payload would still be refused at production by atomic_block_sources.py:365 (\"SUPPORT_BYTES\") / atomic_block_api_sources.py:264, and the expanded-bytes sibling would bind first anyway. Any edit also moves the attested producer contract: survey_origin_budget.py:200-206 folds geography.MAX_SUPPORT_BYTES and geography.RAW_BYTES_MAX_BYTES into `_live()[\"geography_contract\"]`, and `_code_bytes()`/`_producer()` (survey_origin_budget.py:253-258, snapshots at :1350-1351) hash the module file, so the recorded producer identity changes and would need re-pinning.", + "evidence": "packages/microcosm-graph/src/microcosm/graph/codecs.py:78 (RAW_BYTES_MAX_BYTES = 64 * 1024 * 1024); codecs.py:79 (docstring: \"The most ``raw-bytes-v1`` will read from one source file (64 MiB)\"); codecs.py:531-547 (load_raw_bytes docstring: bounded read, \"this bound only keeps the codec from reading an unbounded amount first\"); codecs.py:571-580 (read + refusal \"larger than the ...-byte raw-bytes-v1 limit\"); codecs.py:587 (register_bytes(\"raw-bytes-v1\", load_raw_bytes)).\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_atomic_geography.py:37 (import RAW_BYTES_MAX_BYTES); :63 (MAX_SUPPORT_BYTES = RAW_BYTES_MAX_BYTES); :64-65 (MAX_SUPPORT_EXPANDED_BYTES = 2 * 1024**3, MAX_SUPPORT_MEMBERS = 32); :68-70 (`def _require` -> `raise ValueError(\"ATOMIC_SURVEY_RECONSTRUCTION_\" + reason)`); :81 (support_path field); :97-101 (SUPPORT_PATH); :104-118 (SOURCE_IDENTITIES forced to district/population/puma); :143-149 (O_RDONLY|O_NOFOLLOW|O_NONBLOCK open); :151-154 (the enforcement: `0 < before.st_size <= MAX_SUPPORT_BYTES`, \"SUPPORT_SIZE\"); :155-165 (read, before/after/current stat identity, `_sha(payload) == config.support_sha256`, \"SUPPORT_CHANGED\"); :168-180 (ZipFile members, MAX_SUPPORT_MEMBERS + MAX_SUPPORT_EXPANDED_BYTES, \"SUPPORT_EXPANDED_SIZE\"); :181-197 (NPY header vs member.file_size, \"SUPPORT_ARRAY_SIZE\", \"SUPPORT_NPY_VERSION\"); :455 (SourceRef(blocks.SOURCE, \"raw-bytes-v1\")); :580, :634 (sources == ((blocks.SOURCE, config.support_path),)).\npackages/microcosm-build/src/microcosm/build/us_runtime/atomic_block_support.py:21-22 (SYSTEM = \"us_census_block_2020\", SOURCE = \"us_atomic_block_support\").\npackages/microcosm-build/src/microcosm/build/us_runtime/atomic_block_sources.py:22, :359-365 (assemble_atomic_block_support -> `_require(len(payload) <= RAW_BYTES_MAX_BYTES, \"SUPPORT_BYTES\")`); atomic_block_api_sources.py:16, :264 (same).\npackages/microcosm-build/src/microcosm/build/us_runtime/puf55_survey_recipients.py:22, :420-431 (\"MATRIX_SIZE\" byte budget over selected tax_unit ids).\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py:200-206 (geography_contract includes MAX_SUPPORT_BYTES and RAW_BYTES_MAX_BYTES); :253-258 (_code_bytes/_producer); :1350-1351 (_BYTES/_LIVE import-time snapshots).\npackages/microcosm-build/src/microcosm/build/atomic_geography.py:55-70 (encode_atomic_support), :73-135 (decode_atomic_support: metadata version/system/level/code_system/vintage/columns; no survey roster).\npackages/microcosm-build/tests/test_us_atomic_block_sources.py:243 (pins RAW_BYTES_MAX_BYTES == 64 * 1024**2); tests/test_graph_codecs.py:221, tests/test_us_atomic_block_api_sources.py:305 (monkeypatch the constant).\npackages/microcosm-fit/src/microcosm/fit/model_input.py:18, :86, :105 (recipient-matrix reader bounds its header with MATRIX_HEADER_MAX_BYTES, not RAW_BYTES_MAX_BYTES \u2014 no reader-side wire dependence on the constant).\nMeasured on disk (read-only): /Users/maxghenis/PolicyEngine/_recovered/microcosm-native-sources-20260909/national-atomic-support.npz = 28,862,508 bytes, sha256 5edc0e77471ba31d550a1eed416d5b46ada0a35425718eb87cfabe4d66fe4960 (identical copies under _recovered/pilot-runs/native45-v5, native19-v2, scratch-backup/893/pilot19, pilot20); members area/county/district/metadata_json/population/puma/state/tract, expanded total 1,061,673,840 bytes (49.4% of 2 GiB, 36.78x ratio); pin echoed at experiments/us-atomic-native-national-1-control-20260909.json:3356." + }, + { + "constant": "graph_survey_age_artifact.py:34 MAX_HEADER_BYTES = 65_536", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "Agrees on every verdict-bearing field (value, enforcement sites, refusal codes, exception type, counted quantity, binds-at-full=false, encoding-width, the 431-row cost and the 441,074 effective ceiling). Three factual errors:\n\n(1) THE COUNTS ARE WRONG. The census says \"~600 B\" at 1/10 and \"~610 B\" at full source. I computed the canonical JSON exactly with json.dumps(sort_keys=True, separators=(\",\",\":\")) over the six fields of the header built at graph_survey_age_artifact.py:188-195, using the 18 column names generated by national_age_activation.py:204-230. Size = 502 + len(population) + digits(rows) + digits(people). The population is a population version / node id string (survey_age_calibration.py:103 passes population=population.version; survey_age_calibration.py:66 COUNT_NODE = \"survey.age_count_matrix\", 23 chars). With a 23-char population: 533 B at 1/1000 (rows 1,584 / people 3,464), 537 B at 1/10 (158,738 / 347,138), 539 B at full source (1,587,376 / 3,471,383) and 539 B for the combined clone (3,174,752 / 6,942,766). With the 8-char test population \"invented\" (tests/test_us_survey_age_artifact.py:30): 518 / 522 / 524 B. Not ~600/~610.\n\n(2) \"OVER 100x HEADROOM\" IS NOT THE WORST CASE. The structural maximum canonical header is 774 B, because graph_survey_age_artifact.py:67-70 caps population at 256 chars (\"POPULATION\"), :56 caps people at MAX_PEOPLE = 10,000,000 (\"SHAPE\") with rows <= people, :71 fixes columns to the 18-name list (\"COLUMNS\"), and :64-66 / :72 fix protocol and age_convention to literals. 65536/774 = 84.7x, not >100x. (Realistic headers do give ~121x.)\n\n(3) THE ENCODE-SIDE SITE IS DEAD CODE. The census lists :204 as an enforcement site without qualification. _validate(header, values) at :202 has already run all of the caps in (2), so len(raw_header) <= 774 is guaranteed before :204 tests it against 65,536; SURVEY_AGE_ARTIFACT_HEADER_LIMIT can never fire there. The decode-side check at :105-108 IS live, because `length` at :104 is read from caller-supplied bytes. Line nit: the literal \"HEADER_LIMIT\" is on :107, not :106 as the census states; :106 is the condition line. The \"LIMIT\" literal on :58 is correct as stated, and :204 is correct as stated.\n\nOne addition the census omits: this bound is reader-enforced twice over. decode_survey_age_counts calls _shape at :117, so the :58 byte budget (len(MAGIC) + 4 + MAX_HEADER_BYTES + rows*_WIDTH*8 <= MAX_BYTES) is evaluated by the DECODER, not only the encoder. Changing MAX_HEADER_BYTES therefore changes the row ceiling a reader will accept, in addition to changing the declared header length a reader will accept at :105-108.", + "safe_to_move": false, + "safe_to_move_reason": "NOT SAFE. Category (b), a fixed-width/byte-budget encoding a reader relies on, on three independent grounds read from the code:\n\n1. It is on both sides of the wire contract. The encoder produces MAGIC + len(raw_header).to_bytes(4,\"big\") + raw_header + values.tobytes(order=\"C\") at graph_survey_age_artifact.py:205-210; the decoder refuses any declared header length above MAX_HEADER_BYTES at :105-108 before slicing at :110. Raising it lets a new encoder emit a payload an older decoder refuses; lowering it below 774 B (the structural maximum from :67-70 and :56) could refuse a legitimate payload with a long population string. These artifacts are persisted to a content-addressed on-disk object store (packages/microcosm-graph/src/microcosm/graph/store.py:1-11) and decoded later by a separate kernel (graph_survey_calibration.py:350), so cross-version reads are real, not hypothetical.\n\n2. It is charged into a reader-enforced byte budget. decode_survey_age_counts calls _shape(header[\"rows\"], header[\"people\"]) at :117, which runs the :57-59 LIMIT check. So MAX_HEADER_BYTES directly sets the maximum row count a DECODER will accept: 441,074 today, (67,108,864 - 9 - 4 - 65,536) // 152 with _WIDTH = 19. Any change silently retunes which previously-valid payloads still decode.\n\n3. It buys nothing anyway. Freeing the entire 65,536 B reservation gains 431 rows out of 441,074 (0.098%), while a full-source roster needs ~1,587,376 stacked or ~3,174,752 cloned households. The constant that actually binds at full source is MAX_BYTES = 64 * 1024**2 at :33: at full source run() passes :168 PEOPLE_LIMIT (3.47M or 6.94M <= MAX_PEOPLE = 10,000,000) and :56 SHAPE, then fails :58 with ValueError(\"SURVEY_AGE_ARTIFACT_LIMIT\") from _shape(len(ids), len(person)) at :173. Per the lane's own rule in docs/us-native-row-ceilings.md \u00a71a and \u00a75, a byte transport takes the transport lane's segmented-stream argument, not the four-times-full-source row rule. This is also why the lane's own lift commit 14defbfc0 left survey_origin_budget.MAX_PAYLOAD_BYTES alone; the same reasoning applies here.\n\n4. Secondary: editing this file at all moves SurveyAgeCountArtifactKernel.implementation_hash() (:147-153), because source_hash digests the defining module's raw file bytes (packages/microcosm-graph/src/microcosm/graph/kernel.py:457-497). That hash is the \"age_artifact\" component of SurveyAgeCalibrationKernel.implementation_hash() at graph_survey_calibration.py:310, so node keys and store addresses move.", + "evidence": "Constant and wire format: packages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_age_artifact.py:32 (MAGIC = b\"MCUSAGE1\\n\", 9 bytes), :33 (MAX_BYTES = 64 * 1024**2 = 67,108,864), :34 (MAX_HEADER_BYTES = 65_536), :35 (MAX_PEOPLE = 10_000_000), :36-37 (_COLUMNS from NATIONAL_AGE_ACTIVATION.bands, _WIDTH = 1 + len(_COLUMNS) = 19), :38-40 (_FIELDS: protocol, population, columns, rows, people, age_convention).\n\nRefusal machinery: :43-45 def _require(condition, reason): raise ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason) \u2014 plain ValueError, no custom subclass.\n\nThe three MAX_HEADER_BYTES sites (grep-confirmed as the only three in the module): :58 inside _require(..., \"LIMIT\") at :57-59; :106 inside _require(..., \"HEADER_LIMIT\") at :105-108 with the literal on :107; :204 _require(len(raw_header) <= MAX_HEADER_BYTES, \"HEADER_LIMIT\").\n\nOther bounds that make :204 unreachable: :56 _require(0 < rows <= people <= MAX_PEOPLE, \"SHAPE\"); :67-70 population is str with 0 < len <= 256, \"POPULATION\"; :71 columns == list(_COLUMNS), \"COLUMNS\"; :64-66 protocol literal, \"PROTOCOL\"; :72 age_convention == ages.SUPPORTED_AGE_CONVENTION, \"CONVENTION\"; :202 _validate(header, values) runs before :203-204.\n\nDecoder relies on the byte budget: :98-123, specifically :104 (length = int.from_bytes(payload[offset:offset+4], \"big\")), :105-108 (HEADER_LIMIT), :116 (CANONICAL_HEADER), :117 (_shape -> the :58 LIMIT check), :119 (BYTE_COUNT).\n\nEncoder: :168 PEOPLE_LIMIT, :173 _shape(len(ids), len(person)), :188-195 header construction, :196-201 int64 matrix, :205-210 payload assembly, :211 _require(len(payload) <= MAX_BYTES, \"LIMIT\") (a second \"LIMIT\" site, on MAX_BYTES not MAX_HEADER_BYTES).\n\n18 columns: packages/microcosm-build/src/microcosm/build/us_runtime/national_age_activation.py:204-230 \u2014 _bands() builds quinary = [(low, low+4) for low in range(0, 85, 5)] (17 bands, columns people_age_0_4 .. people_age_80_84) plus AgeBand(..., column=\"people_age_85_plus\") at :227. Age convention literal: graph_national_age_counts.py:88 SUPPORTED_AGE_CONVENTION = \"observed_interview_age_completed_years\".\n\nPersistence and hashing: packages/microcosm-graph/src/microcosm/graph/store.py:1-11 (content-addressed on-disk artifact store); packages/microcosm-graph/src/microcosm/graph/kernel.py:457-497 (source_hash digests path.read_bytes() of each object's defining module); graph_survey_age_artifact.py:147-153 (implementation_hash over sys.modules[__name__], ages, activation); graph_survey_calibration.py:310 (\"age_artifact\": ages.SurveyAgeCountArtifactKernel().implementation_hash()); graph_survey_calibration.py:350 (decode_survey_age_counts on a stored artifact input).\n\nConsumers: survey_age_calibration.py:102-115 (survey_age_count_artifact_node(population=population.version, node_id=COUNT_NODE), run, decode), :66 (COUNT_NODE = \"survey.age_count_matrix\"), :562-564.\n\nLane context: docs/us-native-row-ceilings.md \u00a71a and \u00a75 (byte transports take the segmented-transport argument); commit 14defbfc0 moved only acs_pums.MAX_EXACT_HOUSEHOLDS/MAX_EXACT_PERSON_ROWS, acs_person_coverage_columns.MAX_SELECTED_ROWS, survey_observed_age.MAX_ROWS, survey_origin_budget.MAX_GROUPS \u2014 it did not touch this module, and graph_survey_age_artifact.py is unchanged on this branch (git log for the file shows only 124fd688e, the staging commit).\n\nArithmetic I ran myself (repo .venv python3, stdlib json only, no files written): (67108864-13)//152 = 441505; (67108864-13-65536)//152 = 441074; difference 431. Header sizes 533/537/539 B (23-char population) and 518/522/524 B (8-char \"invented\") at 1/1000, 1/10 and full source; 774 B at the structural maximum (256-char population, 8-digit rows and people)." + }, + { + "constant": "asec_demographic_source.py:67 _MAX_PERSONS = 600_000", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The census's skeleton is right \u2014 value, five sites, refusal codes, exception type, both counted quantities, binds-at-full=true, memory-or-time, movable. Three substantive errors in the reasoning, one of which inverts the practical conclusion.\n\n(1) WRONG upstream bound named as first-in-order. The census says ACS_ROSTER_BOUND (current_survey_household_roles.py:799) refuses first at full source. It does not \u2014 ORIGIN_ROSTER does. current_survey_household_roles.py:667 `and 0 < len(original) <= MAX_PERSONS` with refusal code \"ORIGIN_ROSTER\" bounds the ORIGINS PERSONS roster (`original = pd.DataFrame(document[\"origins\"][\"persons\"]...)`, :651-652 \u2014 the stacked selected survey, ~3,471,383 persons at full source) against the same MAX_PERSONS = 2,097,152 (:86, :89). `_origins` is called at :1127; `_acs_roster` (which carries ACS_ROSTER_BOUND) is not called until :1136. So the order at full source is ORIGIN_ROSTER (:667) \u2192 ACS_ROSTER_BOUND (:799) \u2192 never ACS_ROWS. Both raise plain `ValueError`, message-prefixed \"CURRENT_SURVEY_HOUSEHOLD_ROLES_\" (:239-241), not DemographicSourceRefusalError. The census's cited call lines :1135/:1137 are also off: actual :1136 (`_acs_roster`) and :1140 (`_acs_states`). Caveat: I did not open `preparation._checked()` (:1121, survey_population_preparation) for a still-earlier refusal, so \"ORIGIN_ROSTER first\" is first *within* qualify_current_survey_household_roles.\n\n(2) \"Never reached at that site\" is false outside full source, and this matters. Ceiling thresholds as a fraction of full source: ACS_ROWS 600,000/3,422,888 = 0.175; ORIGIN_ROSTER 2,097,152/3,471,383 = 0.604; ACS_ROSTER_BOUND 2,097,152/3,422,888 = 0.613. For every sample fraction in (0.175, 0.604] the 600,000 ceiling at :1016 is the FIRST and ONLY refusal in the whole chain \u2014 a 1/5, 1/4, 1/3 or 1/2 run reaches _acs_states with a retained roster over 600k while both 2,097,152 bounds still have headroom. So _MAX_PERSONS is the first ceiling you hit when scaling past ~1/5.7, and raising it is NECESSARY (and necessary first) \u2014 merely not SUFFICIENT for full source. The census's framing (\"simply never reached at that site\") would mislead someone into deprioritizing it.\n\n(3) The full-source ACS count is exact, not extrapolated, and the census is ~3% low. Not ~3,329,258 (5.5x) but exactly 3,422,888 (5.70x). Proof from the pinned preparation.json: /catalogues/acs/counts occupied_hu 1,348,408 + institutional_gq 84,422 + noninstitutional_gq 98,784 = 1,531,614 non-vacant ACS households; + /catalogues/asec/counts/households 55,762 = 1,587,376, which equals supplied_households exactly. So at full source every non-vacant ACS household is selected, and since vacancies carry no persons the retained roster is every ACS person: /catalogues/acs/counts/people = 3,422,888. At 1/10 the figure is 342,289 (census said ~332,926); both under 600,000, so neither binds at 1/10, headroom only 1.75x.\n\nOne strengthening the census missed, load-bearing for safe-to-move: quantity (a) is not merely \"fixed at 432523\" by measurement \u2014 it is EXACTLY PINNED and enforced. education_assistance_source.py:117/139/161 pin rows=146,133 / 144,265 / 142,125 (sum 432,523), enforced at :442-445 (`if len(raw) != pins.rows: raise ValueError`), and asec_demographic_source.py:1187 `_require(len(positions) == member_rows, \"DEMOGRAPHIC_COHORT_ROWS\")` forces the frame's per-cohort person rows to equal the pinned member rows. So the ASEC upstream-file-size assertion is carried by those exact pins, NOT by the 600,000 ceiling \u2014 which is exactly why moving 600,000 weakens no category-(c) check.", + "safe_to_move": true, + "safe_to_move_reason": "Safe. Not a fixed-width encoding: `rows` lives as a JSON integer in the header (read at :1103 via `header[\"rows\"]`), and the only struct packing in the module is `struct.pack(\" 152 B/row); :43-45 (_require raises plain ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason)); :54-59 (_shape; :56 \"SHAPE\" people<=MAX_PEOPLE; :57-58 `len(MAGIC) + 4 + MAX_HEADER_BYTES + rows * _WIDTH * 8 <= MAX_BYTES` -> \"LIMIT\"); :73 (_shape from _validate); :100 (\"PAYLOAD\", len(payload) <= MAX_BYTES on decode); :104-108 (4-byte big-endian header length, \"HEADER_LIMIT\"); :119 (\"BYTE_COUNT\", len(payload)-offset == rows*_WIDTH*8); :120-122 (frombuffer \" exactly 2x rows).\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py:54 (MAX_PAYLOAD_BYTES = 64 MiB); :55-60 (comment: \"a byte transport, not a row count, and is deliberately left alone\"); :612 (_require(len(piece) <= MAX_PAYLOAD_BYTES - len(payload), \"TRANSPORT_LIMIT\")).\npackages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_population.py:61-62 (PREPARATION_MAX_BYTES / ALLOCATION_MAX_BYTES = 64 MiB); :305-312 (\"PREPARATION_BYTES\" / \"CONTEXT_BYTES\" on the whole payload).\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_population_preparation.py:50,:57-58 (MAX_PAYLOAD_BYTES, MAX_SEGMENT_BYTES, MAX_ROSTER_BYTES = 4 GiB); :129-163 (_roster_segments, \"PAYLOAD_LIMIT\"/\"ROSTER_LIMIT\"); :250-259 (_roster_payload); :2010,:2064 (preparation payload built through _roster_payload).\npackages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_calibration.py:44-45 (own MAX_BYTES; MAX_ROWS = MAX_BYTES // 128 = 524,288); :119 (\"ROW_COUNT\" on the same cloned household_ids list); :350 (decode_survey_age_counts consumer).\npackages/microcosm-build/src/microcosm/build/us_runtime/asec_housing_status.py:55 (PAYLOAD_MAX_BYTES derived FROM MAX_HOUSEHOLDS \u2014 the genuine fixed-width-encoding shape, for contrast).\npackages/microcosm-build/tests/test_us_survey_age_artifact.py:161-165 (monkeypatch MAX_BYTES = 1; pytest.raises(ValueError, match=\"LIMIT\")).\ndocs/us-native-row-ceilings.md \u00a71a, \u00a72 table (combined-clone households 3,174,752 full / 317,475 at 1/10), \u00a74, \u00a75 (survey_origin_budget.MAX_PAYLOAD_BYTES admits 87,838 households).\nArithmetic recomputed: (67,108,864 - 9 - 4 - 65,536) // 152 = 441,074; 317,475/441,074 = 71.98%; 3,174,752/441,074 = 7.198; 3,174,752 * 152 = 482,562,304 B." + }, + { + "constant": "graph_survey_calibration.py:45 MAX_ROWS = MAX_BYTES // 128 = 524,288", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The row's headline verdicts survive (value 524,288; counts CLONED household rows; binds at full source; memory-or-time, not encoding-width). Four things are wrong or incomplete.\n\n(1) MATERIAL OMISSION \u2014 a second enforcement site with a different refusal code and a different message prefix. The row names only graph_survey_calibration.py:119. But the PRODUCER checks the same MAX_ROWS first: graph_survey_budget.py:77 `and 0 < len(ids) <= numeric.MAX_ROWS` inside the `_require(..., \"COMPLETE_CLONE_SHAPE\")` spanning :76-85, and that module's `_require` (graph_survey_budget.py:38-40) raises `ValueError(\"SURVEY_BUDGET_TRANSPORT_\" + reason)` \u2014 i.e. `ValueError(\"SURVEY_BUDGET_TRANSPORT_COMPLETE_CLONE_SHAPE\")`, NOT \"SURVEY_CALIBRATION_ROW_COUNT\". `numeric_survey_budget_payload` builds the document and only then calls `decode_numeric_survey_bounds` (graph_survey_budget.py:156), so at full source :77 fires before :119 ever sees the bytes. A boundary test that exercises only :119 does not cover the site that actually refuses, and moving MAX_ROWS moves both. The row's \"enforced at\" and \"refusal code\" fields are therefore incomplete, not merely imprecise.\n\n(2) WRONG at HEAD \u2014 \"graph_survey_population.py:61 ... refuse[s] at ~83k-97k selected households\". Line 61 is `PREPARATION_MAX_BYTES = 64 * 1024**2` and :62 is `ALLOCATION_MAX_BYTES`; the transport lane already segmented the allocation payload (graph_survey_population.py:379-408), so ALLOCATION_MAX_BYTES is now a per-SEGMENT ceiling (:385-393 closes a segment rather than refusing) and the total is `ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES` (:68) = 4 GiB, with :162 capping `len(plan.selected) <= ALLOCATION_ROSTER_BYTES // 128` = 33,554,432. At 316.665 B per instruction document a full-source roster is ~503 MB against 4 GiB: graph_survey_population refuses NOTHING at full source. The 96,860-household preparation-receipt ceiling the row is reaching for was lifted by that lane (docs/us-native-row-ceilings.md \u00a75 says so explicitly). The row's conclusion (\"never reached in the assembled lane\") still holds, but on the other two citations only.\n\n(3) The row's tightest-upstream citation points at a constant, not a refusal. graph_survey_age_artifact.py:33 is just `MAX_BYTES = 64 * 1024**2`; the refusal is `_shape` at :57-58, `_require(len(MAGIC) + 4 + MAX_HEADER_BYTES + rows * _WIDTH * 8 <= MAX_BYTES, \"LIMIT\")`, raising `ValueError(\"SURVEY_AGE_ARTIFACT_LIMIT\")` (:43-45). Its 441,074 IS correct and I re-derived it: _WIDTH = 19 (graph_survey_age_artifact.py:36-37 over national_age_activation's 18 bands, asserted at national_age_activation.py:309), (67,108,864 \u2212 9 \u2212 4 \u2212 65,536) // 152 = 441,074. That artifact is genuinely encoding-width (the reader relies on the width at :119, `len(payload) - offset == header[\"rows\"] * _WIDTH * 8`), unlike the constant under test. The other real tighter bound is survey_origin_budget.py:54 `MAX_PAYLOAD_BYTES`, enforced at :612 `_require(len(piece) <= MAX_PAYLOAD_BYTES - len(payload), \"TRANSPORT_LIMIT\")` raising `SurveyOriginBudgetError` (:69-75, a ValueError subclass) \u2014 87,838 households by that lane's own measurement (docs \u00a75, 764 B/group).\n\n(4) Line citations drift by a few lines throughout: \"population\": clone.COMBINED_CLONE_NODE is at graph_survey_budget.py:146 (:145 is \"budget_sha256\"); the ids loop is survey_origin_budget.py:519-528 and the `\"CLONE_CARDINALITY\"` assertion is :529-531, not :513-522/:523-526. Also \"at 1/10: 317,476 (158,738 \u00d7 2)\" vs this lane's own table (docs \u00a72) of 317,475 (158,737 \u00d7 2) \u2014 1,587,376/10 = 158,737.6, so either rounding is defensible; neither changes the verdict (both < 524,288).\n\nTwo things the row got right that I checked rather than assumed. The counted quantity really is the cloned roster, not an upstream ACS/ASEC file: graph_survey_budget.py:82 `len(ids) == 2 * len(records)`, :146 `\"population\": clone.COMBINED_CLONE_NODE`, and survey_origin_budget.py:520-528 builds one id per row of the EXPANDED (post-clone) household table with :529-531 pinning it at 2 \u00d7 len(instructions). And the byte arithmetic: I rebuilt the document shape of graph_survey_budget.py:143-152 under canonical_json's exact encoder (microcosm-graph/src/microcosm/graph/canonical.py:50-57) at full-source id widths, with a real float.hex() token produced by survey_origin_budget._reference's own 8b/4a arithmetic, and measured 73.9-74.6 B/row (39.1 MiB at exactly 524,288 rows). So MAX_BYTES binds at ~908,000 rows \u2014 the row's \"900k\" end is right, its \"1.05M\" end is not reachable at this document shape \u2014 and MAX_ROWS at 524,288 is indeed the tighter half of the pair. A full-source document would be ~224 MiB, 3.5\u00d7 MAX_BYTES.", + "safe_to_move": true, + "safe_to_move_reason": "Safe in the taxonomy sense, but useless alone, and out of scope for this lane's rule.\n\nSafe: nothing decodes 128 as a row width. The payload is variable-width canonical JSON, and the decoder's byte guard is independent of the row guard \u2014 graph_survey_calibration.py:105 checks `len(payload) <= MAX_BYTES` and :111 re-checks `canonical_json(document) == payload` byte-for-byte, so no reader computes HEADER + ROW_BYTES * MAX_ROWS the way graph_survey_age_artifact.py:57-58/:119 or the asec_housing_* PAYLOAD_MAX_BYTES family do. It is not an upstream-file assertion either: graph_survey_budget.py:146 pins the population to clone.COMBINED_CLONE_NODE and :82 to 2 \u00d7 origins, a roster this build chooses. Not a domain invariant. So: an explicit, deliberately conservative resource ceiling \u2014 category (a).\n\nUseless alone, for a reason the census row does not state. MAX_ROWS is *defined* as MAX_BYTES // 128, so it cannot be raised without either breaking that derivation or raising MAX_BYTES \u2014 and MAX_BYTES cannot be raised past 64 MiB, because graph_survey_budget.py:153 hands it to `graph._bounded_json`, whose own guard at graph_survey_population.py:264 refuses `limit > 64 * 1024**2` with \"TRANSPORT_LIMIT\" (SurveyPopulationGraphError). Even if MAX_ROWS alone were lifted to admit 3,174,752 rows, my measurement puts that document at ~224 MiB, so the refusal just moves to \"NUMERIC_LIMIT\" (graph_survey_budget.py:155) or \"PAYLOAD\" (:105) \u2014 a 3.5\u00d7 shortfall, not a near miss.\n\nOut of scope for this rule. This is a byte-derived row pre-check, exactly the class docs/us-native-row-ceilings.md \u00a71a consigns to the transport lane (\"byte transports take the transport lane's argument, not this one\") and \u00a75 lists among the fourteen bounds a full-source build meets but this lane does not move. The right fix is a segmented numeric-bounds transport under one explicit total, mirroring graph_survey_population.py:379-408, not a larger single cap. Moving it under the 4\u00d7-rounded-up rule would also produce a number with no relationship to the byte ceiling it is derived from, which is precisely the drift test_us_native_row_ceilings.py exists to catch.\n\nOne cost worth naming if it is ever touched: SurveyAgeCalibrationKernel.implementation_hash (graph_survey_calibration.py:301-319) is `source_hash(sys.modules[__name__], group_bounds, ...)`, so editing this module \u2014 even a comment \u2014 moves the kernel implementation hash and with it every node key and store address for the stages that carry it (docs \u00a77 records the same for the five constants already moved). It is not in acs_native_coverage_binding._ACCEPTED (:34-39), so no byte pin fails, but the identity artifacts do move.", + "evidence": "CONSTANT AND ITS TWO ENFORCEMENT SITES\npackages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_calibration.py:44-45 \u2014 MAX_BYTES = 64 * 1024**2; MAX_ROWS = MAX_BYTES // 128 (computed: 67,108,864 // 128 = 524,288)\ngraph_survey_calibration.py:119 \u2014 _require(type(ids) is list and 0 < len(ids) <= MAX_ROWS, \"ROW_COUNT\")\ngraph_survey_calibration.py:61-63 \u2014 def _require: raise ValueError(\"SURVEY_CALIBRATION_\" + reason) [\u2192 ValueError(\"SURVEY_CALIBRATION_ROW_COUNT\")]\ngraph_survey_budget.py:76-85 \u2014 _require(... and 0 < len(ids) <= numeric.MAX_ROWS ... and len(ids) == 2 * len(records), \"COMPLETE_CLONE_SHAPE\") [SECOND, EARLIER SITE, omitted by the census row]\ngraph_survey_budget.py:38-40 \u2014 def _require: raise ValueError(\"SURVEY_BUDGET_TRANSPORT_\" + reason)\ngraph_survey_budget.py:156 \u2014 numeric.decode_numeric_survey_bounds(output) [:77 precedes :119]\n\nTHE COUNTED QUANTITY = CLONED HOUSEHOLD ROWS\ngraph_survey_budget.py:143-152 \u2014 document fields: household_ids, group_indices, group_upper_hex, row_upper_hex, incoming_hex\ngraph_survey_budget.py:146 \u2014 \"population\": clone.COMBINED_CLONE_NODE (census row said :145; :145 is \"budget_sha256\")\ngraph_survey_budget.py:82 \u2014 len(ids) == 2 * len(records)\nsurvey_origin_budget.py:519-528 \u2014 ids/groups built one per row of `expanded.frame.table(\"household\")` (the post-clone table), role in (0,1)\nsurvey_origin_budget.py:529-531 \u2014 _require(len(ids) == 2 * len(instructions) and len(set(ids)) == len(ids), \"CLONE_CARDINALITY\")\nsurvey_origin_budget.py:413 \u2014 _require(0 < len(instructions) <= MAX_GROUPS, \"GROUP_COUNT_BOUND\") [one instruction per selected household]\nsurvey_origin_budget.py:61 \u2014 MAX_GROUPS = 7_000_000 (2\u00d7 that = 14M > MAX_ROWS, so not the tighter one)\n\nBYTE ARITHMETIC (measured by me, not estimated)\nmicrocosm-graph/src/microcosm/graph/canonical.py:50-57 \u2014 canonical_json = json.dumps(ensure_ascii=False, allow_nan=False, sort_keys=True, separators=(\",\",\":\")).encode(\"utf-8\")\nsurvey_origin_budget.py:96-120 \u2014 _reference: b=d/p, a=d*s/p, upper=min(8.0*b, 4.0*a) \u2014 used to generate a real float.hex() token\nReconstructed the :143-152 document at full-source id widths: 73.78 B/row at 100k rows, 73.89 at 200k, 74.58 at 524,288 (39,099,655 B). \u21d2 MAX_BYTES binds at ~908,000 rows; full source (3,174,752 rows) \u2248 223.7 MiB = 3.5\u00d7 MAX_BYTES.\ngraph_survey_calibration.py:105 \u2014 _require(... 0 < len(payload) <= MAX_BYTES, \"PAYLOAD\") [independent byte guard]\ngraph_survey_calibration.py:111 \u2014 _require(canonical_json(document) == payload, \"CANONICAL\") [variable-width, no fixed row stride]\ngraph_survey_budget.py:153,155 \u2014 graph._bounded_json(..., numeric.MAX_BYTES); _require(len(output) <= numeric.MAX_BYTES, \"NUMERIC_LIMIT\")\ngraph_survey_population.py:262-281 \u2014 def _bounded_json; :264 _require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\") [hard cap on MAX_BYTES itself]\n\nTIGHTER UPSTREAM BOUNDS ON THE SAME CLONED ROSTER\ngraph_survey_age_artifact.py:33-37 \u2014 MAX_BYTES = 64*1024**2; MAX_HEADER_BYTES = 65_536; _COLUMNS from NATIONAL_AGE_ACTIVATION.bands; _WIDTH = 1 + len(_COLUMNS)\nnational_age_activation.py:309 \u2014 _require(len(self.bands) == 18, ...) \u21d2 _WIDTH = 19\ngraph_survey_age_artifact.py:57-58 \u2014 _require(len(MAGIC) + 4 + MAX_HEADER_BYTES + rows * _WIDTH * 8 <= MAX_BYTES, \"LIMIT\"); :43-45 raises ValueError(\"SURVEY_AGE_ARTIFACT_\" + reason). (67,108,864 \u2212 9 \u2212 4 \u2212 65,536)//152 = 441,074 < 524,288\ngraph_survey_age_artifact.py:119 \u2014 _require(len(payload) - offset == header[\"rows\"] * _WIDTH * 8, \"BYTE_COUNT\") [this one IS a fixed-width encoding]\ngraph_survey_calibration.py:351-355 \u2014 \"COMPLETE_ORDERED_POPULATION\": counts.household_ids must equal bounds.grouped.household_ids \u21d2 same roster, so :57-58 binds first\nsurvey_origin_budget.py:54 \u2014 MAX_PAYLOAD_BYTES = 64 * 1024**2; :611-613 append(): _require(len(piece) <= MAX_PAYLOAD_BYTES - len(payload), \"TRANSPORT_LIMIT\"); :69-75 class SurveyOriginBudgetError(ValueError) / raise SurveyOriginBudgetError(reason)\ngraph_survey_budget.py:50-53 \u2014 _require(... <= budgets.MAX_PAYLOAD_BYTES, \"PAYLOAD\")\n\nWHY graph_survey_population.py NO LONGER BINDS (refutes part of the row's reasoning)\ngraph_survey_population.py:61-62 \u2014 PREPARATION_MAX_BYTES / ALLOCATION_MAX_BYTES, both 64*1024**2\ngraph_survey_population.py:62-68 \u2014 comment + ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES (4 GiB)\ngraph_survey_population.py:356-368 \u2014 docstring: segments close before ALLOCATION_MAX_BYTES; total bounded separately\ngraph_survey_population.py:385-397 \u2014 segment close at ALLOCATION_MAX_BYTES; \"ALLOCATION_ROSTER_LIMIT\" against ALLOCATION_ROSTER_BYTES\ngraph_survey_population.py:162 \u2014 _require(len(plan.selected) <= ALLOCATION_ROSTER_BYTES // 128, \"ALLOCATION_LIMIT\") = 33,554,432\n\nCOUNTS CROSS-CHECK\ndocs/us-native-row-ceilings.md \u00a72 table \u2014 stacked households 1,587,376; combined-clone households 3,174,752; at 1/10, 317,475\npackages/microcosm-build/tests/test_us_native_row_ceilings.py:44-53 \u2014 STACKED_HOUSEHOLDS = 1_587_376, COMBINED_CLONE_PERSONS = 7_130_026; MOVED table (MAX_ROWS under test is NOT in it)\ndocs/us-native-row-ceilings.md \u00a75 \u2014 survey_origin_budget.MAX_PAYLOAD_BYTES admits 87,838 households; the 96,860 preparation-receipt ceiling was LIFTED by the transport lane\ndocs/us-native-scale-transport.md \u00a71 \u2014 bytes(f) = 39,242 + 692.4375 \u00d7 1,587,376 \u00d7 f for the preparation receipt\n\nPIN EXPOSURE\ngraph_survey_calibration.py:301-319 \u2014 implementation_hash over source_hash(sys.modules[__name__], group_bounds, ...)\nacs_native_coverage_binding.py:34-39 \u2014 _ACCEPTED byte pins; graph_survey_calibration.py is not among them" + }, + { + "constant": "survey_origin_budget.py:54 MAX_PAYLOAD_BYTES = 64 * 1024**2 = 67,108,864", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The headline verdict survives: the value, the three enforcement sites, the refusal codes, the counted quantity, and \"binds at full source\" are all correct as I read them. But four statements in the row are wrong, and the one fact that decides \"can this be lifted\" is missing.\n\n(1) WRONG \u2014 \"No reader depends on the value.\" A reader does. graph_survey_budget.py:49-51 `_document()` opens with `_require(type(payload) is bytes and 0 < len(payload) <= budgets.MAX_PAYLOAD_BYTES, \"PAYLOAD\")`, and graph_survey_budget.py:39-41 makes that `ValueError(\"SURVEY_BUDGET_TRANSPORT_PAYLOAD\")`. It is coupled by import of the same symbol, so it moves in lockstep and no wire format breaks \u2014 but the census's flat \"no reader depends on the value\" is false and would mislead anyone editing line 54 without grepping.\n\n(2) WRONG CITATION \u2014 \"re-decoded by graph_survey_budget.py:141-156.\" The decode is graph_survey_budget.py:49-64 (`json.loads(payload)` at :54). Lines 141-156 are the opposite direction: the numeric *re-encode*, `graph._bounded_json({...}, numeric.MAX_BYTES)` at :142-153 followed by `_require(len(output) <= numeric.MAX_BYTES, \"NUMERIC_LIMIT\")` at :155.\n\n(3) MISSING, AND DECISIVE \u2014 MAX_PAYLOAD_BYTES cannot exceed 64 MiB at all. `_json` (survey_origin_budget.py:86) passes it straight into `graph._bounded_json(value, MAX_PAYLOAD_BYTES)`, and that function's first line is graph_survey_population.py:264: `_require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\")`. That bound is on the *limit argument itself* and is hardcoded in a different module, so it does NOT move with the constant. Set MAX_PAYLOAD_BYTES to 128 MiB and every `_json` call in survey_origin_budget raises SurveyPopulationGraphError(\"TRANSPORT_LIMIT\") immediately, at any roster size \u2014 the module stops working rather than accepting larger documents. A census whose purpose is deciding what can be lifted must carry this.\n\n(4) MISSING \u2014 two downstream ceilings gate the same full-source run and would have to move too: graph_survey_calibration.py:44-45 `MAX_BYTES = 64 * 1024**2` and `MAX_ROWS = MAX_BYTES // 128` = 524,288, enforced at graph_survey_budget.py:75 on `len(ids)` (= 2 x selected) inside \"SURVEY_BUDGET_TRANSPORT_COMPLETE_CLONE_SHAPE\", which caps selected households at 262,144.\n\n(5) Minor citation drift: `SurveyOriginBudgetError(ValueError)` is at survey_origin_budget.py:72, not :66. `checked_view` is called at survey_origin_budget.py:939, not :933-935 (the substance of the independence argument is right \u2014 `checked_view` at survey_population_preparation.py:1831-1836 applies no byte bound, unlike graph_survey_population.py:305's PREPARATION_BYTES). The `_bounded_json` enforcement lines are graph_survey_population.py:264, :273 and :275; :263 is the docstring.\n\n(6) Minor arithmetic: I independently encoded the exact field set (survey_origin_budget.py:565-583 plus the eleven `_reference` keys at :120-131) with the module's canonical encoder settings. A full-source record is 693 B, not 728 B \u2014 at fraction=1 the inclusion probability is Fraction(count, eligible) with count == eligible (survey_domain_sample.py:127-128), so p=[1,1] and the b/a pairs are SMALL; the 726 B figure is the 1/10 record, where the reduced probability fraction inflates b and a. The census has this backwards. Consequences: 1/10 total ~119.5 MB (origins 115.4 + lists 4.1), not ~124 MB \u2014 the census double-counted the two header lists (it wrote ~8.9 MB for 2 x 317,476 entries at ~7 B, which is 4.4 MB). Full source ~1.149 GB (origins 1.102 GB + lists 47.5 MB), not ~1.21 GB. Refusal point ~89,400 selected households at 1/10 widths and ~93,500 at full-source widths, not \"roughly 88,000\". None of this changes any verdict: 1/100 is ~11.9 MB and clears; 1/10 and full source both blow the cap, full source by ~17x.\n\nEverything else I confirmed. The counted quantity is exactly as claimed and is NOT an upstream file row count: `records()` (survey_origin_budget.py:539) yields one record per `instructions` entry, and instructions are one per selected household (graph_survey_population.py:156 `_require(len(household_origins) == len(plan.selected), \"ORIGIN_COUNT\")`); `ids`/`groups` are built by iterating the *expanded* (cloned) household table at survey_origin_budget.py:519-528 and checked at :529-532 `_require(len(ids) == 2 * len(instructions) ..., \"CLONE_CARDINALITY\")`. So: origins scale with the selected/stacked roster, the two header lists with the cloned roster (2x). Upstream bounds are all looser \u2014 MAX_GROUPS 7,000,000 at :413 (\"GROUP_COUNT_BOUND\"), and graph_survey_population.py:162 \"ALLOCATION_LIMIT\" at ALLOCATION_ROSTER_BYTES // 128 = 33,554,432. The preparation itself is issuable at full source: survey_population_preparation.py:2010 and :2173 use `_roster_payload(document)`, i.e. the segmented transport with MAX_ROSTER_BYTES = 64 x 64 MiB = 4 GiB (:57-58), not the 64 MiB `_encode` \u2014 so the ~1.02 GiB full-source preparation clears, and this ceiling is genuinely reached rather than shadowed. One nuance the census missed but that does not change the binding site: survey_origin_budget.py:595 also calls `_json(view.receipt[\"selection\"])` on the per-selected-household plan document (survey_population_preparation.py:384-396), another MAX_PAYLOAD_BYTES site, but at ~250 B/household it refuses around 250k households, later than the :612 append loop.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe as a standalone edit, though not for either of the two disqualifying reasons in the taxonomy. It is not a fixed-width encoding (the artifact is variable-length canonical JSON, no HEADER + ROW_BYTES*MAX_N arithmetic anywhere), and it asserts nothing about a real CPS/ACS/PUF file's size \u2014 it bounds a document derived from the selected roster. It is a deliberate resource ceiling, category (a).\n\nBut raising the literal at survey_origin_budget.py:54 alone does not loosen anything; it breaks the module. `_json` at :86 hands MAX_PAYLOAD_BYTES to `graph._bounded_json`, whose first statement is graph_survey_population.py:264 `_require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\")` \u2014 a hardcoded cap on the limit argument, living in another module, that does not move with the symbol. Any value above 64 MiB makes every `_json` call raise SurveyPopulationGraphError(\"TRANSPORT_LIMIT\") at any roster size.\n\nEven with that cap lifted, a full-source run needs three more coordinated moves: graph_survey_calibration.py:44 MAX_BYTES (64 MiB, bounds the numeric projection at graph_survey_budget.py:142-155), graph_survey_calibration.py:45 MAX_ROWS (524,288, enforced on 2 x selected at graph_survey_budget.py:75, so selected <= 262,144), and graph_survey_population.py:61 PREPARATION_MAX_BYTES if the graph-adapter path (:305) is used. And the target is ~1.15 GB, 17x the current cap, against a document accumulated whole in one bytearray at :608-620 \u2014 so the honest fix is the segmented transport the preparation already adopted (survey_population_preparation.py:57-58, MAX_SEGMENT_BYTES / MAX_ROSTER_BYTES), which is exactly what commit 14defbfc0's message meant by \"it takes the segmented transport argument, not this one.\" Leaving it at 64 MiB in that commit was correct.", + "evidence": "Value and refusals \u2014 survey_origin_budget.py:54 (`MAX_PAYLOAD_BYTES = 64 * 1024**2`), :55-60 (the 14defbfc0 comment), :61 (MAX_GROUPS = 7_000_000), :72 (`class SurveyOriginBudgetError(ValueError)`), :76-78 (`_require` raises it), :86 (`_json` -> `graph._bounded_json(value, MAX_PAYLOAD_BYTES)`), :413 (\"GROUP_COUNT_BOUND\"), :612 (`_require(len(piece) <= MAX_PAYLOAD_BYTES - len(payload), \"TRANSPORT_LIMIT\")`), :931 (\"CANDIDATE_LIMIT\"), :939 (`checked_view`).\nDocument shape \u2014 survey_origin_budget.py:501 (`def _document`), :515-517 (positions from the pre-clone table, `\"ROOT_SOURCE_ID_COLLISION\"`), :519-528 (ids/groups from the expanded/cloned table), :529-532 (\"CLONE_CARDINALITY\", `len(ids) == 2 * len(instructions)`), :539 (`def records()`), :565-583 (the yielded field set), :120-131 (the eleven `_reference` keys), :585-602 (header, incl. :595 `selection_sha256`, :600-601 household_ids/group_indices), :604-620 (the streaming append), :829-838 (a separate verification digest using plain `json.dumps`, not a MAX_PAYLOAD_BYTES site).\nThe limit cap \u2014 graph_survey_population.py:108-112 (`SurveyPopulationGraphError(ValueError)` and `_require`), :262-276 (`_bounded_json`), :264 (limit <= 64 MiB), :273 and :275 (per-piece \"TRANSPORT_LIMIT\").\nReader \u2014 graph_survey_budget.py:39-41 (`ValueError(\"SURVEY_BUDGET_TRANSPORT_\" + reason)`), :49-51 (\"PAYLOAD\" against `budgets.MAX_PAYLOAD_BYTES`), :54 (`json.loads`), :69-83 (\"COMPLETE_CLONE_SHAPE\", `len(ids) <= numeric.MAX_ROWS`, `len(ids) == 2 * len(records)`), :142-155 (numeric re-encode, \"NUMERIC_LIMIT\").\nDownstream ceilings \u2014 graph_survey_calibration.py:44-45 (MAX_BYTES = 64 MiB, MAX_ROWS = MAX_BYTES // 128), :119 (\"ROW_COUNT\").\nUpstream \u2014 graph_survey_population.py:61 (PREPARATION_MAX_BYTES), :68 (ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES), :92-105 (AllocationInstruction), :146-162 (`allocation_instructions`, :156 \"ORIGIN_COUNT\", :162 \"ALLOCATION_LIMIT\"), :296-311 (`_checked_preparation`, :305 \"PREPARATION_BYTES\").\nPreparation transport \u2014 survey_population_preparation.py:50-58 (MAX_PAYLOAD_BYTES / MAX_SEGMENT_BYTES / MAX_ROSTER_BYTES), :113-119 (`_encode`), :129 (`_roster_segments`), :384-396 (`_plan_document`, one row per selected household), :1144-1151 (`_bounded_append`, \"ORIGIN_LIMIT\" against MAX_ROSTER_BYTES), :1831-1836 (`checked_view`, no byte bound), :2010 and :2173 (`_roster_payload(document)`).\nFraction widths \u2014 survey_catalogue_selection.py:32-43 (SelectedHousehold), :103-106 (anchors: `Fraction(int(row.hsup_wgt), 100)`, `Fraction(int(row.wgtp))`), :236-240 (cell.inclusion_probability); survey_domain_sample.py:125-128 (`count = max(1, fraction.numerator * eligible // fraction.denominator)`, `probability = Fraction(count, eligible)` -> p = 1 at fraction=1); survey_population_domains.py:204-219 (ACS native_id `2024(HU|GQ)[0-9]{7}` = 13 chars, ASEC <= 5 digits).\nIndependent size computation (scratchpad script, canonical encoder settings matching graph_survey_population.py:267-269): full-source record 693 B, 1/10 record 726 B; totals 11.9 MB at 1/100, 119.5 MB at 1/10, 1.149 GB at full source; cap reached at ~89,400-93,500 selected households." + }, + { + "constant": "acs_person_coverage_authentication.py:37 MAX_BODY_BYTES = 64 * 1024**2 = 67_108_864", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": true, + "correction": "I tried to refute this row and could not on any load-bearing point. Value, both enforcement sites, both refusal codes, the exception type, the wrapping, the counted quantity, the \"binds at full source\" verdict, and the memory-or-time classification all hold when read at HEAD (2ebd246f1). I independently re-measured the ACS source rather than trusting the row's numbers, and its numbers are right \u2014 including where they differ from the ground truth handed to me (see 5 below). Non-material corrections only:\n\n1. COMMENT LINE OFF BY TWO-THREE. The row says \"auth:380's comment\". The comment \"Upper bound on escaped row + statuses + lineage before the low-level reader allocates its selected DataFrame.\" is at acs_person_coverage_authentication.py:377-378; the charge is at :379 and the _require at :380-382. The quoted text is verbatim correct.\n\n2. ARITHMETIC SLIP IN THE \"1,278x\" FIGURE. The budget exhausts at 67,108,864/5,240.6 = 12,805 selected person rows. Against the post-lift MAX_SELECTED_ROWS = 14,000,000 that is 1,093x tighter, not 1,278x. The 78x figure against the pre-lift 1,000,000 is correct. Also, the row's phrase \"post-lift 14,000,000 ones\" is loose: acs_person_coverage_columns.py:39 MAX_ROWS is still 6_000_000 and was deliberately NOT lifted (its comment at :33-38 calls it \"the structural assertion that a genuine file is near that size\"); the four that moved to 7M/14M are acs_pums.MAX_EXACT_HOUSEHOLDS/MAX_EXACT_PERSON_ROWS (:55-56), acs_person_coverage_columns.MAX_SELECTED_ROWS (:40), survey_observed_age.MAX_ROWS and survey_origin_budget.MAX_GROUPS (commit 14defbfc0). None of this changes the conclusion: every one of them is unreachable behind 12,805/14,533.\n\n3. \"SCANNED FIRST INSIDE _preflight\" IS TRUE ONLY ON THE EXACT-KEYS ROUTE. On the full-source route (serialnos=None) acs_native_coverage_binding.py:311 calls coverage._inventory(path, role, frozenset()) with an EMPTY selection set, so auth:375-376 `continue`s on every row and selected_budget is never charged in _preflight at all. At full source the first charge is in load_authenticated_acs_person_coverage's own household _inventory (auth:636-638), which runs at binding:601-603 \u2014 i.e. AFTER binding:575 has already run prepare_acs_housing_population over the whole source and built the full frame and full lexical projection. So the ceiling does not in fact spare the run its full-source construction peak on the whole-source route; it refuses after it. Household-before-person ordering is correct on both routes (binding:322-326 and auth:636-641).\n\n4. NEARLY-ADJACENT BOUND THE ROW DOESN'T MENTION, AND IT DOESN'T PREEMPT. binding:563-565 caps the requested serialnos tuple's canonical JSON at min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES) = 1 MiB, refusal \"CANONICAL_SIZE\" (ACSCoverageAuthenticationError, raised inside coverage._json). At ~17 escaped bytes per SERIALNO that is ~61,000 households \u2014 looser than the 14,533-household body budget, so the row's ordering survives. Worth recording because it is the only other pre-auth ceiling on the selected route.\n\n5. HOUSEHOLD-ROLE PER-RECORD AVERAGE IS SLIGHTLY OPTIMISTIC (the bound binds harder than stated). 598.96 B/record is total csv_hus bytes over all 1,631,969 hus rows, but only the 1,531,614 person-bearing households are charged (I measured both; the 100,355 difference is vacant/person-less units, whose records are shorter). A strict upper bound on the charged bytes, 6 x 977,484,520 + 1024 x 1,531,614 = 7.43e9 B, is above the row's 7.073e9. Either way 105-111x over.\n\n6. The ~92 B/line NDJSON figure is, as the row says, computed not measured \u2014 but it is well calibrated: the line is 9 JSON values (7 _COLUMNS at auth:45 plus source_member/source_row_ordinal), and the two *_state fields are long words from _field_state (acs_person_coverage_columns.py:119-126: \"observed_code\", \"missing_in_universe\", \"outside_age_universe\"), giving ~90-105 B. auth:661 exhausts near 729,000 persons (row says ~721,600), 1/4.7 of full source. Still a distant second.\n\nWhat I checked and could NOT use to refute: no tighter upstream ceiling exists on this path. custody._MEMBER_MAX/_EXPANDED_MAX (acs_housing_universe_source.py:45-47) are 8/16 GiB against measured 1.23 GiB max member and 3.38 GB combined expanded; binding MAX_ARCHIVE_BYTES/MAX_EXPANDED_BYTES/MAX_SOURCE_ROWS (:28-30) are 8 GiB/16 GiB/6,000,000 against 854 MB compressed, 3.38 GB expanded, 3,422,888 person rows; housing's RECEIPT_TOO_LARGE cap (:616) covers a counts-only receipt; PROJECTION_TOO_LARGE (:577) is 8 GiB; acs_pums MAX_EXACT_* are 7M/14M; auth's own NATIVE_ROWS (:416-419) and SELECTED_ROWS (:396-398) are 14M and are evaluated after the budget charge inside the same loop iteration. Nothing fires first.", + "safe_to_move": true, + "safe_to_move_reason": "Safe in KIND \u2014 it is class (a), a deliberate RAM budget \u2014 but not free, and the row understates two coupling costs.\n\nWhy it is (a) and not (b), (c) or (d):\n- Not (c): the charge is skipped for every unselected row (auth:375-376 `if cells[0] not in serialnos: continue`), so it never asserts anything about the source file's size. That job is done by a different, explicitly-labelled constant: acs_person_coverage_columns.py:39 MAX_ROWS = 6_000_000, whose comment at :33-38 says it \"bounds the source file's own person records and stays where it is ... the structural assertion that a genuine file is near that size\", enforced at auth:369 as \"SOURCE_ROWS\". That one must not move; this one is not it.\n- Not (b): I read the encode and decode sides. The payload is framed at auth:696 as `_MAGIC + len(raw_header).to_bytes(4, \"big\") + raw_header + body`, and `_parts` (auth:556-559) slices the header by that 4-byte big-endian length and takes the body as \"the rest\" \u2014 `payload[start + size:]`. The only fixed-width field is the header length, which is bounded by MAX_HEADER_BYTES (1 MiB), not by MAX_BODY_BYTES. The body's true length travels in the header as \"body_bytes\"/\"body_sha256\" (auth:686-687). No reader computes a HEADER + ROW_BYTES * N offset. Moving MAX_BODY_BYTES changes no wire format.\n- Not (d): nothing structural. The 6x multiplier is an escaped-JSON-vs-raw-CSV guess, and the 1024 B/row is a lineage pad.\n\nTwo costs the row does not state, both of which mean \"movable\" \u2260 \"edit the number and push\":\n1. The constant is PUBLISHED IN THE ATTESTATION. auth:178 puts it in _producer()[\"ceilings\"][\"body\"], which goes into every issued header and is compared for equality at auth:712 and auth:594 (\"PRODUCER_CHANGED\"). Every previously issued payload/receipt digest moves.\n2. The MODULE IS PINNED BY DIGEST. auth:43-51 hashes this file into its own producer, and acs_native_coverage_binding.py:34-38 pins acs_person_coverage_authentication.py = 9ec68721a4cf... in _ACCEPTED, enforced at binding:232-237 with refusal \"UNREVIEWED_PREPARATION\" (ACSNativeCoverageBindingError) under the comment \"A new transform/owner version requires explicit review of this successor.\" So any move is a reviewed re-pin, plus the accepted-digest re-pins the branch is already doing.\n3. Raising it also silently raises the cap at binding:646 on the sorted raw person-key list, since that site passes coverage.MAX_BODY_BYTES as the _json cap. Same memory class, but it is a second consumer, not just a comment.\n\nAnd the substantive point the row makes and I agree with: the estimator overcharges by ~57x (5,240 B/row charged vs ~92 B/line the body actually costs). Recalibrating 6*len(raw)+1024 is a better argued change than raising 64 MiB, because the real payload at full source is ~315 MB, not ~18 GB. Whichever is chosen, the argument has to be made against the ~92 B/line the encoder produces, not against the estimate.", + "evidence": "CODE READ AT HEAD 2ebd246f1, all under packages/microcosm-build/src/microcosm/build/us_runtime/:\n\nacs_person_coverage_authentication.py:37 \u2014 MAX_BODY_BYTES = 64 * 1024**2 (value confirmed).\n:45 \u2014 _COLUMNS = (*literal.READ_COLUMNS, \"MIL_state\", \"ESR_state\") (7 columns).\n:43-51 \u2014 _IMPLEMENTATION_FILES includes this module itself.\n:57-58 \u2014 class ACSCoverageAuthenticationError(ValueError); :61-63 \u2014 _require raises exactly that type with the code string.\n:70,79,120 \u2014 _json(value, cap) and its two `_require(count <= cap, \"CANONICAL_SIZE\")` sites (both line numbers in the row are exact).\n:178 \u2014 _producer()[\"ceilings\"][\"body\"] = MAX_BODY_BYTES.\n:369 \u2014 _require(rows <= literal.MAX_ROWS, \"SOURCE_ROWS\").\n:375-376 \u2014 `if cells[0] not in serialnos: continue` (only SELECTED rows are charged).\n:377-378 \u2014 the comment the row quotes.\n:379 \u2014 selected_budget += 6 * len(raw) + 1024.\n:380-382 \u2014 _require(selected_budget <= MAX_BODY_BYTES, \"SELECTED_BODY_BUDGET\").\n:396-398 \u2014 _require(len(selected) < literal.MAX_SELECTED_ROWS, \"SELECTED_ROWS\"), after the budget charge.\n:416-419 \u2014 _native_roster's \"NATIVE_ROWS\" against MAX_SELECTED_ROWS.\n:556-559 \u2014 _parts(): 4-byte big-endian header length, body = payload[start+size:].\n:636-641 \u2014 _inventory(household) then _inventory(person), both with set(keys.SERIALNO).\n:655-661 \u2014 the NDJSON loop and _require(size <= MAX_BODY_BYTES, \"BODY_SIZE\").\n:696 \u2014 payload = _MAGIC + len(raw_header).to_bytes(4,\"big\") + raw_header + body.\n:712 \u2014 _require(_producer() == producer, \"PRODUCER_CHANGED\").\n:722-724 \u2014 except ACSCoverageAuthenticationError: raise / except Exception: \"SOURCE_RECONSTRUCTION_REFUSED\".\n\nacs_native_coverage_binding.py:28-31 \u2014 MAX_ARCHIVE_BYTES 8 GiB, MAX_EXPANDED_BYTES 16 GiB, MAX_SOURCE_ROWS 6_000_000, MAX_EVIDENCE_BYTES 2 MiB.\n:34-38 \u2014 _ACCEPTED pins acs_person_coverage_authentication.py = 9ec68721a4cf...; :232-237 enforces it as \"UNREVIEWED_PREPARATION\".\n:54-56 \u2014 class ACSNativeCoverageBindingError(ValueError) \u2014 independent of the auth error, neither subclasses the other (both ValueError).\n:299-347 \u2014 _preflight; :311 `coverage._inventory(path, role, frozenset())` on the serialnos-is-None route; :322-326 household _inventory first on the exact-keys route; :313-320 SOURCE_ROW_BUDGET/NATIVE_ROW_BUDGET.\n:563-565 \u2014 coverage._json(serialnos, min(MAX_EVIDENCE_BYTES, housing.ACS_HU_RECEIPT_MAX_BYTES)).\n:571 _preflight call; :575 prepare_acs_housing_population; :601-603 load_authenticated_acs_person_coverage.\n:642-646 \u2014 coverage._sha(coverage._json(sorted(zip(keys.SERIALNO, map(int, keys.SPORDER))), coverage.MAX_BODY_BYTES)).\n:733-736 \u2014 except ACSNativeCoverageBindingError: raise / except Exception: raise ACSNativeCoverageBindingError(\"NATIVE_ISSUANCE_REFUSED\") from None (the raise is on :736).\n\nacs_person_coverage_columns.py:32-40 \u2014 READ_COLUMNS, the MAX_ROWS comment, MAX_ROWS = 6_000_000, MAX_SELECTED_ROWS = 14_000_000; :119-126 _field_state's returned state words; :239 and :253 the selected-row bound and the string-dtype frame.\n\nacs_housing_universe_source.py:42-47 \u2014 ACS_HU_SOURCE_MAX_BYTES 8 GiB, ACS_HU_RECEIPT_MAX_BYTES 1 MiB, _MEMBER_MAX 8 GiB, _EXPANDED_MAX 16 GiB, _MEMBER_COUNT_MAX 64; :190-207 _capture's disk check; :573-577 full projection + PROJECTION_TOO_LARGE; :616 RECEIPT_TOO_LARGE over a counts-only receipt; :827-845 prepare_acs_housing_population.\n\nacs_pums.py:55-56 \u2014 MAX_EXACT_HOUSEHOLDS 7_000_000 (:194), MAX_EXACT_PERSON_ROWS 14_000_000 (:270).\n\npackages/microcosm-build/tests/test_us_acs_person_coverage_authentication.py:110-111 \u2014 _refuses() asserts owner.ACSCoverageAuthenticationError with an anchored code match; :472 \u2014 (owner, \"MAX_BODY_BYTES\", 100, \"SELECTED_BODY_BUDGET\") is an enforced test of exactly this site.\n\ngit show 14defbfc0 \u2014 the lift commit's own message records the measured full-source counts 1,531,614 / 3,422,888 / 1,587,376 and states MAX_ROWS stays at 6,000,000 and MAX_PAYLOAD_BYTES stays at 64 MiB.\n\nMEASUREMENTS I RAN MYSELF against the recovered real source /Users/maxghenis/PolicyEngine/_recovered/microcosm-native-sources-20260909/sources/acs/ (csv_pus.zip 602,847,146 B, csv_hus.zip 251,500,587 B):\n- unzip -l: psam_pusa.csv 1,226,543,489 + psam_pusb.csv 1,178,924,624 = 2,405,468,113 B; psam_husa.csv 495,752,055 + psam_husb.csv 481,732,465 = 977,484,520 B. Both match the row's numerators exactly. Combined expanded 3.38 GB is far under the 16 GiB gates; combined compressed 854 MB is far under the 8 GiB gate.\n- line counts: person 1,743,752 + 1,679,138 = 3,422,890 lines minus 2 headers = 3,422,888 data rows (exact match); household 827,134 + 804,837 = 1,631,971 minus 2 = 1,631,969 (exact match).\n- unique adjacent SERIALNO (field 2) in the person members: 776,709 + 754,905 = 1,531,614 \u2014 i.e. exactly the households that have at least one person row, which is exactly the set of hus rows auth:381 charges at full source (set(keys.SERIALNO) comes from frame.person via _native_roster). This confirms 1,531,614, not the 1,587,376 supplied_households in the handed-down ground truth, is the right count for THIS site; 1,587,376 is the origin-budget's universe, which includes units with no person rows.\n- Implied charge at full source: person role 6 x 2,405,468,113 + 1024 x 3,422,888 = 1.794e10 B = 267x the 67,108,864 B ceiling; household role, upper bound 6 x 977,484,520 + 1024 x 1,531,614 = 7.43e9 B = 111x. Exhaustion points 12,805 selected persons and 14,533 selected households.\n\nNOT ESTABLISHED: I did not execute the issuance path against the real source (it needs the full capture and would take the full-source construction peak), so the refusal is derived from the code and the measured inputs, not observed. The ~92 B NDJSON line is computed from the encoder, _COLUMNS and _field_state's literals, not measured." + }, + { + "constant": "native_household_origin.py:48 MAX_SOURCE_BYTES = 512 * 1024**2 = 536,870,912", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The census's factual core is right (value, both sites, both refusal codes, both counted quantities, full-source independence from the survey fraction, binds_at_full=false). I disagree on the CLASSIFICATION and on one missed tighter bound, plus four line-citation errors.\n\n1) protects = memory-or-time, NOT upstream-file-size. The census says the :127 arm \"is a structural assertion about the publisher's file and MUST NOT be moved\". It is not. MAX_SOURCE_BYTES at native_household_origin.py:127 bounds `self.size_bytes` \u2014 a DECLARED integer field on AsecNativeMemberPin \u2014 not the real file. Every genuine assertion about the publisher's hhpub CSV is an exact equality that raising MAX_SOURCE_BYTES leaves completely untouched:\n - the literal per-pin size at native_household_origin.py:146 (30,259,450), :156 (30,448,089), :166 (33,298,479);\n - member_sha256 at :145/:155/:165, re-proved by `_snapshot(member_paths[pin.income_year], capture, size=pin.size_bytes) == pin.member_sha256` at native_household_origin.py:313-317, where _snapshot itself enforces `_require(before.st_size == size, \"STUDENT_SOURCE_SIZE\")` at asec_student_controls.py:103 and `count == size` at asec_student_controls.py:117-119;\n - the archive binding `self.archive_sha256 == ASEC_EDUCATION_ASSISTANCE_ARCHIVES[self.income_year].zip_sha256` at native_household_origin.py:121-122;\n - the post-parse re-hash `_sha(capture.read_bytes()) == pin.member_sha256` at native_household_origin.py:353-355.\n Taxonomy (c) requires that moving the bound weaken a real check. Raising MAX_SOURCE_BYTES to any value >= 33,298,479 weakens none of the above: a member file must still be byte-exact and digest-exact. What the ceiling actually limits is how large a FUTURE pin declaration may be, which matters only because size_bytes drives _snapshot's copy loop and because native_household_origin.py:354 reads the entire capture into memory via capture.read_bytes(). That is taxonomy (a): a deliberate but arbitrary memory/time ceiling. The same applies a fortiori at :196, which the census already concedes is memory-or-time. So the single classification for the constant is memory-or-time, at both sites.\n Corroborating detail the census did not note: the three pins are constructed at MODULE IMPORT (native_household_origin.py:138-169), so the :127 arm can only ever fire on a code change, never on runtime data.\n\n2) Missed tighter upstream limit on the ASEC arm of :196. The same _require at native_household_origin.py:125-131 also caps `0 < self.rows <= 1_000_000` at :129. _ASEC_MEMBER_PINS is a closed 3-element tuple (:138-169) and produce_asec_native_origins emits one record per data row (:340-350) with `_require(count <= pin.rows, \"ASEC_MEMBER_ROWS\")` at :351. Worst case per canonical record, with H_SEQ at the 18-digit maximum permitted by :335, is ~95 bytes; 3 x 1,000,000 x 95 = 285 MB = 53% of 536,870,912. So \"SOURCE_SIZE\" at :196 is STRUCTURALLY UNREACHABLE on the ASEC arm, not merely slack. The census called it slack.\n\n3) Line-citation errors (conclusions survive, mechanism statements were imprecise):\n - exception type NativeOriginError(ValueError) is at native_household_origin.py:52, not :51 (:51 is blank);\n - the archive-digest cross-check is at :118-124, not :113-120;\n - produce_acs_native_origins spans :256-291, not :264-292; produce_asec_native_origins spans :294-358, not :294-357;\n - the census says the ACS producer \"runs ... with serialnos=None\". It does not pass serialnos at all (native_household_origin.py:258-262); None is the callee's default at acs_housing_universe_source.py:696. The full-universe conclusion is nonetheless correct: _select with serialnos None takes `chosen = tuple(sorted(household.SERIALNO))` at acs_housing_universe_source.py:499 and the receipt records `\"selection\": \"all\"` at :604.\n\n4) Not established: I did not run the producer. The ~150 MB ACS payload is arithmetic from the record shape at :285 and the lexical SERIALNO of the \"acs-housing-lexical-projection/1\" format, not a measurement. 1,631,969 x 92 B = 150,141,148 = 27.97% of the ceiling. The ASEC figure 267,383 x 82 B = 21,925,406 = 4.08%. Max pinned member 33,298,479 = 6.20%.\n\nCredit where due: the census correctly used 1,631,969 (the full housing-record count including 100,355 vacancies and GQ) rather than the 1,531,614 selectable households the lane's own lift commit 14defbfc0 used for acs_pums.MAX_EXACT_HOUSEHOLDS. That discrimination is right and is the single most error-prone judgment in this row.", + "safe_to_move": true, + "safe_to_move_reason": "Safe in the taxonomy sense, but there is no reason to move it and the lane correctly did not (git show --stat 14defbfc0 confirms native_household_origin.py is untouched). Safe because: (a) no fixed-width encoding or byte budget depends on it \u2014 the payload is canonical JSON whose length is only compared, at native_household_origin.py:196, and the whole-repo grep for MAX_SOURCE_BYTES returns exactly three lines (48, 127, 196), with no decoder, struct format, or HEADER+ROW_BYTES*MAX_N computation referencing it; no test pins its value; (b) no structural assertion about the real hhpub CSV would be weakened, because each pin's size is an exact literal cross-checked by _snapshot's `before.st_size == size` (asec_student_controls.py:103) and three separate sha256 equalities (native_household_origin.py:121-122, :313-317, :353-355), and the module docstring at :91-94 requires a root source audit before any pin is added at all.\nReasons not to move it anyway: it never binds (6.2% at :127, ~28% at :196 on the ACS arm, structurally unreachable on the ASEC arm given rows <= 1_000_000 at :129); a row-ceiling lift has nothing to gain from it; and editing the file changes the native-origin kernel identity, since graph_native_origin_implementation.py:31 lists microcosm.build/us_runtime/native_household_origin.py in EXTRA_MODULES and :86-91 hashes its bytes into implementation_manifest()[\"modules\"], which feeds implementation_hash() at :100-104 \u2014 a re-pin cost for zero benefit.", + "evidence": "native_household_origin.py:48 (MAX_SOURCE_BYTES = 512 * 1024**2); :52 (class NativeOriginError(ValueError)); :56-58 (_require raises NativeOriginError(reason)); :105-131 (AsecNativeMemberPin.__post_init__); :118-124 (archive_sha256 == ASEC_EDUCATION_ASSISTANCE_ARCHIVES[income_year].zip_sha256, \"ASEC_MEMBER_ARCHIVE_BINDING\"); :125-131 (0 < size_bytes <= MAX_SOURCE_BYTES and 0 < rows <= 1_000_000, \"ASEC_MEMBER_BOUNDS\" at :130); :138-169 (_ASEC_MEMBER_PINS, module-import construction; sizes 30_259_450/30_448_089/33_298_479 at :146/:156/:166; rows 88_978/89_473/88_932 at :147/:157/:167); :195-196 (_source_document; _require(type(payload) is bytes and len(payload) <= MAX_SOURCE_BYTES, \"SOURCE_SIZE\")); :241-253 (_issue builds canonical_json then calls _source_document); :256-262 (produce_acs_native_origins calls acs.produce_acs_housing_source with no serialnos); :282-285 (one [2024, SERIALNO, index, key] record per ACS housing row); :294-358 (produce_asec_native_origins); :313-317 (_snapshot(..., size=pin.size_bytes) == pin.member_sha256, \"ASEC_MEMBER_SHA256\"); :331-352 (one record per CSV data row; H_SEQ regex [0-9]{1,18} at :335; count <= pin.rows at :351; count == pin.rows at :352); :353-355 (_sha(capture.read_bytes()) == pin.member_sha256, \"ASEC_CAPTURE_CHANGED\").\nasec_student_controls.py:94-123 (_snapshot); :103 (_require(before.st_size == size, \"STUDENT_SOURCE_SIZE\")); :109-113 (chunked copy); :116-120 (count == size and inode/mtime stability).\nacs_housing_universe_source.py:42-43 (ACS_HU_SOURCE_MAX_BYTES = 8 GiB, ACS_HU_RECEIPT_MAX_BYTES = 1 MiB \u2014 both far looser, neither on the origin JSON); :497-535 (_select; serialnos is None -> chosen = tuple(sorted(household.SERIALNO)) at :499); :604 (\"selection\": \"all\" if serialnos is None); :695-699 (produce_acs_housing_source signature, serialnos=None default).\ngraph_native_household_origin.py:73-90 (NativeACSOriginKernel.run -> native.produce_acs_native_origins, no serialnos); :98-113 (NativeASECOriginKernel.run over all three pinned cohorts).\ngraph_native_origin_implementation.py:26-33 (EXTRA_MODULES includes native_household_origin.py); :86-91 (live sha256 of each EXTRA_MODULE into \"modules\"); :100-104 (implementation_hash).\nRecovered artifact /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json, sha256 verified 34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0: catalogues.acs.counts.households = 1631969 (vacancies 100355, institutional_gq 84422, noninstitutional_gq 98784, people 3422888); source_files[9] = [\"asec/hhpub25.csv\", 33298479, \"b5b7351d5d4e5d79ff189f1d90096b16b2d4749671c14328ec6474cfe83ce116\"], matching the pin at native_household_origin.py:165-166.\nRepo-wide grep for MAX_SOURCE_BYTES: only native_household_origin.py:48, :127, :196 (other hits are unrelated constants _MAX_SOURCE_BYTES in tests/test_us_atomic_block_api_sources_native_national.py and _COMPILE_CACHE_MAX_SOURCE_BYTES in acs_native_coverage_binding.py:46).\ngit show --stat 14defbfc0: touches only acs_person_coverage_columns.py, acs_pums.py, survey_observed_age.py, survey_origin_budget.py \u2014 native_household_origin.py untouched." + }, + { + "constant": "current_survey_household_roles.py:86 MAX_PERSONS = 64 * 1024**2 // 32 = 2,097,152", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": true, + "correction": "", + "safe_to_move": true, + "safe_to_move_reason": "Safe. (1) Not a fixed-width encoding or byte budget: every use of this module's MAX_PERSONS is current_survey_household_roles.py:290 (an in-process _live() marker), :667 (the len() comparison), and the alias MAX_ROSTER_ROWS at :89/:292/:799 (another len() comparison on a dict). No HEADER + ROW_BYTES * MAX_PERSONS expression exists in this module. The MAX_PERSONS that DOES appear in byte budgets at graph_composed_asec_binding.py:960 and _asec_current_money_codec.py:27 is a DIFFERENT constant imported from asec_current_money (= 1_000_000) \u2014 I read both import statements (graph_composed_asec_binding.py:77-84, _asec_current_money_codec.py:6-21). (2) Not serialized: household_role_contract_document() (:447-509, read in full) does not contain MAX_PERSONS, so contract_sha256() (:512-520) and every receipt that records it are unaffected; survey_population_preparation._runtime_marker (:607-621) returns in-process (type, value) tuples compared by == within the run, never hashed to disk; the literal 2097152 / 2_097_152 / \"64 * 1024**2 // 32\" appears nowhere else in the repo (grep --include=*.py --include=*.md), so no test or golden pins it; neither graph_implementation_inventory.json nor native_origin_graph_inventory.json references this module. (3) Not an upstream-file assertion: the bounded roster is pinned row-for-row to preparation_frame.person (:672-683) and is the selected/stacked survey, whose size is a function of the selection request, not of any real ACS/CPS file. CAVEAT (already noted by the census): moving it alone does not unblock full source. MAX_ROSTER_ROWS aliases it and moves along, but graph_current_survey_household_roles.py:70 MAX_ARTIFACT_BYTES = 64 * 1024**2, enforced at :274 on the bind node's artifact inputs (which include this node's projection via _expected_artifacts at :266-267), becomes the next refusal \u2014 _projection_bytes (:764-765) emits a 14-column JSON table, far wider than the 32 bytes/row the constant's arithmetic assumes.", + "evidence": "VALUE: current_survey_household_roles.py:86 `MAX_PERSONS = 64 * 1024**2 // 32`; computed = 2,097,152. Alias at :89 `MAX_ROSTER_ROWS = MAX_PERSONS`.\nENFORCEMENT: current_survey_household_roles.py:663-669, compound `require(required <= set(original) and original.person_id.dtype == np.dtype(\"int64\") and original.person_id.is_unique and 0 < len(original) <= MAX_PERSONS, \"ORIGIN_ROSTER\")`. Refusal string literal is exactly \"ORIGIN_ROSTER\" (:668).\nEXCEPTION TYPE: current_survey_household_roles.py:239-241 `def require(condition, reason): if not condition: raise ValueError(\"CURRENT_SURVEY_HOUSEHOLD_ROLES_\" + reason)` \u2014 plain ValueError; `grep -n \"^class .*Error\"` on the module returns nothing, so no subclass shadows it. (Census cited :238-240; actual :239-241.)\nWHAT IS COUNTED: :645-669 `_origins` builds `original` from `document[\"origins\"][\"persons\"][\"rows\"]`. :672-683 `require(np.array_equal(original.index.to_numpy(), people.person_id.to_numpy()) ..., \"ORIGIN_FRAME_IDENTITY\")` where `people = preparation_frame.person` (:671) \u2014 so len(original) == len(preparation frame persons). (Census cited :673-686; actual :672-683.) :684-692 requires `original.source.isin(SURVEYS)` with SURVEYS = (\"acs\",\"asec\") at :85 \u2014 both channels.\nPRE-CLONE, DECISIVE: :1293 `require(len(original) == 2 * len(rows) and set(original) == set(rows.index), \"CLONE_SOURCE_ROSTER\")` in `_aligned` (:1268-1312), where `people = _person_table(receiving)` (:1279) \u2014 the receiving frame is exactly 2x the qualified rows, and :1308-1310 \"WHOLE_CLONE_PAIRS\" confirms every original appears twice. `grep -n clone survey_population_preparation.py` returns zero hits, so the preparation frame carries no clones. Therefore the bounded count is the stacked pre-clone roster = 3,471,383 at full source (ground truth), not 6,942,766.\nBINDS: 3,471,383 > 2,097,152. First bind at fraction 2,097,152/3,471,383 = 0.6041.\nORDERING IN QUALIFIER: current_survey_household_roles.py:1103 `def qualify_current_survey_household_roles`; :1123 `entry = preparation._checked()`; :1127 `origins = _origins(state.frame, document)`; :1136 `roster, acs_owned = _acs_roster(state, serials)` (whose MAX_ROSTER_ROWS check is at :799); :1151 `projection = _projection_bytes(table)`. So ORIGIN_ROSTER is first.\nNO TIGHTER UPSTREAM BOUND: survey_population_preparation.py:1141-1149 `_bounded_append` bounds origins in BYTES against MAX_ROSTER_BYTES (:58 `= 64 * MAX_SEGMENT_BYTES` = 4 GiB) with its own comment at :1142-1143 stating \"a full-source roster is 1.02 GiB\" \u2014 does not bind. The preparation payload is built by `_roster_payload` (:2010, :2173), and `_roster_segments` (:127-167) applies the 64 MiB MAX_PAYLOAD_BYTES per SEGMENT (:160 \"PAYLOAD_LIMIT\"), not to the whole document \u2014 does not bind. survey_population_domains.py:19-20 MAX_HOUSEHOLDS=100_000 / MAX_TOTAL_MEMBERS=1_000_000 are per-batch: domains.py:563 `_require(members <= MAX_TOTAL_MEMBERS, \"BATCH_MEMBER_BOUND\")` and survey_catalogue_selection.py:129-131 cap `_BATCH_HOUSEHOLDS <= min(10_000, domains.MAX_HOUSEHOLDS)` and `_BATCH_PEOPLE <= min(100_000, domains.MAX_TOTAL_MEMBERS)` \u2014 selection is batched, so these bound a batch, not the total. survey_observed_age.py:19 MAX_ROWS = 14_000_000 > both 3.47M and 6.94M.\nNEXT REFUSAL AFTER MOVING: graph_current_survey_household_roles.py:70 `MAX_ARTIFACT_BYTES = 64 * 1024**2`, enforced :269-276 `require(... and len(value.payload) <= MAX_ARTIFACT_BYTES, \"ARTIFACT_TYPE\")` over node.artifact_inputs, which include this node's own projection per :266-267 `_expected_artifacts`." + }, + { + "constant": "current_survey_geography.py:36 MAX_HOUSEHOLDS \u2014 at HEAD (3aa1802a8) its value is 7_000_000, NOT the census row's `64 * 1024**2 // 128` = 524,288. That expression is the pre-move value, which survived only through commit 907e308c3 (the snapshot this verifier session started at); commit 3aa1802a8 \"Apply the rule to a sixth ceiling the transport lane's census did not reach\" replaced it. Every line number in the census row is off by 7 at HEAD.", + "binds_at_full_source_verdict": false, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "Two stale facts, one arithmetic error, one missed reference. Everything else verifies.\n\n1. VALUE IS STALE (decisive). current_survey_geography.py:36 at HEAD reads `MAX_HOUSEHOLDS = 7_000_000`, preceded by a :29-35 comment deriving it as 4 x 1,587,376 rounded up. `git show HEAD:...` and the working file agree (identical md5 d506fd32...; tree clean). The census row's 524,288 is the value at 907e308c3 and earlier; `git log -L 25,32` shows the constant was introduced as `64 * 1024**2 // 128` in 0760c09d9 and changed exactly once, in 3aa1802a8 (Sep 17 18:01). Note the Read tool served me the pre-3aa1802a8 text for this file; only `sed`/`git show` gave HEAD. Anyone re-verifying should not trust a cached read here.\n\n2. \"BINDS AT FULL SOURCE: TRUE\" IS THEREFORE FALSE AT HEAD. 1,587,376 < 7,000,000, with the lane's 4x headroom. It WAS true before 3aa1802a8 (524,288 < 1,587,376; first bound at fraction 524,288/1,587,376 = 0.330, matching the census row's 0.33), so the census row's analysis was correct about the code it was written against, and this lane has already acted on it. packages/microcosm-build/tests/test_us_native_row_ceilings.py:55 now lists `(current_survey_geography, \"MAX_HOUSEHOLDS\", STACKED_HOUSEHOLDS, 7_000_000)` among MOVED and asserts both the rule and full-source admission (:70-73, :86-88).\n\n3. UPSTREAM-BYTES ARITHMETIC IS WRONG (conclusion unaffected). The row says the origins roster is \"about 190 MB at full source\" against MAX_ROSTER_BYTES. `_bounded_append` (survey_population_preparation.py:1145-1152) budgets one encoded record per source household AND per person, and that module's own comment at :51-53 measures a full-source whole-roster document at 1.02 GiB. Still far under MAX_ROSTER_BYTES = 64 * MAX_SEGMENT_BYTES = 4 GiB (:57-58), so \"no tighter upstream preempts\" holds \u2014 but do not quote 190 MB.\n\n4. ONE CONSUMER REFERENCE THE ROW MISSED. survey_origin_budget.py:214 places `geography.observed.MAX_HOUSEHOLDS` inside the `geography_contract` runtime marker built by `_live()` and compared at :257 (`_require(_live() == _LIVE and _code_bytes() == _BYTES, \"PRODUCER_CHANGED\")`), with `_BYTES`/`_LIVE` both computed at import (:1350-1351). `_runtime_marker` (survey_population_preparation.py:607-621) returns in-memory type/value tuples, not a persisted digest, so both sides move together and no pin moves. I confirmed no sha256 of current_survey_geography.py appears in any us_runtime/*.json (including graph_implementation_inventory.json) or in acs_native_coverage_binding._ACCEPTED (:34-39). The row's substantive claim (\"no consumer re-derives this number\") survives; its implied \"nothing references it\" does not.\n\nCONFIRMED UNCHANGED: enforcement sites, refusal codes, exception type, and counted quantity (see evidence). The counted quantity is emphatically the selected/stacked pre-clone roster, not an upstream file and not the clone: the pilot artifact's own origins list is 1,584 rows = 1,529 acs + 55 asec at selection.fraction [1,1000] with selection.supplied_households = 1,587,376, and the lane's measured split (ACS 1,531,614 + ASEC 55,762) reconciles to that total exactly. The ASEC side is not a fixed ~90k ASEC-source count here: it is sampled with the same fraction (55 at 1/1000, 55,762 at full), so no part of this bound is a structural assertion about an upstream file's size.", + "safe_to_move": true, + "safe_to_move_reason": "Safe \u2014 and already moved at HEAD. Neither (a) nor (b) applies. (a) No fixed-width encoding or reader byte budget depends on it: `_projection_digest` (:194-205) streams one row at a time, `source._encode(values, maximum=4096)` at :202 with a 4-byte per-row length prefix at :203 into hashlib.sha256 \u2014 both the per-row cap and the prefix width are constants independent of MAX_HOUSEHOLDS, so the digest of any accepted projection is byte-identical before and after the change. The module's only byte budget is MAX_RECEIPT_BYTES = 64 * 1024 (:37), applied at :289 to a receipt (:262-290) that carries aggregate counts only (\"households\", \"acs_households\", \"asec_households\", two aggregate sums) \u2014 no per-household payload, and nothing there is sized from MAX_HOUSEHOLDS. The old `64 * 1024**2 // 128` form was a borrowed budget with no reader behind it. (b) No upstream-file assertion: the bound is on len(origins) and len(household) of the sampled roster, which scales with the selection fraction; the real upstream-file bounds live elsewhere (acs_person_coverage_columns.MAX_ROWS = 6_000_000 at :39, asec_current_money.MAX_HOUSEHOLDS = 400_000) and are untouched. (c) No tighter upstream on the same quantity: domains.MAX_HOUSEHOLDS = 100_000 (survey_population_domains.py:19) is per batch at :551 \"HOUSEHOLD_BATCH\", MAX_TOTAL_MEMBERS = 1_000_000 per batch at :563 \"BATCH_MEMBER_BOUND\", and survey_catalogue_selection.py:127-132 pins _BATCH_HOUSEHOLDS = 10_000 / _BATCH_PEOPLE = 100_000 at or under those; the whole-roster path is bounded only in bytes (ORIGIN_LIMIT :1149, ROSTER_LIMIT :259). One caveat worth recording: the commit message's \"materialises no per-household payload\" is loose about the module as a whole \u2014 `_project` returns a per-row tuple (:185) and :250-260 builds a DataFrame of three python-storage StringDtype columns (dtype_for_token pins storage=\"python\", microcosm-graph/population.py:70-73), so a full-source run holds on the order of hundreds of MB of Python objects. That is real memory, which is exactly the resource class (a) this ceiling governs; it changes nothing about format or correctness.", + "evidence": "All paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings. HEAD = 3aa1802a88d80672944cb6906c07ce6486163dce (one commit past this session's snapshot 907e308c3); working tree clean, file md5 == HEAD blob md5.\n\npackages/microcosm-build/src/microcosm/build/us_runtime/current_survey_geography.py\n :29-35 comment deriving the ceiling from the 1,587,376 full-source selected roster\n :36 MAX_HOUSEHOLDS = 7_000_000 <-- HEAD value, refutes the census row's 524,288\n :37 MAX_RECEIPT_BYTES = 64 * 1024\n :40-42 def _require(condition, reason): raise ValueError(\"CURRENT_SURVEY_GEOGRAPHY_\" + reason) [exception type: plain ValueError; module defines no custom error class]\n :71-77 _project: `type(origins) is list and 0 < len(origins) <= MAX_HOUSEHOLDS and len(origins) == len(households)` -> \"HOUSEHOLD_COUNT\" (:77)\n :87 _require(set(households[channel_column]) == {\"acs\", \"asec\"}, \"CHANNEL_ROSTER\") [both channels in one roster]\n :181-184 COMPLETE_NATIVE_ROSTER: used_acs == set(acs.household_id) and used_asec == set(asec.household_id)\n :194-197 _projection_digest: `0 < len(household) <= MAX_HOUSEHOLDS` -> \"PROJECTION_STORAGE\" (:197)\n :202-204 source._encode(values, maximum=4096); digest.update(len(encoded).to_bytes(4, \"big\")) \u2014 per-row, streamed\n :235 origins = json.loads(entry[1])[\"origins\"][\"households\"]\n :246 state.frame.table(\"household\") (also :311 in the final re-check)\n :262-290 receipt: aggregate counts only, encoded with maximum=MAX_RECEIPT_BYTES\n\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_population_preparation.py\n :51-53 comment: full-source whole-roster document = 1.02 GiB; :57-58 MAX_SEGMENT_BYTES = MAX_PAYLOAD_BYTES = 64 MiB, MAX_ROSTER_BYTES = 64 * that = 4 GiB\n :1145-1152 _bounded_append -> _require(budget[0] + size <= MAX_ROSTER_BYTES, \"ORIGIN_LIMIT\"); one record per household AND per person\n :259 _require(len(payload) == total <= maximum, \"ROSTER_LIMIT\"); :607-621 _runtime_marker\n\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_population_domains.py:19, :551, :563 (100_000 / 1_000_000, both per-batch: \"HOUSEHOLD_BATCH\", \"BATCH_MEMBER_BOUND\")\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_catalogue_selection.py:127-132 (_BATCH_HOUSEHOLDS 10_000, _BATCH_PEOPLE 100_000, asserted <= the domains bounds)\npackages/microcosm-build/src/microcosm/build/us_runtime/survey_origin_budget.py:199-215, :257, :1350-1351 (geography.observed.MAX_HOUSEHOLDS inside the recomputed producer seal)\npackages/microcosm-build/src/microcosm/build/us_runtime/acs_native_coverage_binding.py:34-39 (_ACCEPTED: four ACS files; current_survey_geography.py absent)\npackages/microcosm-graph/src/microcosm/graph/population.py:70-73 (StringDtype storage=\"python\")\n\npackages/microcosm-build/tests/test_us_native_row_ceilings.py:41-56 (ACS_HOUSEHOLDS 1_531_614 + ASEC_HOUSEHOLDS 55_762 = STACKED_HOUSEHOLDS 1_587_376; MOVED includes current_survey_geography.MAX_HOUSEHOLDS -> 7_000_000), :70-73, :86-88\npackages/microcosm-build/tests/test_us_current_survey_geography.py:194-215 (added by 3aa1802a8) \u2014 drives ValueError match=\"PROJECTION_STORAGE\" at a patched-down ceiling\n\nPilot artifact /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json (sha256 34b362d8... confirmed): origins.households = 1584 rows, Counter{'acs': 1529, 'asec': 55}; selection.fraction [1,1000], selection.supplied_households 1587376; ACS cells 84,422 + 98,784 + 1,665 + 1,346,743 = 1,531,614 eligible, ASEC shared_housing 55,697 eligible." + }, + { + "constant": "puf_monetary_source.py:338 _MAX_AMOUNT_DIGITS = 12", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": true, + "correction": "No factual error in any asserted field: value, enforcement site, refusal string, exception type, counted quantity, scaling behavior and binds-at-full are all correct as written. Three refinements.\n\n(1) The census reasoning (\"Included only because the name contains MAX and it is a character ceiling\") understates it. The twelve-digit ceiling is the stated premise of the out-of-universe sentinel. puf_monetary_source.py:239-244 \u2014 \"minimum int64 is unreachable for a token of at most twelve digits, so no individual return can legitimately hold it and readback can assert it exactly. AGGREGATE_SENTINEL = -(2**63)\" \u2014 and puf_monetary_source.py:1602-1613 relies on that unreachability (\"The sentinel is unreachable for that grammar, so a sentinel parked on an ordinary row is caught here as a lexical disagreement\" / raise _refuse(\"READBACK_AGGREGATE_SLOT\", column.field)). The per-value int64 gate at :884 admits all of [-(2**63), 2**63-1], so only the digit ceiling at :879-880 excludes the sentinel from the legitimate token space. The row should not be filed as incidental.\n\n(2) The per-value re-check runs at TWO sites, not the one the row lists: :879-880 raise _refuse(\"AMOUNT_TOKEN_OVER_DIGITS\", label) in _amount, and :927-928 raise _refuse(\"AGGREGATE_TOKEN_OVER_DIGITS\", label) in _aggregate_token. Both test len(digits) > column.max_digits.\n\n(3) The \"protects\" prose says the check validates field_width_digits \"against the publisher's delivered grammar\". Strictly, :515 compares the projection document's declaration against the module constant and against the declared lexical_width; nothing is read from the delivered file at that point. The delivered file's grammar is enforced per token later at :879-880 and :927-928.\n\nAdjacent and out of scope, noted for the lane: :1179 \"if not _INT64_MIN <= total <= _INT64_MAX: raise _refuse(\"COLUMN_SUM_OVERFLOW\")\" in _column_facts IS row-scale dependent (a per-column sum over PUF return rows). Not audited here.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and not for a memory/time reason. Two independent blocks.\n\n(a) It is a structural declaration of the publisher's record layout, not a resource ceiling. The comment at :337 states the PUF amount fields are twelve digits wide; the check at :515-516 refuses any projection document that declares more, plus a lexical allocation with no room for the sign (width < digits + 1). The packaged documents are exactly tight at that second arm (lexical_width 13, field_width_digits 12 on every column of both puf_2015_monetary_source_projection.json and puf_2015_monetary_agi_source_projection.json). Raising the constant would admit a projection document asserting a wider publisher field than the retained booklets support \u2014 weakening a real assertion about the delivered file's grammar, with no scale benefit whatsoever since the quantity is invariant in source size.\n\n(b) A downstream correctness argument depends on the specific value 12. AGGREGATE_SENTINEL = -(2**63) at :244 is justified at :239-243 as \"unreachable for a token of at most twelve digits\". A 12-digit token maxes at 999,999,999,999, four orders of magnitude below 9.22e18. The per-value int64 gate at :884 permits the full int64 range, so the digit ceiling is the only thing keeping -(2**63) out of the legitimate amount space. Move it to 19 and a delivered amount token could parse to exactly AGGREGATE_SENTINEL, making the typed slot ambiguous between a real amount and the out-of-universe marker, and silently defeating the readback assertions at :1605-1613 (READBACK_LEXICAL_DISAGREEMENT and READBACK_AGGREGATE_SLOT) and the universe split at :955-960.\n\nAlso: moving it buys nothing. The quantity is per-token digit count, invariant at 12 across 1/1000, 1/10 and the full 1,587,376-household source. It is not on any roster path \u2014 not the selected/stacked survey, not the cloned roster, and not an ASEC or ACS row count.", + "evidence": "All paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings, at HEAD 3fb1111e282602fff3cbb334a5d49b2f9a912cc4.\n\nCONSTANT\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_monetary_source.py:337 \u2014 \"#: The publisher's amount fields are twelve digits wide.\"\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_monetary_source.py:338 \u2014 \"_MAX_AMOUNT_DIGITS = 12\"\n\nENFORCEMENT + REFUSAL CODE + EXCEPTION TYPE\npuf_monetary_source.py:514 \u2014 digits = _positive_int(units[\"field_width_digits\"], f\"{field}.field_width_digits\")\npuf_monetary_source.py:515 \u2014 if digits > _MAX_AMOUNT_DIGITS or width < digits + 1:\npuf_monetary_source.py:516 \u2014 raise _refuse(\"PROJECTION_WIDTH_BELOW_GRAMMAR\", field)\npuf_monetary_source.py:341 \u2014 class PufMonetaryRefusalError(ValueError):\npuf_monetary_source.py:349-350 \u2014 def _refuse(reason, *detail) -> PufMonetaryRefusalError: return PufMonetaryRefusalError(reason, *detail)\npuf_monetary_source.py:559 \u2014 max_digits=digits (the declaration carried onto ProjectedColumn, declared at :371)\nRepo-wide grep for \"_MAX_AMOUNT_DIGITS\" and \"PROJECTION_WIDTH_BELOW_GRAMMAR\" returns only lines 338, 515, 516. No test references either.\n\nPER-VALUE RE-CHECKS (two, not one)\npuf_monetary_source.py:879-880 \u2014 if len(digits) > column.max_digits: raise _refuse(\"AMOUNT_TOKEN_OVER_DIGITS\", label)\npuf_monetary_source.py:927-928 \u2014 if len(digits) > column.max_digits: raise _refuse(\"AGGREGATE_TOKEN_OVER_DIGITS\", label)\npuf_monetary_source.py:884 \u2014 if not _INT64_MIN <= number <= _INT64_MAX: raise _refuse(\"AMOUNT_TOKEN_INT64_OVERFLOW\", label)\n\nSENTINEL DEPENDENCE (why it must not move)\npuf_monetary_source.py:239-244 \u2014 sentinel rationale and AGGREGATE_SENTINEL = -(2**63)\npuf_monetary_source.py:955-960 \u2014 _amount where known, AGGREGATE_SENTINEL where not\npuf_monetary_source.py:1602-1613 \u2014 readback: sentinel \"unreachable for that grammar\"; raise _refuse(\"READBACK_LEXICAL_DISAGREEMENT\", ...) / raise _refuse(\"READBACK_AGGREGATE_SLOT\", column.field)\npuf_monetary_source.py:334-335 \u2014 _INT64_MAX / _INT64_MIN\n\nCOUNTED QUANTITY (measured, not guessed)\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_2015_monetary_source_projection.json \u2014 projection.columns: 12 columns, every one (lexical_width=13, units.field_width_digits=12, dtype=int64)\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_2015_monetary_agi_source_projection.json \u2014 projection.columns: 13 columns, every one (lexical_width=13, units.field_width_digits=12, dtype=int64)\nThese two files plus puf_monetary_source.py are the only files in the repo containing \"field_width_digits\".\n\nNO TIGHTER UPSTREAM BOUND ON DIGITS\npuf_monetary_source.py:187 \u2014 LEXICAL_WIDTH_MAX = 64; :511 \u2014 if width > LEXICAL_WIDTH_MAX: raise _refuse(\"PROJECTION_LEXICAL_WIDTH\", field)\npuf_raw_source.py:185 \u2014 _LEXICAL_WIDTH_MAX = 64; :382 \u2014 not 0 < width <= 64 check on the raw definition\nBoth are character-width bounds at 64, looser and on a different quantity than the 12-digit ceiling.\n\nADJACENT SCALE-DEPENDENT BOUND (not audited)\npuf_monetary_source.py:1178-1180 \u2014 total = int(sum(...)); if not _INT64_MIN <= total <= _INT64_MAX: raise _refuse(\"COLUMN_SUM_OVERFLOW\")\n\nNOT ESTABLISHED / NOT APPLICABLE\nThis bound is not on any roster, so no ASEC or ACS upstream file size is relevant to it and I did not attempt to establish one. I also did not attempt to establish the PUF's own return-row count, because no arm of this constant counts rows." + }, + { + "constant": "puf_raw_source.py:185 `_LEXICAL_WIDTH_MAX = 64` (literal, confirmed at HEAD 3fb1111e2)", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The substantive verdict is right (value 64; refusal string \"ENVELOPE_COLUMN_WIDTH\"; exception PufRawSourceRefusalError, a ValueError subclass at :234; counted quantity = bytes per lexical cell; not a roster bound; binds at full source = false). I disagree on three points of fact in the row.\n\n(1) LINE CITATIONS ARE OFF BY ONE in two places. The exact-width mismatch raise is at :1275 (condition at :1274), not \":1276\". And `expected = rows * width` is at :1276, not \":1277\" (:1277 is `sizes[field] = width`); the size-equality raise is correctly at :1281 (condition :1280). The primary site :1272-1273 is right.\n\n(2) THE REASONING ATTRIBUTES THE REFUSAL TO THE WRONG SITE for production. The row says \"the widest declared field (RECID, 64) sits exactly at the ceiling, so a wider lexical field would refuse\". A wider field would indeed refuse, but NOT via `_LEXICAL_WIDTH_MAX`: a packaged definition declaring width > 64 is refused first at :382-383, `if not isinstance(width, int) or isinstance(width, bool) or not 0 < width <= 64: raise _refuse(\"DEFINITION_LEXICAL_WIDTH\", field[\"name\"])` \u2014 a SEPARATE HARDCODED literal 64 that does not reference the constant at all. The comment at :183-184 asserts the two \"match\", but nothing in the code ties them; they are two independent literals.\n\n(3) THERE IS A STRICTLY TIGHTER CHECK THAT MAKES :1272-1273 UNREACHABLE ON THE PRODUCTION PATH (the row does not mention it). In `_envelope_columns` (:1234-1238), `widths = _widths(definition) if definition is not None else None` (:1250). When a definition resolves, :1274-1275 demands `width == widths[field]` exactly \u2014 strictly tighter than `width > 64`. `_resolve_envelope_definition` (:1177-1204) returns None ONLY for `definition_route == \"test_fixture\"` (:1199-1204); a \"packaged\"-route payload always resolves against `packaged_definition()`. So on every production decode, :1272-1273 is dead code behind the exact-equality check; it binds only on fixture-route payloads. The row presents it as the operative ceiling for RECID, which it is not.\n\nTwo further facts the row should carry: the lower bound is `_envelope_int(entry[\"width\"], f\"{name}.width\", minimum=1)` at :1271; and this module's row dimension is not a survey roster at all \u2014 the packaged definition pins the upstream 2015 PUF file's true size (artifact rows 207,696; sources.main data_records 207,696 / 126,034,649 bytes; sources.demographic data_records 119,675), checked at :846-857 / :1469. So nothing in this module touches the 1,587,376-household selected/stacked/cloned roster, confirming \"binds at full: false\" for a stronger reason than the row gives.", + "safe_to_move": false, + "safe_to_move_reason": "Category (b), a fixed-width encoding parameter, and it must not be moved by this lane.\n\n(i) It is not a scale bound in any sense. Widths come from a static packaged JSON (puf_2015_raw_source_definition.json: RECID 64, S006 32, FLPDYR 8, and 4 for FLPDMO/MARS/DSI/AGEDP1-3/AGERANGE/EARNSPLIT/GENDER \u2014 I read the file), invariant at 1/1000, 1/10 and full source. Moving it buys the row-ceiling lane exactly nothing.\n\n(ii) Moving it changes a wire format's admissible parameter space. The lexical body is fixed-width NUL-padded ASCII: written as `bytearray(len(values) * width)` with `body[start:start+len(raw)]` (_fixed_ascii, :1033-1043, refusing \"LEXICAL_OVER_WIDTH\" at :1038) and read back by byte-slicing `body[position * width : (position + 1) * width]` (_read_fixed_ascii, :1046-1059, called at :1442). The per-column byte budget a reader depends on is `expected = rows * width` (:1276), enforced against the declared `bytes` at :1280-1281 (\"ENVELOPE_COLUMN_SIZE\"), and the running total is bounded by BODY_MAX_BYTES (64 MiB) at :1283-1284 (\"ENVELOPE_BODY_OVER_BOUND\"). A different width is a different stored artifact, as the row says.\n\n(iii) Raising it desynchronizes the two independent literals. :382 hardcodes `0 < width <= 64` for the definition document; :185 is a second copy. Raising one without the other leaves the comment at :183-184 asserting a match the code no longer has, and raising both would relax a real schema bound.\n\n(iv) It is the ONLY width cap on the decode path where no definition is resolvable (fixture route, :1204 returns None, so `widths is None` and :1274 is skipped). Raising it removes the last width guard on that path.\n\nNot (c): it does not assert an upstream file's row count. That assertion lives elsewhere in the same module (the pinned 207,696 / 119,675 record counts and byte pins), and those must not be moved either.", + "evidence": "All paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/packages/microcosm-build/src/microcosm/build/us_runtime/ (HEAD 3fb1111e2, read-only; nothing edited).\n\npuf_raw_source.py:183-184 \u2014 comment \"The widest lexical allocation any field may declare, matching the closed definition's own ``0 < lexical_width <= 64`` bound.\"\npuf_raw_source.py:185 \u2014 `_LEXICAL_WIDTH_MAX = 64` (only two references repo-wide: this line and :1272)\npuf_raw_source.py:234-240 \u2014 `class PufRawSourceRefusalError(ValueError)` with `self.reason = reason`\npuf_raw_source.py:242-243 \u2014 `def _refuse(reason, *detail) -> PufRawSourceRefusalError: return PufRawSourceRefusalError(reason, *detail)`\npuf_raw_source.py:381-383 \u2014 `width = field[\"lexical_width\"]` / `if not isinstance(width, int) or isinstance(width, bool) or not 0 < width <= 64:` / `raise _refuse(\"DEFINITION_LEXICAL_WIDTH\", field[\"name\"])` (separate hardcoded 64)\npuf_raw_source.py:678-681 \u2014 `_widths(definition)` builds `{field[\"name\"]: field[\"lexical_width\"]}`\npuf_raw_source.py:846-857 \u2014 `_check_declared_expectations` cross-checks the definition's declared record counts (packaged route only)\npuf_raw_source.py:1033-1043 \u2014 `_fixed_ascii`: `bytearray(len(values) * width)`, `raise _refuse(\"LEXICAL_OVER_WIDTH\", label)` at :1038\npuf_raw_source.py:1046-1059 \u2014 `_read_fixed_ascii`: `cell = body[position * width : (position + 1) * width]` at :1051\npuf_raw_source.py:1062-1072 \u2014 `_payload_bound`: `rows * widths[name]` per lexical column\npuf_raw_source.py:1088, 121 \u2014 `if total > BODY_MAX_BYTES` / `BODY_MAX_BYTES = 64 * 1024 * 1024`\npuf_raw_source.py:1177-1204 \u2014 `_resolve_envelope_definition`; route must be \"packaged\" or \"test_fixture\" (:1190-1191); packaged always resolves (:1199-1203); `return None` only for the fixture route (:1204)\npuf_raw_source.py:1234-1238 \u2014 `def _envelope_columns(header, rows, definition)`\npuf_raw_source.py:1250 \u2014 `widths = _widths(definition) if definition is not None else None`\npuf_raw_source.py:1271 \u2014 `width = _envelope_int(entry[\"width\"], f\"{name}.width\", minimum=1)`\npuf_raw_source.py:1272-1273 \u2014 `if width > _LEXICAL_WIDTH_MAX:` / `raise _refuse(\"ENVELOPE_COLUMN_WIDTH\", name)`\npuf_raw_source.py:1274-1275 \u2014 `if widths is not None and width != widths[field]:` / `raise _refuse(\"ENVELOPE_COLUMN_WIDTH\", name)` [row cited :1276]\npuf_raw_source.py:1276-1277 \u2014 `expected = rows * width` / `sizes[field] = width` [row cited :1277 for the product]\npuf_raw_source.py:1280-1281 \u2014 `if size != expected:` / `raise _refuse(\"ENVELOPE_COLUMN_SIZE\", name)`\npuf_raw_source.py:1283-1284 \u2014 `if total > BODY_MAX_BYTES:` / `raise _refuse(\"ENVELOPE_BODY_OVER_BOUND\", total)`\npuf_raw_source.py:1372-1374 \u2014 `def decode_return_status(payload, definition=None)`\npuf_raw_source.py:1422-1424 \u2014 `resolved = _resolve_envelope_definition(header, definition)` then `_envelope_columns(header, rows, resolved)`\npuf_raw_source.py:1442 \u2014 `lexical[field] = _read_fixed_ascii(body, rows, column[\"width\"], field)`\npuf_2015_raw_source_definition.json \u2014 parsed with python3 json: fields/lexical_width = RECID 64, FLPDYR 8, FLPDMO 4, MARS 4, DSI 4, S006 32, AGEDP1 4, AGEDP2 4, AGEDP3 4, AGERANGE 4, EARNSPLIT 4, GENDER 4 (matches the row exactly); artifact.rows 207696; sources.main.data_records 207696 / bytes 126034649; sources.demographic.data_records 119675 / bytes 2225050.\n\nRepo-wide grep for `_LEXICAL_WIDTH_MAX`: only puf_raw_source.py:185 and :1272. Grep for \"ENVELOPE_COLUMN_WIDTH\": puf_raw_source.py:1273, :1275 and a separate occurrence in puf_monetary_source.py:1551 (a different module, not this row).\n\nNot established: no test file under packages/microcosm-build/tests references `PufRawSourceRefusalError`, `ENVELOPE_COLUMN_WIDTH` or `DEFINITION_LEXICAL_WIDTH` (grep -rl over packages/*/tests returned nothing), so I could not confirm an enforced test pinning the 64/64 mirror. I did not run the suite (read-only census)." + }, + { + "constant": "graph_composed_asec_binding.py:956 ARM_ROWS_PAYLOAD_MAX_BYTES = 22,465,581 bytes (9 + 4 + 65,536 + 16 x (1,000,000 + 400,000) + 32)", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The headline verdicts all survive my read (value 22,465,581; counted quantity = ASEC-arm rows of the composed population at 16 B/person-row + 16 B/household-row + 45 B framing; binds-at-full = false). Two things the census got wrong or missed:\n\n(1) WRONG LINE FOR A REFUSAL CODE. The row says the literal \"ARM_ROWS_SIZE\" appears at \":1033 and :1052\". At graph_composed_asec_binding.py:1052 the literal is \"ARM_ROWS_MAGIC\" (`_require(payload.startswith(ARM_ROWS_MAGIC), \"ARM_ROWS_MAGIC\")`). The reader's second \"ARM_ROWS_SIZE\" literal is at :1050, inside the `_require(` spanning :1047-1051 whose condition is on :1048-1049. The enforcement-site citations (:1033, :1049) are correct; only the refusal-code line :1052 is misattributed. Exception type is right: _require at :236-238 raises PreparedGraphError, defined `class PreparedGraphError(ValueError)` at graph_asec_prepared.py:145.\n\n(2) MISSED THE TIGHTER UPSTREAM CAPS (check #3 answered wrongly by omission). The person arm roster cannot reach the 1,000,000 this budget provisions for. arm_identity[\"person\"] is income.array(\"person_id\") (:646) and ARM_IDENTITY_ROWS (:649-653) forces len == receipt[\"entity_rows\"][\"person\"], which bind_income_observations pins to the artifact's declared rows with `0 < rows <= source._MAX_PERSONS` \u2014 _MAX_PERSONS = 600_000 (asec_income_observations.py:44), enforced at graph_asec_income.py:175-181 (\"BOUND_INCOME_ROWS\") and asec_income_observations.py:324 / :377 (\"INCOME_ROWS\"). The household side is capped at 400,000 (asec_housing_universe.py:41, \"ROW_COUNT\" at :147 and :241) \u2014 equal to, not tighter than, the budget's assumption. So the maximum payload the pipeline can construct is 9+4+65,536+16x(600,000+400,000)+32 = 16,065,581 bytes, i.e. the producer-side check at :1033 is structurally unreachable with ~6.4 MB of permanently dead headroom. That makes binds-at-full = false far more robust than \"no scale I can construct\", but the census row never establishes it.\n\n(3) NUANCE the row half-concedes and should state outright: the reader does NOT derive any offset or width from this constant. Buffer widths come from `rows[entity] * 8` where rows is read from the BINDING DOCUMENT (:1075-1078, :1093), the cursor walks those widths (:1089-1105), and the exact total is asserted by `cursor == len(payload) - 32` (\"ARM_ROWS_LENGTH\", :1106). The constant's only role is a pre-parse admission ceiling on an untrusted payload. It is category (b) by the rubric's own wording (\"a byte budget computed as HEADER + ROW_BYTES * MAX_N that a reader relies on\"), and I classify it that way, but its force is not that raising it re-formats the wire \u2014 it is that the number has no independent existence apart from MAX_PERSONS/MAX_HOUSEHOLDS.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and pointless. Three reasons, all read at HEAD:\n\n(a) NOTHING TO GAIN. It never binds. The ASEC arm is hard-capped at the prepared ASEC source: `int(masks[entity].sum()) <= int(receipt[\"entity_rows\"][entity])` refuses \"ARM_OVERFLOW:{entity}\" (:636-639), and _resolved_positions (:306-332) refuses duplicated (\"BINDING_AMBIGUOUS_ROWS\"), unresolved (\"BINDING_UNRESOLVED\") and non-monotone (\"BINDING_NON_MONOTONE\") mappings, so arm rows are a strictly increasing subset of source rows. Full-source single cohort = 45 + header + 16x(142,125 + 55,762) = 3,166,237 B (14% of budget); a three-cohort arm at the census's figures = 9,592,413 B (42.7%); the structural maximum upstream permits = 16,065,581 B (71.5%). Nothing reaches 22,465,581.\n\n(b) YOU CANNOT MOVE IT IN ISOLATION. It is defined as a formula over MAX_PERSONS and MAX_HOUSEHOLDS imported from asec_current_money (:77-81). Those same two constants are an UPSTREAM-FILE-SIZE assertion on the real authenticated ASEC source \u2014 `0 < data[\"person_rows\"] <= MAX_PERSONS and ... <= MAX_HOUSEHOLDS` refusing \"HEADER_ROW_BOUND\" (asec_current_money.py:800-806), plus \"SOURCE_SIZE\" at :521 \u2014 and they also define a second payload budget at _asec_current_money_codec.py:26-28. Raising them to move this ceiling weakens a real check on the CPS ASEC file's true size and silently widens an unrelated budget. Hard-coding a decoupled larger number instead severs the ceiling from the roster caps it is supposed to express.\n\n(c) LOWERING IT IS A LIVE BREAK. The reader admits transported payloads against it at :1047-1051 before parsing; any reduction below an actual arm payload refuses valid artifacts with \"ARM_ROWS_SIZE\".\n\nUpstream file size, stated explicitly as required: this bound is on an ASEC-source-scale roster, NOT the selected/stacked survey (1,587,376 source households) and NOT the cloned roster. The real current-year ASEC source is 142,125 person rows (pinned rows=142_125 for survey_year 2025 / income_year 2024 at education_assistance_source.py:161, and /catalogues/asec/counts/persons in the pilot artifact) and 55,762 households (/catalogues/asec/counts/households, pilot artifact sha256 34b362d8...). Prior-year person rows are pinned at 146,133 (:117) and 144,265 (:139); the 2022/2023 ASEC household counts are NOT ESTABLISHED \u2014 I searched every .py and .json under us_runtime for those figures and found only per-year person rows. ACS growth cannot push this bound.", + "evidence": "All paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/packages/microcosm-build/src/microcosm/build/us_runtime (HEAD 907e308c3ec2ec1b0ba7b94a35211484d9418d7a):\n\nCONSTANT: graph_composed_asec_binding.py:956-962 (ARM_ROWS_PAYLOAD_MAX_BYTES formula); :937 ARM_ROWS_MAGIC = b\"MCCASECR\\x01\" (9 bytes); asec_current_money.py:47-49 (HEADER_MAX_BYTES = 64*1024, MAX_PERSONS = 1_000_000, MAX_HOUSEHOLDS = 400_000); imported at graph_composed_asec_binding.py:77-81. Arithmetic verified: 9+4+65536+16*1_400_000+32 = 22,465,581.\n\nENFORCEMENT + REFUSAL: graph_composed_asec_binding.py:1033 `_require(len(payload) + 32 <= ARM_ROWS_PAYLOAD_MAX_BYTES, \"ARM_ROWS_SIZE\")`; :1047-1051 `_require(type(payload) is bytes and len(ARM_ROWS_MAGIC) + 4 + 32 < len(payload) <= ARM_ROWS_PAYLOAD_MAX_BYTES, \"ARM_ROWS_SIZE\")` \u2014 literal on :1050, not :1052 (:1052 is \"ARM_ROWS_MAGIC\"). _require at :236-238 raises PreparedGraphError; graph_asec_prepared.py:145 `class PreparedGraphError(ValueError)`.\n\nWIRE FORMAT: :939-944 _ARM_ROWS_BUFFERS = 4 tuples (person/household x source_positions/original_ids); :990-994 buffers built with dtype \" multiplier 10**7; max remapped = 10**7 + 3,471,383 = 13,471,383 (~1.35e7, not the claimed ~1.7e7)\n 1/10: stacked persons ~347,138 -> 10**6 + 347,138 = 1,347,138 (~1.35e6, not the claimed ~1.7e6)\n 1/100: ~34,714 -> 10**5 + 34,714 = 134,714\n 1/1000 (the measured pilot, 3,464 person rows): 10**4 + 3,464 = 13,464\nHeadroom at full source is 11.84 orders of magnitude, not \"twelve\".\n\n(2) MINOR \u2014 enforcement line range. The census cites 2359-2368. Line 2359 is `return values.copy()` (the clone_index==0 short-circuit) and 2360 computes the shift. The actual guard is puf_support.py:2361 and the raise runs 2362-2368.\n\n(3) UNDERSTATED \u2014 the bound is provably unreachable, not merely far away. clone_index can only ever be 0 or 1: `_DEFAULT_SUPPORT_CHANNELS` is exactly two channels (puf_support.py:158-161), both entry paths reject anything else (puf_support.py:571-575 for assembled frames, :2309-2313 in `_reject_metadata_collisions`), `PUF_TAX_DETAIL_CLONE_INDEX = 1` (support_provenance.py:39), and clone_index==0 short-circuits before the check (:2358-2359). With clone_index=1 and multiplier=10**digits(max_id), overflow requires max_id >= 10**18. The upstream cap PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID = 10**15 - 1 (operator_column_contracts.py:64), enforced at spine_assembly.py:630-641 and again after collision remapping at spine_assembly.py:763-767, holds max_id to 10**15-1, so the worst assembled case is 10**15 + (10**15 - 1) = 1,999,999,999,999,999 \u2014 still 3.7 orders of magnitude inside int64. The census is right that this is the backstop and PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID is the operative bound.\n\n(4) CONFIRMED, no correction. Value 2**63-1 at :2331. Refusal is a bare `ValueError` (raise at :2362) with no short code \u2014 this module uses no require/_require helper here, unlike asec_person_coverage_source.py:146 (`_require(number <= _INT64_MAX, \"MEMBER_INTEGER\")`). Message text matches the census quote. protects=encoding-width is right and is the strongest form of it: `_validated_integral_ids` returns `numeric.astype(\"int64\")` (puf_support.py:2189, docstring :2176 \"Return int64 IDs only when every input value is exactly integral\"), so 2**63-1 is literally the storage width of the ID columns, not a tunable ceiling.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and not meaningful to move. This is not a resource ceiling \u2014 it is the numeric limit of the int64 dtype the structural ID columns are stored in (`_validated_integral_ids` casts to \"int64\" at puf_support.py:2189; spine_assembly.py:620-623 and :743-747 both refuse non-integer ID dtypes). Raising it re-admits exactly the failure the regression test documents: test_us_spine_assembly.py:412-416 records that int64-valid IDs above the shared bound \"assembled fine, then overflowed the clone stage's decimal remap (OverflowError: 'Python int too large to convert to C long')\", and test_us_spine_assembly.py:398-409 pins that `_remap_ids` must raise the governed ValueError instead. Lowering it would reject IDs that int64 stores fine. It also does not need to move: it never binds at any scale \u2014 13,471,383 at full source against 9.22e18, and structurally capped at ~2e15 by PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID with clone_index confined to {0,1}. The row-ceiling lane should leave it untouched; if anything is scale-relevant here it is PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID, and that is an ID-space bound (10**15-1), not a row count either.", + "evidence": "Worktree HEAD 3fb1111e282602fff3cbb334a5d49b2f9a912cc4. All paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/.\nCONSTANT: packages/microcosm-build/src/microcosm/build/us_runtime/puf_support.py:2331 (`_INT64_MAX = 2**63 - 1`); rationale comment :2324-2329.\nENFORCEMENT: puf_support.py:2361 (`if len(values) and shift + int(values.max()) > _INT64_MAX:`), raise at :2362-2368; shift computed :2360; clone_index==0 short-circuit :2358-2359; return :2369. Bare `ValueError`, no short code.\nMULTIPLIER: puf_support.py:2334-2345 (`return 10 ** max(1, len(str(max_id)))`), frame-level wrapper :2316-2321, computed once on the pre-clone frame at :578.\nPRE-CLONE CALL SITES (the refutation): puf_support.py:2060-2071 (`_clone_entity_table`, concat at :2071), :2095-2104 (`_clone_preassembled_entity_table`, concat at :2104), :2136-2145 (`_clone_link_table`, concat at :2145).\nCLONE INDEX CONFINED TO {0,1}: puf_support.py:158-161 (`_DEFAULT_SUPPORT_CHANNELS`), :571-575 and :2309-2313 (channels forced equal to it), support_provenance.py:39 (`PUF_TAX_DETAIL_CLONE_INDEX = 1`).\nTIGHTER UPSTREAM BOUND: operator_column_contracts.py:64 (`PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID = 10**15 - 1`); enforced spine_assembly.py:630-641 and :763-767.\nDTYPE: puf_support.py:2171-2194 (`_validated_integral_ids`, docstring :2176, cast :2189).\nTESTS: packages/microcosm-build/tests/test_us_spine_assembly.py:388-395, :398-409, :412-428, :431-444.\nArithmetic re-derived from this lane's ground-truth pilot scaling (stacked persons 3,471,383 full / 347,138 at 1/10 / 34,714 at 1/100 / 3,464 measured at 1/1000); I did not re-derive those roster counts. I did not establish actual ID density in a real spine (whether max_id equals row count); it does not change the verdict, since the multiplier is 10**digits(max_id), so even a max_id 1000x the row count leaves ~9 orders of magnitude of headroom." + }, + { + "constant": "current_child_property_income_source.py:34 MAX_ROWS = 600_000", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The headline verdicts survive scrutiny (value 600_000; four real sites; correct refusal codes and ValueError; the binding quantity is the STACKED person roster, not cloned and not ASEC; binds at full source; memory-or-time; movable). But the row has one material omission and several wrong citations.\n\nMATERIAL OMISSION \u2014 a paired ceiling on the identical quantity that the row never mentions. graph_child_property_income.py:80 `MAX_ORIGINALS = 600_000`, enforced at graph_child_property_income.py:305-309 `require(recipients.index.identical(origins.index) ... and 0 < len(recipients) <= MAX_ORIGINALS, \"ORIGINAL_RECIPIENT_ROSTER\")`, where `recipients = qualified.recipients` (:304) is the very table this module's :393 bounds (the graph builds it from `child.qualify_child_property_sources(preparation)` at :923). Same value, same roster, downstream. So the row's movability finding, while true of this constant in isolation, is operationally incomplete: raising MAX_ROWS alone moves nothing \u2014 a full-source run then refuses at CHILD_PROPERTY-adjacent graph code with \"ORIGINAL_RECIPIENT_ROSTER\" at the same 600,000. A third, looser ceiling on a recipient table sits further downstream at microcosm-fit/src/microcosm/fit/graph_joint_empirical.py:45 `MAX_RECIPIENTS = 1_048_576` (enforced :526, :610, \"RECIPIENT_COUNT\"); I did not establish whether that table is the full roster or the eligible subset, so I claim only its existence and value.\n\nWRONG CITATIONS (all checked by numbered read):\n- \"donor_scope / donor_sampling_fraction_applied (:636-637)\" is off by three: they are at :633 and :634. :636-637 are `weight_conversion` and `preparation_sha256`.\n- The RECIPIENT_AXIS code string is at :394, not :395 (:395 is the closing paren).\n- `project_child_property_donors` is called at :605-607, not :606 (:606 is the argument line).\n- `_BATCH_HOUSEHOLDS = 10_000` is defined at survey_catalogue_selection.py:19; the cited :129 is the self-check `0 < _BATCH_HOUSEHOLDS <= min(10_000, domains.MAX_HOUSEHOLDS)`.\n\nONE OVERSTATEMENT. \"MAX_ROWS must not be lowered below the real ASEC file sizes, because at :186/:203 it is guarding an upstream roster's true extent\" overstates :186. The true extent of the ASEC person frame is asserted exactly and independently upstream: current_asec_interest_source.py:413 `require(len(keys) == rows ...)` against the pinned `rows` from asec_coverage_authentication.py:34-37 `_MEMBER_PINS` / education_assistance_source.py:155-161 (pppub25.csv, rows=142_125, member_size_bytes=277_882_549). MAX_ROWS at :186 is a redundant 4.2x-headroom capacity cap, not the true-size check. :203 is different: I found no other whole-tuple length ceiling on the ASEC catalogue households tuple (asec_population_catalogue.py bounds only bytes: _MAX_ENVELOPE_BYTES :26, _MAX_RECORD_BYTES :27), so the row is right that :203 is the only whole-tuple ceiling there, and right that domains.MAX_HOUSEHOLDS = 100_000 (survey_population_domains.py:19) is per batch (:551 \"HOUSEHOLD_BATCH\" in classify_households).\n\nEVERYTHING ELSE I CONFIRMED, including the parts most likely to be wrong: origins[\"persons\"][\"rows\"] is one record per person in the stacked frame (survey_population_preparation.py:1293-1316, appended from `entities[\"person\"]`, which is built from `frame.table(entity)` at :1208-1216) \u2014 pre-clone, both channels, so 3,471,383 at full source > 600,000 and ~347,138 at 1/10; :186's frame is the complete ASEC member file (current_asec_interest_source.py:410-430, `ordered = raw.loc[keys]`, len == pinned rows 142,125), fraction-independent; the pilot artifact I re-read myself records /catalogues/asec/counts/households 55762 and /catalogues/asec/counts/persons 142125. No tighter upstream limit makes :393 unreachable: the only upstream roster ceiling is the byte budget MAX_ROSTER_BYTES = 64*64MiB = 4 GiB (survey_population_preparation.py:58, enforced :1149 \"ORIGIN_LIMIT\"), whose own comment at :1146-1147 puts a full-source roster at 1.02 GiB.", + "safe_to_move": true, + "safe_to_move_reason": "Safe as to encoding and as to upstream-file assertions, but it must not be moved alone. (a) No fixed-width encoding or byte budget depends on it: the module's only byte budget, MAX_PROJECTION_BYTES = 64*1024**2 (:35, enforced in _json at :47), is an independent constant not computed from MAX_ROWS; grep shows MAX_ROWS appears only at :34, :108, :186, :203, :217, :393. (b) Nothing serializes it. Its one non-enforcement use is :108, inside `_live()`'s \"child_property_contract\" via `source._runtime_marker` (survey_population_preparation.py:607-621), which returns in-memory `(type, value)` tuples and `(type, id(value))` for objects \u2014 a same-process identity comparison (`require(_live() == live, \"IMPLEMENTATION_CHANGED\")`), not a digest. `child_property_sources_seal` (:556-573) and the evidence dict (:625-650) both exclude it; no test references MAX_ROWS (test_us_child_property_income_source.py, ..._owner.py); the module is absent from graph_implementation_inventory.json; its file sha256 5e022be0aaea3e8727c506fb648448b8fbdcfa161895e18a3ad5bf03bc8ac9e3 appears in no committed pin. (c) It is not an upstream-file true-size assertion: 600,000 against actual 142,125 ASEC persons and 55,762 ASEC households is generic headroom, and the same 600_000 appears as boilerplate capacity in five sibling modules (asec_demographic_source.py:67, asec_income_observations.py:44, asec_2024_native_population.py:40, asec_coverage_authentication.py:43, asec_student_controls.py:53, asec_person_coverage_source.py:32). Raising it is harmless at :217 too, since H_NUMPER is independently pinned by `== len(household.persons)` (:216-219) with MAX_MEMBERS = 20 (survey_population_domains.py:18). CONSTRAINTS: do not lower it below ~142,125 (:186) / ~55,762 (:203); keep :203 tight enough to remain a real guard on the ASEC catalogue tuple, which has no other whole-tuple ceiling; and raise graph_child_property_income.py:80 MAX_ORIGINALS in the same change, or the lane still refuses at 600,000 with \"ORIGINAL_RECIPIENT_ROSTER\" (graph_child_property_income.py:308-309), with graph_joint_empirical.py:45 MAX_RECIPIENTS = 1_048_576 next in line.", + "evidence": "Value and refusal machinery: packages/microcosm-build/src/microcosm/build/us_runtime/current_child_property_income_source.py:34 (MAX_ROWS = 600_000); :41-43 (`def require(condition, code)` / `raise ValueError(\"CHILD_PROPERTY_\" + code)`).\nFour sites: :186 `and 0 < len(frame) <= MAX_ROWS` with code \"FULL_SOURCE_AXIS\" at :188; :203 `require(type(households) is tuple and len(households) <= MAX_ROWS, \"CATALOGUE_TYPE\")`; :217 `_integer_literal(household.h_numper, \"H_NUMPER\", MAX_ROWS)` -> :175 `require(0 <= result <= maximum, name + \"_RANGE\")` = \"H_NUMPER_RANGE\" (and \"H_NUMPER_LITERAL\" at :167-173); :393 `and len(table) <= MAX_ROWS` with code \"RECIPIENT_AXIS\" at :394.\nBinding quantity: :372-378 (`payload = origins[\"persons\"]`, `table = pd.DataFrame(payload[\"rows\"], ...)`); :612-618 (origins from `json.loads(entry[1])[\"origins\"]`, passed to project_child_property_recipients); survey_population_preparation.py:1154, :1201, :1208-1216, :1293-1316 (origins[\"persons\"][\"rows\"] = one row per stacked-frame person, both channels, pre-clone).\nNon-binding ASEC sites: current_asec_interest_source.py:389-413 (pins -> `require(len(keys) == rows ...)`, \"COMPLETE_SOURCE_JOIN\"), :430-447 (`literals = ordered.copy()` is the complete member frame); asec_coverage_authentication.py:34-37 (_MEMBER_PINS from ASEC_EDUCATION_ASSISTANCE_ARCHIVES); education_assistance_source.py:155-161 (pppub25.csv, member_size_bytes=277_882_549, rows=142_125); pilot artifact /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json sha256 34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0, keys /catalogues/asec/counts/households 55762 and /catalogues/asec/counts/persons 142125 (re-read by me).\nNo tighter upstream ceiling: survey_population_preparation.py:57-58 (MAX_SEGMENT_BYTES / MAX_ROSTER_BYTES = 64 * 64MiB), :1144-1151 (_bounded_append, \"ORIGIN_LIMIT\", comment \"a full-source roster is 1.02 GiB\"); survey_population_domains.py:19 MAX_HOUSEHOLDS = 100_000 enforced per batch at :551 \"HOUSEHOLD_BATCH\"; survey_catalogue_selection.py:19 _BATCH_HOUSEHOLDS = 10_000, checked :128-129; asec_population_catalogue.py:26-27 (byte-only record bounds); current_acs_income_anchor_source.py:113-132 + :372-376 (_scan `maximum=total_rows - count`, i.e. an upstream-file row bound, \"SOURCE_ROW_SHAPE\", not a stacked-roster cap).\nPaired downstream ceiling the census missed: graph_child_property_income.py:80 MAX_ORIGINALS = 600_000; :302-309 (`recipients = qualified.recipients`; `0 < len(recipients) <= MAX_ORIGINALS`, \"ORIGINAL_RECIPIENT_ROSTER\"); :923-925 (qualified built from child.qualify_child_property_sources); packages/microcosm-fit/src/microcosm/fit/graph_joint_empirical.py:45, :526, :610 (MAX_RECIPIENTS = 1_048_576, \"RECIPIENT_COUNT\").\nSerialization check: current_child_property_income_source.py:100-152 (_live's child_property_contract via source._runtime_marker) vs survey_population_preparation.py:607-621 (_runtime_marker returns (type, value) / (type, id(value))); :556-573 (seal excludes it); :625-650 (evidence excludes it)." + }, + { + "constant": "puf_full_source.py:94 AGGREGATE_LEXEME_MAX_CHARACTERS = 64", + "binds_at_full_source_verdict": false, + "protects_verdict": "structural-invariant", + "agrees_with_census": false, + "correction": "Three corrections. Two are citation errors, one is a classification error; the census's headline conclusions (value 64, dead name, 236 tokens, does not bind at full source) survive.\n\n(1) EVERY line citation in the row is wrong at HEAD (3fb1111e2; puf_full_source.py is unmodified on this branch since 815730051, so these were never right here). Actual sites, read myself:\n - decode of aggregate rows: `_AGGREGATE.fullmatch` at puf_full_source.py:220, inside `_require(...)` spanning :219-222, refusal literal \"FULL_AGGREGATE_LEXICAL_GRAMMAR\" at puf_full_source.py:221. The census's \":211-214\" is the record-loop header (`raw._records(...)` at :211) and `_require(row_index < rows, \"FULL_EXCESS_RECORD\")` at :213 \u2014 a different bound with a different refusal code.\n - artifact decode: `_AGGREGATE.fullmatch(v)` at puf_full_source.py:395, `_require(...)` spanning :391-398, refusal literal \"FULL_AGGREGATE_LEXICAL_GRAMMAR\" at puf_full_source.py:397. The census's \":386-391\" lands mostly on `_require(..., \"FULL_AGGREGATE_TOKEN_KEYS\")` at :386-389 \u2014 again a different refusal code.\n - `_require` is defined at puf_full_source.py:98-100 (`raise ValueError(code)` at :100), not :97-99. Exception type ValueError is correct.\n\n(2) MISCLASSIFICATION: \"encoding-width\" is wrong under the taxonomy. Nothing stores these tokens in a fixed-width field and no reader computes a byte budget from 64. The tokens are emitted as ordinary variable-length JSON strings into the artifact header (puf_full_source.py:280-282), the header is length-prefixed with `struct.pack(\" 59\n :190-191 _require len==pin.bytes \"FULL_SOURCE_SIZE\"; sha256==pin.sha256 \"FULL_SOURCE_SHA256\"\n :216-223 aggregate branch; :219-222 _require(all(_AGGREGATE.fullmatch(token) ...)) with literal \"FULL_AGGREGATE_LEXICAL_GRAMMAR\" at :221\n :254 _HEADER_LIMIT = 128 * 1024\n :272-288 json.dumps header incl. \"aggregate_tokens\" at :280-282\n :289 _require(len(header) <= _HEADER_LIMIT, \"FULL_HEADER_LIMIT\")\n :291 _MAGIC + struct.pack(\" 0, \"FULL_HEADER_LIMIT\")\n :386-389 _require(..., \"FULL_AGGREGATE_TOKEN_KEYS\") [what the census mis-cited]\n :390-398 for fields in tokens.values(): _require(... _AGGREGATE.fullmatch(v) at :395 ...) with literal \"FULL_AGGREGATE_LEXICAL_GRAMMAR\" at :397\n\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_raw_source.py:\n :141 PUF_AGGREGATE_RECIDS = (999996, 999997, 999998, 999999)\n :183-185 comment + _LEXICAL_WIDTH_MAX = 64 (\"widest lexical allocation any field may declare\")\n :189 _LEXICAL_KIND = \"lexical_fixed_ascii_nul_padded\"\n :382-383 raise _refuse(\"DEFINITION_LEXICAL_WIDTH\", ...) when not 0 < width <= 64\n :386-387 raise _refuse(\"DEFINITION_AGGREGATE_RECIDS\") when recids mismatch\n :625-629 field_size_limit pin checks (\"DEFINITION_CSV_FIELD_SIZE_LIMIT\", \"CSV_FIELD_SIZE_LIMIT\")\n :724-742 _records: \"RECORD_CAP\", \"RECORD_WIDTH\", \"RECORD_CHARACTER_CAP\"\n :870-887 _check_declared_expectations -> _refuse(\"DECLARED_EXPECTATION\", label) incl. \"aggregate_records\"\n :1272-1276 if width > _LEXICAL_WIDTH_MAX: _refuse(\"ENVELOPE_COLUMN_WIDTH\"); expected = rows * width [the real fixed-width site]\n :29-31, :38-40 module docstring: 128 MiB cap because pinned main delivery is 126,034,649 bytes; four publisher aggregate records; 88,021 unmatched returns\n\npackages/microcosm-build/src/microcosm/build/us_runtime/puf_2015_raw_source_definition.json:\n aggregates.expected_records = 4; aggregates.recids = [999996..999999]\n csv_profile: field_size_limit 131072, logical_record_character_cap 1048576, record_cap_per_file 1000000, header_field_cap 1024\n sources.main: bytes 126034649, data_records 207696, sha256 0a7fd643edb1acc55c507db795914b41d232922be78c149b58d111f4672499df\n sources.demographic: bytes 2225050, data_records 119675\n join.expected_main_unmatched_records 88021, join.expected_matched_keys 119675\n\nRepo-wide grep for AGGREGATE_LEXEME_MAX_CHARACTERS over packages/, tools/, docs/: single hit, puf_full_source.py:94 (the name is dead \u2014 census correct)." + }, + { + "constant": "puf_monetary_source.py:187 LEXICAL_WIDTH_MAX = 64", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "Every verdict-bearing field of the row survives my own read: the value (64 at :187), both enforcement sites, both refusal-code string literals, the exception type, the counted quantity (declared bytes per lexical cell, not any roster), the 13-vs-64 headroom, invariance under sampling fraction, binds-at-full=false, and the encoding-width classification. But three statements in the row are factually wrong or incomplete at HEAD, and one is a fabricated-looking citation:\n\n(1) WRONG CITATION. The row says \"every reader slices body[position*width:(position+1)*width] (_read_fixed_ascii, :1440-1451)\". `_read_fixed_ascii` is defined at puf_monetary_source.py:1338-1351, and the slice is at :1343 (`cell = body[position * width : (position + 1) * width]`). Lines 1440-1451 are the envelope header's `sources`/`demographic` pin dict inside the encoder (`\"demographic\": {...definition.demographic.sha256...}` at :1440-1447, `\"header_length_endianness\": \"big\"` at :1449) \u2014 no reader, no slicing. The mechanism described is real; the line range cited for it is not.\n\n(2) OFF-BY-ONE CITATION. The row says bytes == rows * width is required \"(:1553, :1565)\". At HEAD, `expected = rows * width` is :1552 and :1553 is the bare `else:`; the comparison is `if size != expected:` at :1564 with `raise _refuse(\"ENVELOPE_COLUMN_SIZE\", name)` at :1565.\n\n(3) INCOMPLETE SCOPE / UNDERCOUNT. The row says the constant governs \"the packaged projection ... twelve columns (puf_2015_monetary_source_projection.json)\". Both enforcement sites also gate a SECOND packaged document: puf_monetary_agi_projection.py imports `_projection_from_document` (:109) and `_envelope_columns` (:103, called at :448) from puf_monetary_source, and `packaged_agi_projection()` at :226-237 parses puf_2015_monetary_agi_source_projection.json, which declares 13 columns, every one lexical_width 13 (I parsed both JSONs). So the constant guards 25 declared widths across two documents, all 13 \u2014 value unchanged, surface larger than stated.\n\nTwo omissions worth recording, both of which strengthen rather than weaken the row's conclusion:\n- There is a lower bound on the same field: puf_monetary_source.py:515 `if digits > _MAX_AMOUNT_DIGITS or width < digits + 1: raise _refuse(\"PROJECTION_WIDTH_BELOW_GRAMMAR\", field)` with `_MAX_AMOUNT_DIGITS = 12` at :338. That is why 13 is the declared width (12 publisher digits + one sign position). There is NO tighter UPPER bound anywhere on the path, so 64 is the operative cap \u2014 but it is unreachable in production because both packaged documents are closed, fully validated, and the envelope must match the document width exactly (:1550 `width != widths[field]`).\n- 64 is not a lone number. puf_raw_source.py:185 defines a near-duplicate `_LEXICAL_WIDTH_MAX = 64` (enforced at :1272, same \"ENVELOPE_COLUMN_WIDTH\" code) and puf_raw_source.py:382 hardcodes `not 0 < width <= 64` \u2192 `raise _refuse(\"DEFINITION_LEXICAL_WIDTH\", ...)` on the raw source definition document, where RECID declares lexical_width exactly 64 (puf_2015_raw_source_definition.json, fields[0]). So the shared schema bound BINDS exactly in the sibling module even though it has 51 bytes of slack here.", + "safe_to_move": false, + "safe_to_move_reason": "Not safe, and moving it would buy nothing at full source. (i) It is not a scale ceiling: the quantity it caps is a per-cell byte width declared in a projection document, which is 13 regardless of whether the build selects 1,584 or 1,587,376 households \u2014 the PUF row count enters only as the `rows` multiplier in `rows * width` (:1552), bounded separately by BODY_MAX_BYTES (:180, checked :1388-1389 and :1567-1568). Raising 64 relieves zero full-source pressure. (ii) Raising it relaxes admissibility on a fixed-width wire-format field. The bodies are fixed-width NUL-padded ASCII (`_LEXICAL_KIND` :204): the encoder allocates `len(values) * width` (`_fixed_ascii` :1325-1335) and every decode slices `body[position*width:(position+1)*width]` (`_read_fixed_ascii` :1343, called at :1704 and at puf_monetary_agi_projection.py:466) after asserting `size == rows * width` (:1552, :1564-1565). A larger cap admits a document declaring a wider column, and the stored layout changes with it; that is exactly the \"must be argued separately\" category, not a movable resource ceiling. (iii) Lowering it is worse: anything below 13 refuses both packaged documents outright at :511-512, and the same 64 is mirrored by puf_raw_source.py:185/:382 where the raw definition's RECID declares exactly 64 \u2014 moving one of the three 64s desynchronizes the shared schema bound. This is also NOT an upstream-file-size assertion (no row-count claim about the real PUF is involved), so nothing is being weakened in that sense; it simply should be left alone.", + "evidence": "HEAD 3aa1802a88d80672944cb6906c07ce6486163dce, tree clean for this module (only docs/us-native-row-ceilings.md and packages/microcosm-build/tests/test_us_native_row_ceilings.py are modified).\nAll paths under /Users/maxghenis/PolicyEngine/_worktrees/microcosm-native-row-ceilings/packages/microcosm-build/src/microcosm/build/us_runtime/.\npuf_monetary_source.py:186 comment \"The widest lexical allocation a projected column may declare\"; :187 `LEXICAL_WIDTH_MAX = 64`; :132 exported in __all__.\nSite 1: :510 `width = _positive_int(entry[\"lexical_width\"], ...)`; :511 `if width > LEXICAL_WIDTH_MAX:`; :512 `raise _refuse(\"PROJECTION_LEXICAL_WIDTH\", field)`. Reached from `_projected_column` via `_projection_from_document` :635, used by `packaged_projection()` :679-688.\nSite 2: :1538 `widths = {column.field: column.lexical_width ...}`; :1549 `_envelope_int(entry[\"width\"], ..., minimum=1)`; :1550 `if width > LEXICAL_WIDTH_MAX or width != widths[field]:`; :1551 `raise _refuse(\"ENVELOPE_COLUMN_WIDTH\", name)`.\nException type: :341 `class PufMonetaryRefusalError(ValueError)`; :349-350 `def _refuse(reason, *detail) -> PufMonetaryRefusalError: return PufMonetaryRefusalError(reason, *detail)`.\nEncoding mechanics: :204 `_LEXICAL_KIND = \"lexical_fixed_ascii_nul_padded\"`; :1325-1335 `_fixed_ascii`; :1338-1351 `_read_fixed_ascii` (slice at :1343); :1354-1366 `payload_bound` (`sizes[f\"{column.field}_lexical\"] = rows * column.lexical_width` at :1365); :1552 `expected = rows * width`; :1564-1565 size equality \u2192 \"ENVELOPE_COLUMN_SIZE\"; :1704 reader call; :1567-1568 \"ENVELOPE_BODY_OVER_BOUND\" against BODY_MAX_BYTES (:180).\nLower bound / grammar: :338 `_MAX_AMOUNT_DIGITS = 12`; :515 `if digits > _MAX_AMOUNT_DIGITS or width < digits + 1:` \u2192 \"PROJECTION_WIDTH_BELOW_GRAMMAR\"; :232 `_GRAMMAR = \"ascii_decimal_signed_integer\"`.\nPackaged data, parsed directly: puf_2015_monetary_source_projection.json \u2192 12 columns, lexical_width 13 and field_width_digits 12 for every one (E00200, E00300, E00400, E00600, E00650, E00900, E01000, E01500, E01700, E02100, P22250, P23250). puf_2015_monetary_agi_source_projection.json \u2192 13 columns, lexical_width 13 for every one.\nSecond consumer: puf_monetary_agi_projection.py:103 imports `_envelope_columns`, :109 `_projection_from_document`, :115 `packaged_projection`; :226-237 `packaged_agi_projection()`; :448 `_envelope_columns(header, rows, resolved)`; :466 `_read_fixed_ascii(body, rows, column[\"width\"], field)`.\nSibling 64s: puf_raw_source.py:183-185 `_LEXICAL_WIDTH_MAX = 64` (\"matching the closed definition's own ``0 < lexical_width <= 64`` bound\"); :1272-1275 `if width > _LEXICAL_WIDTH_MAX: raise _refuse(\"ENVELOPE_COLUMN_WIDTH\", name)`; :381-383 `width = field[\"lexical_width\"] ... not 0 < width <= 64 \u2192 raise _refuse(\"DEFINITION_LEXICAL_WIDTH\", field[\"name\"])`. puf_2015_raw_source_definition.json fields: RECID lexical_width 64, S006 32, FLPDYR 8, others 4.\nRepo-wide grep for LEXICAL_WIDTH_MAX returns only puf_raw_source.py and puf_monetary_source.py; no test file in packages/microcosm-build/tests references LEXICAL_WIDTH_MAX, PROJECTION_LEXICAL_WIDTH, ENVELOPE_COLUMN_WIDTH or PROJECTION_WIDTH_BELOW_GRAMMAR, so no test pins this ceiling.\nNot established (and not needed for the verdict): I did not re-derive the PUF's delivered row count; the module comment at :176 states 207,696 delivered rows. This module is the IRS PUF source lane, not the ASEC/ACS selected-survey lane, so the 1,587,376-household scaling numbers supplied in the brief do not enter this bound at all." + }, + { + "constant": "graph_composed_asec_binding.py:78 HEADER_MAX_BYTES (imported from asec_current_money.py:47, `HEADER_MAX_BYTES = 64 * 1024` = 65,536)", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "The row's three headline verdicts survive scrutiny (value 65,536; binds-at-full = false; encoding-width; do not move), but four things in it are wrong or missing, and one of them is a measured count.\n\n1. THE COUNTS ARE WRONG (~600 bytes is a 25-35% underestimate). I reconstructed the header byte-for-byte from the literal constructor at graph_composed_asec_binding.py:995-1013 using the module's own `_json` encoder (asec_current_money.py:77-81: sort_keys, separators (\",\",\":\"), ensure_ascii) with a 64-hex binding_sha256, a 64-hex per-buffer sha256, and \"bytes\" = rows*8 (the width the reader asserts at :1093). Measured lengths:\n - 1/1000 arm (3,464 person / 1,584 household rows): 794 bytes\n - ASEC-1yr-scale arm (180k/90k): 803 bytes\n - pooled-3yr-scale arm (550k/270k): 806 bytes\n - at the declared ceilings MAX_PERSONS=1,000,000 / MAX_HOUSEHOLDS=400,000: 807 bytes\n - hypothetically at full stacked (3,471,383/1,587,376) or full cloned (6,942,766/3,174,752): 812 bytes\n - at 1000x the full cloned roster: 830 bytes\n So \"at 1/10: roughly 600\" and \"at full source: roughly 600\" should read ~805 and ~810. The conclusion is unchanged (~80x headroom, and the shape is row-count independent: only the decimal digits of four \"bytes\" ints and two \"rows\" ints move), but the census records counts and these are off.\n\n2. TWO LINE CITATIONS ARE WRONG. The refusal literal \"ARM_ROWS_HEADER_SIZE\" is at :1016 (inline on the `_require`, not :1018 \u2014 that line is blank) and at :1058 (not :1059 \u2014 that is the closing paren). Also `struct.pack(\" _asec_current_money_codec.py:8/:27, asec_current_money_selection.py:31/:86, puf_monetary_agi_projection.py:95, graph_composed_asec_binding.py:78/:959) plus roughly a dozen `_parse` limits on unrelated documents (asec_current_money.py:147, :417, :739, :868, :1050; selection :193, :221, :281, :399). A \"move the arm-rows header ceiling\" edit silently moves the ASEC money codec's and selection codec's frames too.\n\n(c) Identity. Editing asec_current_money.py changes that file's own sha256, which is hashed into the execution identity at asec_current_money_resources.py:49 (MODULE_NAMES at :30-34) and asserted at asec_current_money.py:417-433 (\"EXECUTION_MODULES\"). The constant is not a free-standing knob; the module is attested.\n\n(d) No coverage. Nothing outside src/ references arm_rows or any of these refusal codes, so a change would ship unverified.\n\n(e) No need. Measured: the header is 794 bytes at the 1/1000 pilot arm, 807 at the declared MAX_PERSONS/MAX_HOUSEHOLDS ceilings, and 812 even if the arm were the full cloned roster \u2014 80x under 65,536, and it cannot grow with rows because the document is 7 fixed keys, a 2-entity `rows` map (_ID_ENTITIES, :172) and a 4-entry buffer roster (_ARM_ROWS_BUFFERS, :939-944). It is not the binding constraint on this payload at any source scale.\n\nIt is NOT an upstream-file-size assertion, so the \"must not be moved\" category (c) does not apply \u2014 the reason to leave it alone is (a)+(b), plus the absence of any reason to change it.", + "evidence": "Constant and value: packages/microcosm-build/src/microcosm/build/us_runtime/asec_current_money.py:47 (`HEADER_MAX_BYTES = 64 * 1024`), :48-49 (`MAX_PERSONS = 1_000_000`, `MAX_HOUSEHOLDS = 400_000`).\nImport: graph_composed_asec_binding.py:77-84 (HEADER_MAX_BYTES on :78).\nProducer enforcement: graph_composed_asec_binding.py:1016 `_require(len(header) <= HEADER_MAX_BYTES, \"ARM_ROWS_HEADER_SIZE\")` inside `_arm_rows_parts` (:984-1017).\nReader enforcement: graph_composed_asec_binding.py:1054 `size = struct.unpack_from(\" PreparedGraphError; graph_asec_prepared.py:145 `class PreparedGraphError(ValueError)`. asec_current_money.py:68-70 `_require` -> MoneyRefusalError; asec_current_money.py:59 `class MoneyRefusalError(ValueError)`; asec_current_money.py:83-84 `_parse` limit refusal literal \"JSON_SIZE_OR_TYPE\".\nWhich roster is counted: graph_composed_asec_binding.py:299-303 (`_original_ids` reads spine_source_id_column under the arm mask), :641-644, :645-648, :649-653 (\"ARM_IDENTITY_ROWS\"), :636-639 (\"ARM_OVERFLOW\", masks sum <= receipt[\"entity_rows\"]), :585-590 (\"PREPARED_EVIDENCE_ROWS\"), :306-332 (`_resolved_positions`, strictly increasing injection into source order), :1268-1272 (producer call site).\nShared-constant blast radius: _asec_current_money_codec.py:8, :27, :116, :120; asec_current_money_selection.py:31, :86, :193, :221, :276, :281, :399, :504; puf_monetary_agi_projection.py:95, :357, :414. Independent 64-KiB copies: puf_raw_source.py:124, puf_monetary_source.py:184, asec_housing_universe.py:40, asec_housing_status.py:52.\nModule attestation: asec_current_money_resources.py:30-34 (MODULE_NAMES) and :45-50 (per-module sha256 into the execution identity); asec_current_money.py:417-433 (EXECUTION_MODULES assertions).\nTest coverage: `grep -rl \"arm_rows\" --include=\"*.py\"` over the worktree excluding /src/ returns no files; no match for ARM_ROWS_HEADER_SIZE / ARM_ROWS_SIZE / HEADER_MAX_BYTES under packages/microcosm-build/tests.\nByte counts: computed with /opt/homebrew/bin/python3.13 -I -B -S by replaying the exact literal at :995-1013 through the same json.dumps settings as asec_current_money.py:77-81 (no repo import, no file written)." + }, + { + "constant": "puf_diagnostic_consumer.py:671 CURRENT_SURVEY_MAX_BYTES = 64 * 1024**2 = 67,108,864", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": false, + "correction": "The substantive load-bearing sentence is WRONG: \"So it is a deliberate resource ceiling, movable on its own terms.\" It is not movable on its own terms. The encode-side call at graph_puf_diagnostic_consumer.py:403 passes this constant as the `limit` argument to survey_graph._bounded_json, and _bounded_json's FIRST statement hard-caps its own limit: graph_survey_population.py:264 `_require(type(limit) is int and 0 < limit <= 64 * 1024**2, \"TRANSPORT_LIMIT\")`. CURRENT_SURVEY_MAX_BYTES is currently EXACTLY at that shared transport cap, so raising the single constant by even one byte makes the encode side refuse immediately with TRANSPORT_LIMIT (SurveyPopulationGraphError, graph_survey_population.py:108/112-114) before any matrix is built. The census lists graph:403 and cites TRANSPORT_LIMIT at :273/:275 (the per-piece append checks) but misses :264, which is the one that actually blocks the move. Raising it therefore requires either splitting the matrix bound (:833) from the projection bound (:684/:651) into a second constant, or raising the shared _bounded_json cap that PREPARATION_MAX_BYTES (graph_survey_population.py:61) and ALLOCATION_MAX_BYTES (:62) also live under. Everything else in the row I confirmed, with four minor citation slips and one loose characterisation: (1) SURVEY_WAGE_ROSTER is at :755-757, not :756-758; (2) \"release_eligible\": False is at :844, not :846 (:846 is the closing brace); (3) decode_recipient_matrix reads `rows` from the header at model_input.py:114 and bounds only the header at :105 \u2014 the cited \":97-107\" is the function head, not the row read (the census's CONCLUSION here is right: no decoder reads CURRENT_SURVEY_MAX_BYTES back, MATRIX_HEADER_MAX_BYTES = 64*1024 at :18 is a separate bound, header length is a 4-byte field but `rows` is a JSON int with no width limit, and the body check at :124 is `rows*8 + rows*len(columns)*8`); (4) the pppub25.csv row pin is education_assistance_source.py:161 (`rows=142_125`), not :163; the mask is puf_detail_transfer.py:375-376, not :374-375. Loose characterisation of arm (a): the census says the ASEC roster is \"capped by the real upstream file ..., not by the survey fraction.\" It IS drawn at the survey fraction \u2014 preparation.json /selection/cells shows a separate `asec` shared_housing cell with eligible_households 55,697, inclusion_probability 55/55,697, selected_households 55, alongside the ACS cells (84,422 + 98,784 + 1,665 + 1,346,743 = 1,531,614 eligible). The correct statement is that the ASEC channel draws from its OWN eligible universe of 55,697 households, which does not grow with the 1.53M ACS households, so at full source it saturates at \u2248142,125 person rows (catalogue /catalogues/asec/counts/persons, = the pppub25.csv pin) minus the persons in the 65 excluded `asec_nonhousing` households \u2014 the exact full-source native-ASEC person count is NOT established, only that it is just under 142,125. I re-measured arm (a)'s byte estimate with the same canonical encoder settings _bounded_json uses (sort_keys, separators (\",\",\":\")): a typical 7-element row is 52 bytes and a worst-case row (int64-max ids, 255 status bytes) is 75 bytes, so 142,125 rows is 7.5\u201310.8 MB. Arm (a) never binds; that conclusion stands. Arm (b) I confirm exactly: FEATURES has 5 names (puf_detail_transfer.py:29), encode_recipient_matrix writes int64 ids + 5 float64 columns = 48 B/row (model_input.py:88-94), the mask is `units[support_clone_index_column(\"tax_unit\")].eq(1)` (puf_detail_transfer.py:375-376) = the stacked (not cloned) tax-unit count, 2,132,510 \u00d7 48 = 102,360,480 > 67,108,864, so SURVEY_MATRIX_BOUND refuses at full source. No tighter upstream limit intervenes: the only other len() bound on the matrix in the whole tree is :833 itself; ALLOCATION_ROSTER_BYTES is 64 \u00d7 64 MiB = 4 GiB (graph_survey_population.py:68); and RAW_BYTES_MAX_BYTES = 64 MiB (microcosm-graph/codecs.py:78, enforced at :576) applies only to raw-bytes-v1 SOURCE FILE reads, not to kernel artifact payloads, so it does not cap the emitted matrix.", + "safe_to_move": true, + "safe_to_move_reason": "Safe in the census's sense \u2014 it is neither (b) a fixed-width encoding nor (c) an upstream-file-size assertion \u2014 but NOT safe to move by editing this one number. Not a fixed-width encoding: decode_recipient_matrix (model_input.py:97-135) never reads this constant; it bounds only the header with MATRIX_HEADER_MAX_BYTES (:18, :105), reads `rows` as an unbounded JSON int (:114), and validates the body as rows*8 + rows*len(columns)*8 (:124) \u2014 a larger payload changes no wire format. Not an upstream-file assertion: the real upstream check is education_assistance_source.py:161 (`rows=142_125` on pppub25.csv) and the roster-equality require at puf_diagnostic_consumer.py:755-757 (len(rows) == len(native_asec)); the byte ceilings assert nothing about the CPS/ACS files. Not structural: SCOPE is \"development_conditional_support_not_tax_inputs\" (:28) and release_eligible is False (:844), so no released artifact depends on it. Required co-changes before it can actually be raised: (1) graph_survey_population.py:264 caps _bounded_json's `limit` at 64*1024**2, and graph_puf_diagnostic_consumer.py:403 passes this constant as that limit \u2014 raising the constant alone yields TRANSPORT_LIMIT, so the matrix bound at :833 must be split into its own constant (leaving the projection bound at :684/:651 at 64 MiB, which is correct since arm (a) tops out near 10 MB) or the shared transport cap must be raised, which also loosens PREPARATION_MAX_BYTES/ALLOCATION_MAX_BYTES paths. (2) puf_diagnostic_consumer is inside source_hash at graph_puf_diagnostic_consumer.py:66-69, so any edit changes the kernel implementation hash and every node key/spec digest derived from it \u2014 this is a re-pin, not a hazard, but it must be done in the same change.", + "evidence": "packages/microcosm-build/src/microcosm/build/us_runtime/puf_diagnostic_consumer.py:26 (require = detail.require), :28 (SCOPE), :671 (CURRENT_SURVEY_MAX_BYTES = 64 * 1024**2), :682-685 (require -> \"SURVEY_HOST_PROJECTION_BOUND\" at :684), :750-752 (native_asec = clone_index 0 AND channel \"asec\"), :755-757 (\"SURVEY_WAGE_ROSTER\", len(rows) == len(native_asec) > 0), :830-833 (\"SURVEY_MATRIX_BOUND\" at :833), :844 (release_eligible False). packages/microcosm-build/src/microcosm/build/us_runtime/puf_detail_transfer.py:29 (FEATURES, 5 names), :79-82 (def require -> bare ValueError), :272 (def recipient_matrix), :327-330 (HOST_ROLE_PAIR_COVERAGE, len(native)==len(detail)), :375-376 (mask = tax_unit clone_index == 1), :448-452 (encode_recipient_matrix(features, entity=\"tax_unit\", entity_ids=ids), mask). packages/microcosm-build/src/microcosm/build/us_runtime/graph_puf_diagnostic_consumer.py:52 (require = shared.require), :66-69 (source_hash includes the consumer module), :403 (_bounded_json(document, consumer.CURRENT_SURVEY_MAX_BYTES)), :648-652 (\"SURVEY_HOST_PROJECTION_BOUND\" at :651), :661 (kernel re-calls current_survey_recipient_matrix). packages/microcosm-build/src/microcosm/build/us_runtime/graph_puf_detail_transfer.py:49 (require = detail.require -> bare ValueError). packages/microcosm-build/src/microcosm/build/us_runtime/graph_survey_population.py:61-62 (PREPARATION_MAX_BYTES, ALLOCATION_MAX_BYTES = 64*1024**2), :68 (ALLOCATION_ROSTER_BYTES = 64 * ALLOCATION_MAX_BYTES), :108 (class SurveyPopulationGraphError(ValueError)), :112-114 (_require raises it), :262-281 (_bounded_json), :264 (limit <= 64*1024**2, \"TRANSPORT_LIMIT\"), :273, :275. packages/microcosm-fit/src/microcosm/fit/model_input.py:18 (MATRIX_HEADER_MAX_BYTES = 64*1024), :53-94 (encode: entity_ids int64 + values float64), :86 (header bound), :97-135 (decode), :105 (header bound), :114 (rows read from header), :124 (body length = rows*8 + rows*len(columns)*8). packages/microcosm-build/src/microcosm/build/us_runtime/education_assistance_source.py:161 (rows=142_125, pppub25.csv). packages/microcosm-graph/src/microcosm/graph/codecs.py:78 and :576 (RAW_BYTES_MAX_BYTES, raw-bytes-v1 source reads only). packages/microcosm-graph/src/microcosm/graph/store.py:337 (int64 offsets, no width cap). /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json (sha256 re-verified 34b362d8...): /native/asec = {households 55, persons 140}, /native/acs = {households 1529, persons 3324}; /catalogues/asec/counts = {households 55762, persons 142125, unrepresented_households 33170}; /selection/cells asec shared_housing eligible_households 55697, inclusion_probability 55/55697, selected 55; ACS cells eligible 84422 + 98784 + 1665 + 1346743." + }, + { + "constant": "puf55_survey_recipients.py:22 RAW_BYTES_MAX_BYTES (from microcosm.graph.codecs; defined packages/microcosm-graph/src/microcosm/graph/codecs.py:78 = 64 * 1024 * 1024 = 67,108,864)", + "binds_at_full_source_verdict": true, + "protects_verdict": "memory-or-time", + "agrees_with_census": true, + "correction": "Agrees on every load-bearing point. Three immaterial corrections. (1) Line slip: `recipients` is defined at puf55_survey_recipients.py:406-408, not :409-411 (:409 is the `_current_money_surface` call). (2) The constant's own documentation is codecs.py:78-79 (\"The most raw-bytes-v1 will read from one source file (64 MiB)\"); codecs.py:571-580 is where that read limit is enforced, not where it is documented. (3) The stated per-profile capacities (838,860 / 932,067) are the pre-check (body-only) capacities; the post-encode arm at :431 bounds the whole payload, which adds len(_MAGIC)=39 + 4 + the ~300-byte header (model_input.py:19, 88-94), so the real capacities are roughly 838,855 / 932,062. Neither the combined-capacity argument nor the binds-at-full verdict changes. One thing the row should state more explicitly, though it does flag the sharing: the ceiling is movable only via a new module-local constant \u2014 editing RAW_BYTES_MAX_BYTES itself is not safe (see safe_to_move_reason).", + "safe_to_move": true, + "safe_to_move_reason": "Safe to move ONLY by introducing a module-local constant in puf55_survey_recipients.py; NOT safe to change the shared microcosm-graph constant. (a) No fixed-width encoding or reader byte budget depends on it: encode_recipient_matrix (model_input.py:88-94) emits _MAGIC + a 4-byte big-endian HEADER length + header JSON + ids + values, and decode_recipient_matrix (model_input.py:126-135) derives the body split from header[\"rows\"] and len(columns), never from RAW_BYTES_MAX_BYTES. The only fixed-width field is the header length, bounded by MATRIX_HEADER_MAX_BYTES = 64 KiB (model_input.py:18, 86, 105), and the header is O(1) in rows because index_identity is a fixed-size digest (qrf.py:625-638). These payloads leave as graph artifacts (graph_puf55_survey_recipients.py:149-155, 262-273) and are never re-read through load_raw_bytes. (b) No upstream-file-size assertion is weakened by relaxing the matrix arms: the quantity bounded is an in-process encoded matrix over the stacked clone-one roster, not a source file. BUT the value is the shared raw-bytes-v1 per-source-file read limit, enforced on real upstream files at codecs.py:571-580, reused for source payloads at atomic_block_sources.py:365 (\"SUPPORT_BYTES\") and atomic_block_api_sources.py:264, aliased as MAX_SUPPORT_BYTES at survey_atomic_geography.py:63, documented as frozen at puf_raw_source.py:31 (\"RAW_BYTES_MAX_BYTES stays at 64 MiB\"), and pinned by test_us_atomic_block_sources.py:243 (== 64 * 1024**2). Raising it in codecs.py would silently widen every raw-bytes-v1 source read, which is exactly the class-(c)-adjacent protection that must not be moved.", + "evidence": "packages/microcosm-graph/src/microcosm/graph/codecs.py:78-79 (RAW_BYTES_MAX_BYTES = 64 * 1024 * 1024, docstring \"The most raw-bytes-v1 will read from one source file (64 MiB)\"); codecs.py:571-580 (handle.read(RAW_BYTES_MAX_BYTES + 1) and the \"larger than the ...-byte raw-bytes-v1 limit\" refusal). puf55_survey_recipients.py:22 (import); :48 (MAX_RECEIPT_BYTES = 128 * 1024); :51-53 (_require raises bare ValueError(\"PUF55_SURVEY_RECIPIENTS_\" + reason)); :31 (PROFILES = (full.PUF55_SURVEY_SS, full.PUF55_SURVEY_SS_NO_TOTAL)); :356 (def _project); :406-408 (recipients = units[support_clone_index_column(\"tax_unit\")].eq(1)); :414-420 (per-profile loop, mask at :416, ids at :417, empty-route continue at :419-420); :421-424 (pre-encode _require len(ids) * (1 + len(profile.predictors)) * 8 <= RAW_BYTES_MAX_BYTES, \"MATRIX_SIZE\"); :428-430 (encode_recipient_matrix); :431 (post-encode _require len(payload) <= RAW_BYTES_MAX_BYTES, \"MATRIX_SIZE\"); :433-437 (\"RECIPIENT_PARTITION\" requires selected to be exactly the clone-one tax_unit_ids, uniquely); :487 (frame = state.financial_population.frame); :499-501 (_project call); :546 (receipt maximum=MAX_RECEIPT_BYTES). support_provenance.py:39 (PUF_TAX_DETAIL_CLONE_INDEX = 1); :360-363 (support_clone_index_column). graph_combined_clone.py:141-145 (native arm = clone index 0, detail copy = index 1); :259-262 (NATIVE_ROW_COUNT and DETAIL_ROW_COUNT both == len(before_ids), i.e. the detail arm equals the pre-clone stacked roster). graph_atomic_survey_financial.py:857, 1392, 1439 (financial population descends from prefix.clone_population). full_puf_enrichment.py:34 (PREDICTORS = PUF_TAX_DETAIL_DEFAULT_PREDICTORS); :36-40 (PUF59_PREDICTORS = 2 named + PREDICTORS[2:]); :43 (PUF55_SURVEY_SS_PREDICTORS = (*PUF59_PREDICTORS, SURVEY_SS_TOTAL_PREDICTOR)); :73-77 (predictors property: NO_TOTAL gets PUF59_PREDICTORS). puf_support.py:207-216 (PUF_TAX_DETAIL_DEFAULT_PREDICTORS = 8 names, so PUF59 = 8, PUF55_SURVEY_SS = 9; 80 and 72 bytes/row). packages/microcosm-fit/src/microcosm/fit/model_input.py:17-19, 53-94 (encode: MAGIC + 4-byte big-endian header length + header + ids + values), :86 and :105 (MATRIX_HEADER_MAX_BYTES = 64 KiB), :126-137 (decode derives body split from rows/columns). packages/microcosm-fit/src/microcosm/fit/qrf.py:625-638 (_index_identity is a fixed-size digest, so the header is O(1) in rows). No tighter upstream total: survey_population_domains.py:19-20 and :551, :563 are per-batch (\"HOUSEHOLD_BATCH\", \"BATCH_MEMBER_BOUND\"), with survey_catalogue_selection.py:129-131 capping the batch at 10,000 households. Sibling uses of the shared constant: atomic_block_sources.py:365; atomic_block_api_sources.py:264; survey_atomic_geography.py:63; puf_raw_source.py:29-31; packages/microcosm-build/tests/test_us_atomic_block_sources.py:243. HEAD 3fb1111e282602fff3cbb334a5d49b2f9a912cc4." + } + ] +} diff --git a/experiments/native-row-ceilings/consumer-gap.json b/experiments/native-row-ceilings/consumer-gap.json new file mode 100644 index 000000000..3ccbf1928 --- /dev/null +++ b/experiments/native-row-ceilings/consumer-gap.json @@ -0,0 +1,26 @@ +{ + "scope": "Constants and one committed transport-lane receipt. Runs nothing. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "producer": { + "module": "survey_population_preparation", + "MAX_SEGMENT_BYTES": 67108864, + "MAX_ROSTER_BYTES": 4294967296, + "receipt_is_one_joined_payload": "_roster_payload returns b''.join(segments)" + }, + "consumer": { + "module": "graph_survey_population", + "PREPARATION_MAX_BYTES": 67108864, + "checked_at": "_checked_preparation :305 PREPARATION_BYTES, :309 CONTEXT_BYTES; SurveyPopulationAllocationKernel.__init__ :577/:582", + "reached_from": "graph_survey_population :528, :655, :921, :1190; graph_atomic_survey_population :172; survey_age_calibration :467" + }, + "gap": { + "producer_total_over_consumer_cap": 64.0, + "measured_tenth_roster_bytes": 109804304, + "producer_verdict_at_tenth": "accepted", + "tenth_over_consumer_cap": 1.636211633682251, + "measured_full_source_roster_bytes": 1099892722, + "full_source_over_consumer_cap": 16.38967874646187, + "households_the_consumer_cap_admits": 96839 + }, + "source_receipt": "../microcosm-native-scale/experiments/native-scale-transport/ceiling-receipt.json" +} diff --git a/experiments/native-row-ceilings/consumer_gap.py b/experiments/native-row-ceilings/consumer_gap.py new file mode 100644 index 000000000..43ad3b5c6 --- /dev/null +++ b/experiments/native-row-ceilings/consumer_gap.py @@ -0,0 +1,92 @@ +"""Two ceilings the transport lane's change does not reach, stated from constants. + +1. **The preparation receipt's consumer still caps it at 64 MiB.** The transport + lane raised the producer's total to ``MAX_ROSTER_BYTES`` = 64 x 64 MiB, and + ``survey_population_preparation`` builds the receipt as one + ``_roster_payload`` under that total. But the bytes it hands on are checked + again by ``graph_survey_population._checked_preparation`` against + ``PREPARATION_MAX_BYTES``, still 64 MiB, with refusal ``PREPARATION_BYTES``. + The transport lane's own committed ceiling receipt measured a 1/10 roster at + 109,804,304 bytes and recorded it ``accepted`` by the producer; that is 1.64x + the consumer's cap. + +2. **The ACS coverage authentication body budget is tighter than anything else + censused**, and ``selected_body_budget.py`` measures it. + +This reads constants and one committed receipt. It runs nothing, reads no +source archive, and writes only its output path. Not a build, not a +certification, not release eligible. + + python consumer_gap.py +""" + +from __future__ import annotations + +import json +import pathlib +import sys + +ROOT = pathlib.Path(__file__).resolve().parents[2] +sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] + +from microcosm.build.us_runtime import graph_survey_population as consumer +from microcosm.build.us_runtime import survey_population_preparation as producer + +for _module in (consumer, producer): + if not pathlib.Path(_module.__file__).resolve().is_relative_to(ROOT): + raise SystemExit(f"{_module.__name__} resolved outside {ROOT}") + + +def main() -> int: + receipt = json.loads(pathlib.Path(sys.argv[1]).read_bytes()) + out = pathlib.Path(sys.argv[2]) + scales = {tuple(s["fraction"]): s for s in receipt["scales"]} + tenth, full = scales[(1, 10)], scales[(1, 1)] + record = { + "scope": ( + "Constants and one committed transport-lane receipt. Runs nothing. " + "Not a build, not a certification, not release eligible." + ), + "release_eligible": False, + "producer": { + "module": "survey_population_preparation", + "MAX_SEGMENT_BYTES": producer.MAX_SEGMENT_BYTES, + "MAX_ROSTER_BYTES": producer.MAX_ROSTER_BYTES, + "receipt_is_one_joined_payload": "_roster_payload returns b''.join(segments)", + }, + "consumer": { + "module": "graph_survey_population", + "PREPARATION_MAX_BYTES": consumer.PREPARATION_MAX_BYTES, + "checked_at": ( + "_checked_preparation :305 PREPARATION_BYTES, :309 CONTEXT_BYTES; " + "SurveyPopulationAllocationKernel.__init__ :577/:582" + ), + "reached_from": ( + "graph_survey_population :528, :655, :921, :1190; " + "graph_atomic_survey_population :172; survey_age_calibration :467" + ), + }, + "gap": { + "producer_total_over_consumer_cap": producer.MAX_ROSTER_BYTES + / consumer.PREPARATION_MAX_BYTES, + "measured_tenth_roster_bytes": tenth["roster_bytes"], + "producer_verdict_at_tenth": tenth["roster"], + "tenth_over_consumer_cap": tenth["roster_bytes"] + / consumer.PREPARATION_MAX_BYTES, + "measured_full_source_roster_bytes": full["roster_bytes"], + "full_source_over_consumer_cap": full["roster_bytes"] + / consumer.PREPARATION_MAX_BYTES, + "households_the_consumer_cap_admits": int( + consumer.PREPARATION_MAX_BYTES + / (full["roster_bytes"] / full["households"]) + ), + }, + "source_receipt": sys.argv[1], + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(json.dumps(record["gap"], indent=1)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/experiments/native-row-ceilings/selected-body-budget.json b/experiments/native-row-ceilings/selected-body-budget.json new file mode 100644 index 000000000..b074dac76 --- /dev/null +++ b/experiments/native-row-ceilings/selected-body-budget.json @@ -0,0 +1,37 @@ +{ + "scope": "Measurement of acs_person_coverage_authentication's SELECTED_BODY_BUDGET over a bounded prefix of the pilot's captured public ACS PUMS archive. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "max_body_bytes": 67108864, + "charge_expression": "6 * len(raw) + 1024 per selected row", + "refusal": "SELECTED_BODY_BUDGET (ACSCoverageAuthenticationError, a ValueError)", + "roles": { + "person": { + "archive": "/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/snapshots/acs-hu-capture-e5pg8w1z/csv_pus.zip", + "records_sampled": 200000, + "record_bytes_mean": 695.57, + "record_bytes_min": 525, + "record_bytes_max": 852, + "charge_per_selected_row_mean": 5197.43, + "rows_the_64MiB_budget_admits_mean": 12911, + "rows_the_64MiB_budget_admits_worst": 10936, + "full_source_rows": 3422888, + "fraction_of_source_admitted": 0.0037719609873299973, + "full_source_budget_bytes": 17790225775, + "full_source_over_cap": 265.0950219520402 + }, + "household": { + "archive": "/Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/snapshots/acs-hu-capture-e5pg8w1z/csv_hus.zip", + "records_sampled": 200000, + "record_bytes_mean": 589.12, + "record_bytes_min": 363, + "record_bytes_max": 756, + "charge_per_selected_row_mean": 4558.75, + "rows_the_64MiB_budget_admits_mean": 14720, + "rows_the_64MiB_budget_admits_worst": 12069, + "full_source_rows": 1531614, + "fraction_of_source_admitted": 0.009610776605593837, + "full_source_budget_bytes": 6982242749, + "full_source_over_cap": 104.0435247032118 + } + } +} diff --git a/experiments/native-row-ceilings/selected_body_budget.py b/experiments/native-row-ceilings/selected_body_budget.py new file mode 100644 index 000000000..9b14141fa --- /dev/null +++ b/experiments/native-row-ceilings/selected_body_budget.py @@ -0,0 +1,104 @@ +"""Measure the ACS coverage authentication body budget, the tightest ceiling censused. + +``acs_person_coverage_authentication`` charges every *selected* row +``6 * len(raw) + 1024`` bytes against ``MAX_BODY_BYTES`` before the low-level +reader allocates its DataFrame (:378-382, refusal ``SELECTED_BODY_BUDGET``). +The only variable is ``len(raw)``, the record's own encoded bytes, so the +admitted row count is measurable from the real archive rather than guessed. + +Reads a bounded prefix of the pilot's captured public ACS PUMS archive and +reports the record lengths and the household count the 64 MiB budget admits. +Nothing outside the output path is written. Not a build, not a certification, +not release eligible. + + python selected_body_budget.py +""" + +from __future__ import annotations + +import json +import pathlib +import statistics +import sys +import zipfile + +ROOT = pathlib.Path(__file__).resolve().parents[2] +sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] + +from microcosm.build.us_runtime import acs_person_coverage_authentication as auth + +if not pathlib.Path(auth.__file__).resolve().is_relative_to(ROOT): + raise SystemExit(f"acs_person_coverage_authentication resolved outside {ROOT}") + +# Measured full-source ACS counts; see roster-census.json. +ACS_PERSONS = 3_422_888 +ACS_HOUSEHOLDS = 1_531_614 +SAMPLE_RECORDS = 200_000 + + +def _sample(archive: pathlib.Path, prefix: str) -> list[int]: + lengths: list[int] = [] + with zipfile.ZipFile(archive) as zf: + for name in sorted(n for n in zf.namelist() if prefix in n.lower()): + with zf.open(name) as member: + member.readline() # header + for raw in member: + lengths.append(len(raw)) + if len(lengths) >= SAMPLE_RECORDS: + return lengths + return lengths + + +def main() -> int: + snapshot = pathlib.Path(sys.argv[1]) + out = pathlib.Path(sys.argv[2]) + roles = { + "person": (snapshot / "csv_pus.zip", "psam_pus", ACS_PERSONS), + "household": (snapshot / "csv_hus.zip", "psam_hus", ACS_HOUSEHOLDS), + } + measured = {} + for role, (archive, prefix, full_source) in roles.items(): + lengths = _sample(archive, prefix) + if not lengths: + raise SystemExit(f"no records sampled for {role} from {archive}") + mean = statistics.fmean(lengths) + charge_mean = 6 * mean + 1024 + charge_max = 6 * max(lengths) + 1024 + measured[role] = { + "archive": str(archive), + "records_sampled": len(lengths), + "record_bytes_mean": round(mean, 2), + "record_bytes_min": min(lengths), + "record_bytes_max": max(lengths), + "charge_per_selected_row_mean": round(charge_mean, 2), + "rows_the_64MiB_budget_admits_mean": int( + auth.MAX_BODY_BYTES // charge_mean + ), + "rows_the_64MiB_budget_admits_worst": int( + auth.MAX_BODY_BYTES // charge_max + ), + "full_source_rows": full_source, + "fraction_of_source_admitted": (auth.MAX_BODY_BYTES // charge_mean) + / full_source, + "full_source_budget_bytes": int(charge_mean * full_source), + "full_source_over_cap": charge_mean * full_source / auth.MAX_BODY_BYTES, + } + record = { + "scope": ( + "Measurement of acs_person_coverage_authentication's SELECTED_BODY_BUDGET " + "over a bounded prefix of the pilot's captured public ACS PUMS archive. " + "Not a build, not a certification, not release eligible." + ), + "release_eligible": False, + "max_body_bytes": auth.MAX_BODY_BYTES, + "charge_expression": "6 * len(raw) + 1024 per selected row", + "refusal": "SELECTED_BODY_BUDGET (ACSCoverageAuthenticationError, a ValueError)", + "roles": measured, + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(json.dumps(measured, indent=1)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/packages/microcosm-build/tests/test_us_native_row_ceilings.py b/packages/microcosm-build/tests/test_us_native_row_ceilings.py index b56afdc04..37185eae9 100644 --- a/packages/microcosm-build/tests/test_us_native_row_ceilings.py +++ b/packages/microcosm-build/tests/test_us_native_row_ceilings.py @@ -25,6 +25,7 @@ from microcosm.build.us_runtime import ( acs_native_coverage_binding, + acs_person_coverage_authentication, acs_person_coverage_columns, acs_pums, asec_current_money, @@ -33,6 +34,7 @@ graph_survey_population, survey_observed_age, survey_origin_budget, + survey_population_preparation, ) # Measured full-source counts. See the module docstring for the derivation. @@ -162,3 +164,60 @@ def test_the_measured_counts_reconcile(): assert ACS_HOUSEHOLDS + ASEC_HOUSEHOLDS == STACKED_HOUSEHOLDS assert ACS_PERSONS + ASEC_PERSONS == STACKED_PERSONS assert STACKED_PERSONS * 2 == COMBINED_CLONE_PERSONS + + +def test_the_preparation_receipt_ceiling_is_still_enforced_by_its_consumer(): + """Lifted in the producer, left at 64 MiB in the consumer. Pinned so it shows. + + The transport lane raised `survey_population_preparation.MAX_ROSTER_BYTES` to + 64 segments and reported the preparation-receipt ceiling moved from 96,860 + households to 6,206,000. `_roster_payload` still returns one joined payload, + and `graph_survey_population._checked_preparation` checks those same bytes + against `PREPARATION_MAX_BYTES`, still 64 MiB, refusing `PREPARATION_BYTES`. + + The transport lane's own committed ceiling receipt measured a 1/10 roster at + 109,804,304 bytes and recorded it accepted by the producer -- 1.64x this cap + -- and a full-source roster at 1,099,892,722 bytes, which this cap admits + 96,839 households of. That is the ceiling the transport lane lifted, still + standing one module downstream. + + Not this lane's to move: it is a byte transport, and the whole receipt is one + `bytes` because `KernelResult.artifacts` is a mapping of `bytes`. Pinned here + so the next reader meets it in a test rather than in a build. + + See experiments/native-row-ceilings/consumer-gap.json. + """ + assert survey_population_preparation.MAX_ROSTER_BYTES == 64 * 64 * 1024**2 + assert graph_survey_population.PREPARATION_MAX_BYTES == 64 * 1024**2 + assert ( + survey_population_preparation.MAX_ROSTER_BYTES + == 64 * graph_survey_population.PREPARATION_MAX_BYTES + ) + measured_tenth_roster_bytes = 109_804_304 + assert measured_tenth_roster_bytes > graph_survey_population.PREPARATION_MAX_BYTES + + +def test_the_acs_body_budget_is_the_tightest_ceiling_on_the_path(): + """0.38% of source, and neither this lane's argument nor the row family. + + `acs_person_coverage_authentication` charges every selected row + `6 * len(raw) + 1024` against MAX_BODY_BYTES before the reader allocates, + refusing SELECTED_BODY_BUDGET. Measured over 200,000 real records of the + pilot's captured public ACS PUMS archive, a person record averages 695.57 + bytes, so the charge is 5,197 bytes and 64 MiB admits 12,911 selected + persons -- 0.38% of the 3,422,888 a full-source build selects, and 265x + under at full source. + + It is a byte transport, so it takes the segmented-transport argument. It is + also why lifting `acs_person_coverage_columns.MAX_SELECTED_ROWS` in the same + lane is necessary and not sufficient: this refuses 265x earlier. + + See experiments/native-row-ceilings/selected-body-budget.json. + """ + assert acs_person_coverage_authentication.MAX_BODY_BYTES == 64 * 1024**2 + measured_charge_per_row = 5197 + admitted = ( + acs_person_coverage_authentication.MAX_BODY_BYTES // measured_charge_per_row + ) + assert admitted < ACS_PERSONS // 100 + assert acs_person_coverage_columns.MAX_SELECTED_ROWS > ACS_PERSONS diff --git a/pyproject.toml b/pyproject.toml index 2537e24bc..13510f150 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -42,6 +42,8 @@ ignore = ["E501"] "experiments/native-row-ceilings/repin.py" = ["E402"] "experiments/native-row-ceilings/regenerate_inventory_contract.py" = ["E402"] "experiments/native-row-ceilings/origin_budget_size.py" = ["E402"] +"experiments/native-row-ceilings/selected_body_budget.py" = ["E402"] +"experiments/native-row-ceilings/consumer_gap.py" = ["E402"] [tool.ruff.lint.isort] # microcosm is a PEP 420 namespace package (no top-level __init__.py), so From a40666b69da4c595f73342217e00c0d3f7bce287 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:13:26 -0400 Subject: [PATCH 16/22] Commit the census and fold the two consumer ceilings into the argument census.json: 146 bounds across five families, each with its enforcement site, refusal code and exception type, what it protects, its counts at 1/10 and full source, and whether it binds -- plus the 41 adversarial verdicts, none of which overturned a headline binding call. Fourteen bind at full source, six at 1/10. Co-Authored-By: Claude Opus 5 --- docs/us-native-row-ceilings.md | 65 ++++++++++++++++++--- experiments/native-row-ceilings/census.json | 15 ++++- 2 files changed, 70 insertions(+), 10 deletions(-) diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md index eec1cf329..1af350114 100644 --- a/docs/us-native-row-ceilings.md +++ b/docs/us-native-row-ceilings.md @@ -237,9 +237,50 @@ each was read rather than assumed: - **No per-row payload sits behind the binding site.** `:1016` takes two numpy arrays. -### 5b. The loudest one - -It is in a module this lane did touch, so it gets said plainly: +### 5b. Three the owner should see + +The census found **six** bounds that a **1/10** build meets, all of them byte +transports. That contradicts the transport lane's §5 conclusion that "a 1/10 +build meets no ceiling this lane did not lift", and two of them matter enough to +be measured rather than listed. + +**The ACS coverage authentication body budget is the tightest ceiling on the +whole path, and it is not close.** +`acs_person_coverage_authentication` charges every *selected* row +`6 * len(raw) + 1024` bytes against `MAX_BODY_BYTES` (64 MiB) before the reader +allocates its DataFrame, refusing `SELECTED_BODY_BUDGET`. The only variable is +the record's own length, so the admitted count is measurable: over 200,000 real +records of the pilot's captured public ACS PUMS archive, a person record +averages 695.57 bytes, so the charge is 5,197 bytes and the budget admits +**12,911 selected persons — 0.38% of source**, 265× under at full source. The +household role admits 14,720, 0.96%. + +That is below 1/100. It is a byte transport, so it is the transport argument's, +and it is why lifting `acs_person_coverage_columns.MAX_SELECTED_ROWS` in §3 is +**necessary and not sufficient**: this refuses 265× earlier on the same path. + +**The preparation-receipt ceiling the transport lane lifted is still enforced one +module downstream.** That lane raised `survey_population_preparation`'s total to +`MAX_ROSTER_BYTES` = 64 × 64 MiB and reported the receipt ceiling moved from +96,860 households to 6,206,000. `_roster_payload` still returns one joined +payload — it must, because `KernelResult.artifacts` is a mapping of `bytes` — +and `graph_survey_population._checked_preparation` checks those same bytes +against `PREPARATION_MAX_BYTES`, **still 64 MiB**, refusing `PREPARATION_BYTES` +(`:305`, `:309`, and again in the allocation kernel at `:577`/`:582`), reached +from six call sites including `graph_atomic_survey_population:172`. + +From the transport lane's own committed `ceiling-receipt.json`: a 1/10 roster +measured 109,804,304 bytes and was recorded `accepted` by the producer — 1.64× +this cap — and a full-source roster measured 1,099,892,722 bytes, of which this +cap admits **96,839 households**. That is, to within rounding, exactly the 96,860 +the transport lane reported as the ceiling it had lifted. + +Both are pinned in `test_us_native_row_ceilings.py`, so the next reader meets +them in a test rather than in a build. + +### 5c. The loudest one in a module this lane touched + +It gets said plainly: > **`survey_origin_budget.MAX_PAYLOAD_BYTES` (64 MiB) admits 87,838 households — > 5.53% of source.** That is below one tenth, and below the 96,860-household @@ -274,13 +315,21 @@ this module. ## 6. The census -`experiments/native-row-ceilings/` carries the receipts. Every `MAX_*` row, +`experiments/native-row-ceilings/census.json` is the census. Every `MAX_*` row, household, person, group or byte bound in `packages/microcosm-build/src/microcosm/build/us_runtime/` reachable from the -19-node financial graph and the 45-node pilot graph was read at this head — 146 -bounds across five module families — with its enforcement site, refusal code, -what it protects, and its counts at 1/10 and at full source. The lane report -holds the table. +19-node financial graph and the 45-node pilot graph was read at this head — **146 +bounds across five module families** — each with its enforcement site, its +refusal code and exception type, what it protects, the count it meets at 1/10 +and at full source, and whether it binds. Every verdict that claimed a +full-source build meets the bound, or that the bound guards an encoding width or +an upstream file's real size, then went through an adversarial pass that read the +code again and tried to refute it — **41 verdicts**, and no headline verdict was +overturned. + +The result: **14 bounds bind at full source, 6 of them at 1/10.** Seven of the +fourteen are the row counts §3 moved. The other seven are byte transports and +byte-derived row pre-checks; §5a and §5b say which and why. ## 7. Pins diff --git a/experiments/native-row-ceilings/census.json b/experiments/native-row-ceilings/census.json index e4dd1af2f..db69eed37 100644 --- a/experiments/native-row-ceilings/census.json +++ b/experiments/native-row-ceilings/census.json @@ -1,8 +1,9 @@ { - "scope": "Census of every MAX_* row/household/person/group/byte bound in packages/microcosm-build/src/microcosm/build/us_runtime/ reachable from the 19-node financial graph and the 45-node pilot graph, read at this head, with an adversarial verification pass over every verdict that claimed a full-source build meets the bound or that the bound guards an encoding width or an upstream file's size. Not a build, not a certification, not release eligible.", + "scope": "Census of every MAX_* row/household/person/group/byte bound in packages/microcosm-build/src/microcosm/build/us_runtime/ reachable from the 19-node financial graph and the 45-node pilot graph, read at this head, with an adversarial pass over every verdict claiming a full-source build meets the bound or that it guards an encoding width or an upstream file's size. Not a build, not a certification, not release eligible.", "release_eligible": false, "bounds_censused": 146, - "verdicts_verified": 41, + "verdicts_verified": 42, + "verdicts_overturned_on_binding": 0, "binds_at_full_source": [ "acs_housing_universe_source.py:43 ACS_HU_RECEIPT_MAX_BYTES", "acs_native_coverage_binding.py:31 MAX_EVIDENCE_BYTES", @@ -2876,6 +2877,16 @@ "safe_to_move": true, "safe_to_move_reason": "Safe to move ONLY by introducing a module-local constant in puf55_survey_recipients.py; NOT safe to change the shared microcosm-graph constant. (a) No fixed-width encoding or reader byte budget depends on it: encode_recipient_matrix (model_input.py:88-94) emits _MAGIC + a 4-byte big-endian HEADER length + header JSON + ids + values, and decode_recipient_matrix (model_input.py:126-135) derives the body split from header[\"rows\"] and len(columns), never from RAW_BYTES_MAX_BYTES. The only fixed-width field is the header length, bounded by MATRIX_HEADER_MAX_BYTES = 64 KiB (model_input.py:18, 86, 105), and the header is O(1) in rows because index_identity is a fixed-size digest (qrf.py:625-638). These payloads leave as graph artifacts (graph_puf55_survey_recipients.py:149-155, 262-273) and are never re-read through load_raw_bytes. (b) No upstream-file-size assertion is weakened by relaxing the matrix arms: the quantity bounded is an in-process encoded matrix over the stacked clone-one roster, not a source file. BUT the value is the shared raw-bytes-v1 per-source-file read limit, enforced on real upstream files at codecs.py:571-580, reused for source payloads at atomic_block_sources.py:365 (\"SUPPORT_BYTES\") and atomic_block_api_sources.py:264, aliased as MAX_SUPPORT_BYTES at survey_atomic_geography.py:63, documented as frozen at puf_raw_source.py:31 (\"RAW_BYTES_MAX_BYTES stays at 64 MiB\"), and pinned by test_us_atomic_block_sources.py:243 (== 64 * 1024**2). Raising it in codecs.py would silently widen every raw-bytes-v1 source read, which is exactly the class-(c)-adjacent protection that must not be moved.", "evidence": "packages/microcosm-graph/src/microcosm/graph/codecs.py:78-79 (RAW_BYTES_MAX_BYTES = 64 * 1024 * 1024, docstring \"The most raw-bytes-v1 will read from one source file (64 MiB)\"); codecs.py:571-580 (handle.read(RAW_BYTES_MAX_BYTES + 1) and the \"larger than the ...-byte raw-bytes-v1 limit\" refusal). puf55_survey_recipients.py:22 (import); :48 (MAX_RECEIPT_BYTES = 128 * 1024); :51-53 (_require raises bare ValueError(\"PUF55_SURVEY_RECIPIENTS_\" + reason)); :31 (PROFILES = (full.PUF55_SURVEY_SS, full.PUF55_SURVEY_SS_NO_TOTAL)); :356 (def _project); :406-408 (recipients = units[support_clone_index_column(\"tax_unit\")].eq(1)); :414-420 (per-profile loop, mask at :416, ids at :417, empty-route continue at :419-420); :421-424 (pre-encode _require len(ids) * (1 + len(profile.predictors)) * 8 <= RAW_BYTES_MAX_BYTES, \"MATRIX_SIZE\"); :428-430 (encode_recipient_matrix); :431 (post-encode _require len(payload) <= RAW_BYTES_MAX_BYTES, \"MATRIX_SIZE\"); :433-437 (\"RECIPIENT_PARTITION\" requires selected to be exactly the clone-one tax_unit_ids, uniquely); :487 (frame = state.financial_population.frame); :499-501 (_project call); :546 (receipt maximum=MAX_RECEIPT_BYTES). support_provenance.py:39 (PUF_TAX_DETAIL_CLONE_INDEX = 1); :360-363 (support_clone_index_column). graph_combined_clone.py:141-145 (native arm = clone index 0, detail copy = index 1); :259-262 (NATIVE_ROW_COUNT and DETAIL_ROW_COUNT both == len(before_ids), i.e. the detail arm equals the pre-clone stacked roster). graph_atomic_survey_financial.py:857, 1392, 1439 (financial population descends from prefix.clone_population). full_puf_enrichment.py:34 (PREDICTORS = PUF_TAX_DETAIL_DEFAULT_PREDICTORS); :36-40 (PUF59_PREDICTORS = 2 named + PREDICTORS[2:]); :43 (PUF55_SURVEY_SS_PREDICTORS = (*PUF59_PREDICTORS, SURVEY_SS_TOTAL_PREDICTOR)); :73-77 (predictors property: NO_TOTAL gets PUF59_PREDICTORS). puf_support.py:207-216 (PUF_TAX_DETAIL_DEFAULT_PREDICTORS = 8 names, so PUF59 = 8, PUF55_SURVEY_SS = 9; 80 and 72 bytes/row). packages/microcosm-fit/src/microcosm/fit/model_input.py:17-19, 53-94 (encode: MAGIC + 4-byte big-endian header length + header + ids + values), :86 and :105 (MATRIX_HEADER_MAX_BYTES = 64 KiB), :126-137 (decode derives body split from rows/columns). packages/microcosm-fit/src/microcosm/fit/qrf.py:625-638 (_index_identity is a fixed-size digest, so the header is O(1) in rows). No tighter upstream total: survey_population_domains.py:19-20 and :551, :563 are per-batch (\"HOUSEHOLD_BATCH\", \"BATCH_MEMBER_BOUND\"), with survey_catalogue_selection.py:129-131 capping the batch at 10,000 households. Sibling uses of the shared constant: atomic_block_sources.py:365; atomic_block_api_sources.py:264; survey_atomic_geography.py:63; puf_raw_source.py:29-31; packages/microcosm-build/tests/test_us_atomic_block_sources.py:243. HEAD 3fb1111e282602fff3cbb334a5d49b2f9a912cc4." + }, + { + "constant": "PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID = 10**15 - 1 (defined operator_column_contracts.py:64; re-exported puf_support.py:53-55)", + "binds_at_full_source_verdict": false, + "protects_verdict": "encoding-width", + "agrees_with_census": false, + "correction": "I confirm the value, the enforcement sites, the refusal shape (no short code; bare ValueError), the encoding-width classification, and binds-at-full=false. I reject the census's \"counts\" row and its stated derivation.\n\n(1) WRONG QUANTITY. The census says the bound counts \"the cloned person roster size (~6,942,766)\". It does not. Both real enforcement sites run on PRE-CLONE frames:\n - spine_assembly.py:630 sits inside `_validate_source_frame` (def at :596), called at :225 on each INPUT channel spine, i.e. the per-channel ids before any offsetting or cloning.\n - spine_assembly.py:763 sits inside `_id_offsets` (def at :732), called at :235, checking the post-collision-offset assembled id.\nA cloned frame can never re-enter assembly: spine_assembly.py:642-651 refuses any input already carrying the `*_support_channel` / `*_source_id` / `*_spine_source_id` / `*_support_clone_index` columns (\"already carries support provenance; assembly must be the provenance owner\"), which is exactly what the clone stage writes (puf_support.py:2061-2063, :2097).\nThe clone's own ids are governed by a DIFFERENT number: the guard at puf_support.py:2361 tests `shift + int(values.max()) > _INT64_MAX` (`_INT64_MAX = 2**63 - 1`, puf_support.py:2331). PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID appears only in the message string at :2367 \u2014 it is not the predicate there.\n\n(2) COUNTS ARE 2x TOO LARGE. Since ids are dense over the stacked (pre-clone) roster, the governed maximum at full source is the STACKED person count ~3,471,383, and at 1/10 ~347,138 \u2014 not 6,942,766 / 694,277. (Separately, the post-clone maximum the census seems to have in mind is not 6.94M either: the clone shifts by `clone_index * 10**digits(max_id)` (puf_support.py:2360, :2065-2069), so the cloned maximum is ~10**7 + 3.47M \u2248 1.35e7, and that value is checked against _INT64_MAX, not this constant.) The stated multipliers (10**6 at 1/10, 10**7 at full) happen to survive the error only because 347,138 and 694,277 have the same digit count, as do 3,471,383 and 6,942,766.\n\n(3) THE DERIVATION IN \"protects\" IS OFF BY ONE DECADE. `_id_multiplier_for_values` (puf_support.py:2334-2346) returns `10 ** max(1, len(str(max_id)))`. len(str(10**15 - 1)) == 15, so a frame at the bound yields multiplier 10**15, not 10**16, and (2**63-1 - (10**15-1)) // 10**15 = 9222 safe clone indices, not 921. The census reproduced this error faithfully from the code comment at puf_support.py:2324-2329, which the test docstring at test_us_spine_assembly.py:388-395 repeats. The error is conservative (real headroom is 10x larger than documented), so it is a documentation defect, not a safety defect \u2014 but it is asserted as fact in the census row.\n\n(4) LINE-RANGE NITS. The spine_assembly raise closes at :641 (census said 630-640). The puf_support raise is :2362-2368, not 2360-2368 (:2360-2361 are the shift computation and the predicate). The comment the census cites as \"puf_support.py:2318-2324\" is actually at :2324-2329; :2317-2321 is the body of `_id_multiplier_for_frame`.\n\n(5) TIGHTER UPSTREAM LIMITS EXIST (census did not look). Roster size \u2014 and therefore the dense id maximum \u2014 is capped ~8 orders of magnitude earlier: acs_pums.py:270 refuses a selected ACS roster above MAX_EXACT_PERSON_ROWS = 14,000,000 (acs_pums.py:56), and acs_person_coverage_columns.py:39 asserts the ACS source file itself at MAX_ROWS = 6,000,000 (upstream-file-size class \u2014 must not be moved). On the ASEC side asec_current_money.py:48-49 sets MAX_PERSONS = 1,000,000 / MAX_HOUSEHOLDS = 400,000, enforced at :521 with refusal code \"SOURCE_SIZE\". The clone-safe bound is therefore structurally unreachable, independent of the scale argument.\n\n(6) One census premise I checked and CONFIRMED, because it was the likeliest failure mode: ids really are dense over the SELECTED roster, not the full source. ACS selection happens at load time (acs_pums.py:267 filters by `serialnos`) before `serial_to_id = np.arange(1, len(household) + 1)` at acs_pums.py:339-341; ASEC household ids are densified at asec_pool.py:274-277 / :315; person ids are `np.arange(len(result))` at microcosm-frame units.py:219. The recovered pilot artifact agrees: /native/acs/households = 1529 and /native/asec/households = 55 sum to /origins/households = 1584, so the native frames are built already-selected. (Note for the record: the legacy pilot path via `sample_frame_households` -> `Frame.select` (microcosm-frame bundle.py:1199-1270) does NOT renumber, so on that path ids would stay at full-source magnitude. It changes nothing here \u2014 that is still ~3.4M.)", + "safe_to_move": false, + "safe_to_move_reason": "Encoding-width, and there is no reason to move it. The constant is the int64 arithmetic-width budget for the clone stage's decimal remap `id + clone_index * 10**digits(max_id)` (puf_support.py:2360, :2065-2069). Raising it spends the headroom the shift relies on: test_us_spine_assembly.py:412-417 records the exact regression it was added for \u2014 \"int64-valid IDs above the shared bound assembled fine, then overflowed the clone stage's decimal remap (OverflowError: 'Python int too large to convert to C long')\". Because `_id_multiplier_for_values` derives the multiplier from digit count, raising the cap by one decade multiplies the shift by ten and divides the safe clone count by ten. Moving it DOWN is harmless but pointless. Either way the bound has ~8 orders of headroom over the full-source assembled maximum (~3.47M) and is additionally unreachable behind acs_pums.py:270 (14,000,000 person rows) and acs_person_coverage_columns.py:39 (6,000,000 source rows), so it cannot block any full-source build and does not belong on a \"ceilings to raise\" list. It is not a memory/time ceiling and not an upstream-file-size assertion.", + "evidence": "packages/microcosm-build/src/microcosm/build/us_runtime/operator_column_contracts.py:63-64 (definition, with the comment \"Shared structural limit; importing it must not load the PUF donor operations.\"); puf_support.py:53-55 (re-export, aliased name on :54); packages/microcosm-build/tests/test_us_spine_assembly.py:395 (value pinned: `assert PUF_SUPPORT_MAX_CLONE_SAFE_SOURCE_ID == 10**15 - 1`); spine_assembly.py:596 (def _validate_source_frame), :225 (call), :630-641 (oversized-id guard, bare `raise ValueError` at :636, message \"the clone-safe bound {...}\" at :638); spine_assembly.py:732 (def _id_offsets), :235 (call), :752-760 (dtype-overflow guard using np.iinfo), :763-767 (post-offset re-check, bare `raise ValueError` at :764); spine_assembly.py:642-651 (assembly refuses frames already carrying support provenance); puf_support.py:2316-2321 (_id_multiplier_for_frame), :2324-2329 (the derivation comment), :2331 (_INT64_MAX), :2334-2346 (_id_multiplier_for_values, `10 ** max(1, len(str(max_id)))`), :2348-2369 (_remap_ids; predicate on _INT64_MAX at :2361, raise at :2362-2368, constant only in the message at :2367); puf_support.py:2061-2063 and :2097 (clone writes the provenance columns), :2065-2069 and :2099-2103 (the shift); puf_support.py:158-161 (_DEFAULT_SUPPORT_CHANNELS = base ASEC + PUF tax detail); stacked_spine.py:1010-1017 and multispine_pool.py:4489-4493 (the two assemble_spines callers, both ASEC+ACS pre-clone); acs_pums.py:267-273 (serialno selection at load + MAX_EXACT_PERSON_ROWS refusal), :339-341 (dense household ids over the selected set), :49-56 (ceiling comments); acs_person_coverage_columns.py:33-40 (MAX_ROWS = 6,000,000 as an upstream source-file assertion), :198; asec_current_money.py:48-49, :521 (refusal code \"SOURCE_SIZE\"); asec_pool.py:274-277, :315 (dense ASEC household ids), :411-414 (_source_person_id is a string provenance column, not the entity id); packages/microcosm-frame/src/microcosm/frame/units.py:219 (person_id = np.arange(len(result))), :206-213 (dense group ids); packages/microcosm-frame/src/microcosm/frame/bundle.py:1199-1270 (Frame.select subsets without renumbering); packages/microcosm-build/src/microcosm/build/frame_sampling.py:108-175 (sample_frame_households); /Users/maxghenis/PolicyEngine/_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json (/native/acs/households=1529, /native/asec/households=55, /origins/households=1584, /catalogues/acs/counts/people=3422888). Grep confirms no `require`/`_require` refusal-code helper in spine_assembly.py or operator_column_contracts.py; puf_support.py's `_require_*` helpers (:1222, :4118, :3019, :3153) are unrelated to this bound." } ] } From d277ee8d7f02320048ccffabf7a9d0c6a3d788db Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:14:30 -0400 Subject: [PATCH 17/22] Bring the journal and the changelog fragment up to the seven Co-Authored-By: Claude Opus 5 --- PROGRESS-native-row-ceilings.md | 65 ++++++++++++---------- changelog.d/native-row-ceilings.changed.md | 2 +- 2 files changed, 37 insertions(+), 30 deletions(-) diff --git a/PROGRESS-native-row-ceilings.md b/PROGRESS-native-row-ceilings.md index 74807bd59..cf5226c74 100644 --- a/PROGRESS-native-row-ceilings.md +++ b/PROGRESS-native-row-ceilings.md @@ -42,40 +42,47 @@ the reconciliation. ## Done -- Census of every `MAX_*` bound in `us_runtime/` reachable from the two graph - entry points: 146 bounds across five module families, each with its - enforcement site, refusal code, what it protects, and its counts at 1/10 and - full source. **14 bind at full source; 5 of those bind at 1/10 too.** +- **Census**: 146 bounds across five module families, reachable from the 19-node + financial graph and the 45-node pilot graph, each with its enforcement site, + refusal code and exception type, what it protects, its counts at 1/10 and at + full source, and whether it binds. 41 of those verdicts then went through an + adversarial pass; no headline binding call was overturned. + **14 bind at full source, 6 of them at 1/10.** - `survey_origin_budget.MAX_GROUPS` **established**: `allocation_instructions` requires one instruction per selected household, so a full-source budget has - 1,587,376 groups and the 1,000,000 bound binds. -- Five constants lifted under the rule (commit `14defbfc0`). -- Pins re-derived through their generators (`2ebd246f1`): one moved, - `acs_native_coverage_binding._ACCEPTED["acs_pums.py"]`. -- **An inherited break re-pinned** (`55ca820c7`): base-branch commit `b6081efcb` - added `path.read_bytes()` to `_spill_roster` without regenerating + 1,587,376 groups against a 1,000,000 bound. +- **Seven ceilings lifted** under the rule, each with a boundary test. +- Pins re-derived through their generators: one moved + (`acs_native_coverage_binding._ACCEPTED["acs_pums.py"]`). 124 inventory + contracts checked, 10 stage manifests built. +- **An inherited break re-pinned**: base-branch `b6081efcb` added + `path.read_bytes()` to `_spill_roster` without regenerating `survey_population_preparation.py`'s `resource_accesses_sha256`, so - `implementation_manifest()` raised for every stage containing it — including - the one the nineteen-node path runs. Not this branch's file; re-pinned here so - the base is functional and this lane's own manifests can be built. -- Tests: the rule as an executable table, plus a boundary test per moved bound. + `implementation_manifest()` raised for every stage containing it. Proven + pre-existing by recomputing the contract from `origin/native-scale-transport`, + `a64f7b733` and `b6081efcb`'s own blobs, and by running the affected tests + against a base worktree: 7 fail there with that exact error and all pass here. +- `docs/us-native-row-ceilings.md`, and the rule as an executable table. -## The loudest finding +## The three findings the report leads with -`survey_origin_budget.MAX_PAYLOAD_BYTES` (64 MiB) admits **87,838 households, -5.53% of source** — below 1/10, and below the 96,860-household ceiling the -transport lane lifted. Measured through the module's own encoder at full-source -id widths: 764 B per group, 1.13 GiB at full source, 18.07× the cap. It is a -byte transport, so this lane lifts `MAX_GROUPS` and leaves it: at full source the -refusal moves from `GROUP_COUNT_BOUND` to `TRANSPORT_LIMIT`. Necessary, not -sufficient, and the report says so. +1. **`acs_person_coverage_authentication.MAX_BODY_BYTES` admits 12,911 selected + ACS persons — 0.38% of source.** Measured over 200,000 real records of the + pilot's captured public archive. The tightest ceiling on the path, 265× under + at full source, and below 1/100. +2. **The preparation-receipt ceiling the transport lane lifted is still enforced + one module downstream** at `PREPARATION_MAX_BYTES` = 64 MiB, which admits + 96,839 households — to within rounding the exact 96,860 that lane reported as + lifted. +3. **`survey_origin_budget.MAX_PAYLOAD_BYTES` admits 87,838 households, 5.53%**, + and cannot be raised at all in that module: `graph._bounded_json` refuses any + limit above 64 MiB before encoding a byte. + +All three are byte transports and take the transport lane's argument, not this +one. All three are pinned in tests. ## Next -1. Finish the adversarial verification of the 14 binding verdicts. -2. Decide, on that evidence, whether the two pure row-count bounds the transport - lane's census missed (`asec_demographic_source._MAX_PERSONS`, - `current_child_property_income_source.MAX_ROWS`, both 600,000) move here. -3. `docs/us-native-row-ceilings.md`. -4. Tests as CI runs them; `ci_test_groups --verify`; `spec_engine_coverage --check`. -5. Draft PR against `native-scale-transport`. +- Final clean test run, then the draft PR against `native-scale-transport`. +- Open for Max: whether the inherited re-pin stays here or moves to #945; and + whether the byte transports above are one follow-up lane or several. diff --git a/changelog.d/native-row-ceilings.changed.md b/changelog.d/native-row-ceilings.changed.md index 60a5e5c71..16d60b926 100644 --- a/changelog.d/native-row-ceilings.changed.md +++ b/changelog.d/native-row-ceilings.changed.md @@ -1 +1 @@ -Raise the six US native-build row-count ceilings a full-source build meets to four times their measured full-source counts, keeping every refusal code and expression, and re-pin the one implementation digest that moves. +Raise the seven US native-build row-count ceilings a full-source build meets to four times their measured full-source counts, keeping every refusal code and expression, and re-pin the one implementation digest that moves. From 099c83c8a42ed15631b569bfd1af6bfa5eb27052 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:16:53 -0400 Subject: [PATCH 18/22] Write the lane report's census, argument, moved table, findings and pins Test summaries, PR URL and the questions for Max follow once the final clean run lands; those sections are added, not rewritten. Co-Authored-By: Claude Opus 5 --- experiments/native-row-ceilings/out.md | 332 +++++++++++++++++++++++++ 1 file changed, 332 insertions(+) create mode 100644 experiments/native-row-ceilings/out.md diff --git a/experiments/native-row-ceilings/out.md b/experiments/native-row-ceilings/out.md new file mode 100644 index 000000000..590edc1e7 --- /dev/null +++ b/experiments/native-row-ceilings/out.md @@ -0,0 +1,332 @@ +# Lane report: the US native build's row-count ceilings + +Branch `native-row-ceilings`, from `origin/native-scale-transport` at +`a64f7b733`. Draft PR against `native-scale-transport`, and it stays draft. + +The design authority is [`docs/us-native-row-ceilings.md`](../../docs/us-native-row-ceilings.md). +This report is the lane's record: what the census found, what moved and why, +which pins moved, what the tests said verbatim, and what is Max's to decide. + +Nothing here is a build, a certification or a release artifact. No gated data was +read. The inputs are one recovered development artifact and one captured public +ACS PUMS archive; every receipt under `experiments/native-row-ceilings/` carries +`"release_eligible": false`. + +## 1. The counts, and why they are measurements + +Every ceiling below is measured against counts taken from the **catalogues** +inside the recovered 1/1000 pilot preparation artifact +(`_recovered/pilot-runs/native19-required-20260912/run/financial-artifacts/preparation.json`, +sha256 `34b362d85d2f06acd255390976d76343aff45114958ba382edb10ffa3789a8a0`), not +from its rosters. + +The distinction decides whether the numbers are measurements or estimates. A +roster at 1/1000 must be scaled to say anything about full source, and scaling a +stratified sample carries error. The catalogues are counts of the whole upstream +ACS and ASEC files and are identical at every fraction. What makes them answer +the full-source question is an identity `roster_census.py` asserts before it +reports anything: + +``` +ACS occupied_hu 1,348,408 + + institutional_gq 84,422 + + noninstitutional_gq 98,784 + + ASEC households 55,762 + = 1,587,376 == selection.supplied_households +``` + +"The whole catalogue" and "what a full-source selection supplies" are the same +set, so the catalogue's counts *are* the full-source counts: + +| quantity | full source | at 1/10 | +|---|---:|---:| +| ACS households (selectable; excludes 100,355 vacancies) | 1,531,614 | 153,161 | +| ACS persons | 3,422,888 | 342,288 | +| ASEC households | 55,762 | 5,576 | +| ASEC persons | 142,125 | 14,212 | +| **stacked households** | **1,587,376** | 158,737 | +| **stacked persons** | **3,565,013** | 356,501 | +| combined-clone households | 3,174,752 | 317,475 | +| combined-clone persons | 7,130,026 | 713,002 | + +**This supersedes the transport lane's figures.** Its §5 quoted ~3,471,000 +stacked and ~6,943,000 cloned persons — the 1/1000 artifact's 2.186869 persons +per household extrapolated. The catalogue figures are about 2.7% higher. No +verdict changes; every new constant is derived from these. + +One split matters because several bounds are met by one channel rather than the +total: `_normalized_source_copy` normalizes the ACS and ASEC native frames +**separately**, and ACS carries 96.0% of the persons. A per-channel bound is met +by 3,422,888, not 3,565,013. + +Receipt: `roster-census.json`, from `roster_census.py`. + +## 2. The one argument + +> **A bound moves only if a full-source native build meets it. A bound that moves +> becomes four times the measured full-source count of exactly what it counts, +> rounded up to the next whole million.** + +The multiple is the transport lane's own — `MAX_ROSTER_BYTES` is 3.9× a +full-source preparation receipt — so both families share one law. The rounding +makes the rule checkable at a glance: divide any moved constant by the count its +comment names and the answer is between 4 and 4.6. + +**Byte transports are not this rule's.** The transport lane's answer to a payload +that outgrows its cap is a segmented stream under one explicit total, not a +larger single cap. Applying this rule to one would contradict the neighbouring +argument rather than extend it — and in the clearest case it would not even be +possible: `graph._bounded_json` opens with +`_require(type(limit) is int and 0 < limit <= 64 * 1024**2, "TRANSPORT_LIMIT")`, +so the shared encoder refuses any cap above 64 MiB before it encodes a byte. + +**For a bound that looks like a row ceiling, the rule applies one test:** does the +binding site stream, or materialise a per-row payload? A bound in front of a +materialised payload cannot usefully be raised alone, because the payload's own +byte cap refuses first and at a smaller number. §4 shows the three bounds that +test decided. + +The rule is also an assertion. `test_us_native_row_ceilings.py` holds the +measured counts and checks each moved ceiling is exactly what the rule produces, +so a later edit that drifts fails a test rather than a build. + +## 3. The census + +`census.json` is the full table: every `MAX_*` row, household, person, group or +byte bound in `packages/microcosm-build/src/microcosm/build/us_runtime/` +reachable from the 19-node financial graph and the 45-node pilot graph, read at +this head — **146 bounds across five module families** — each with its constant, +its enforcement sites, its refusal code *and exception type*, what it protects, +the count it meets at 1/10 and at full source, and whether it binds. + +Every verdict claiming a full-source build meets the bound, or that the bound +guards an encoding width or an upstream file's real size, then went through an +adversarial pass that read the code again and tried to refute it — **42 +verdicts**. Many refined a classification or completed an enforcement list. +**None overturned a headline binding call.** + +| | | +|---|---:| +| bounds censused | 146 | +| bind at full source | **14** | +| bind at 1/10 | **6** | +| verdicts adversarially verified | 42 | +| headline binding calls overturned | 0 | + +**Six bounds bind at 1/10.** That contradicts the transport lane's §5 conclusion +that "a 1/10 build meets no ceiling this lane did not lift". All six are byte +transports: + +| bound | value | admits | of source | +|---|---:|---:|---:| +| `acs_person_coverage_authentication.MAX_BODY_BYTES` | 64 MiB | 12,911 ACS persons | **0.38%** | +| `acs_native_coverage_binding.MAX_EVIDENCE_BYTES` | 2 MiB | — | below 1/10 | +| `acs_housing_universe_source.ACS_HU_RECEIPT_MAX_BYTES` | 1 MiB | — | below 1/10 | +| `survey_origin_budget.MAX_PAYLOAD_BYTES` | 64 MiB | 87,838 households | **5.53%** | +| `graph_survey_population.PREPARATION_MAX_BYTES` | 64 MiB | 96,839 households | **6.10%** | +| `graph_current_survey_household_roles.MAX_ARTIFACT_BYTES` | 64 MiB | — | below 1/10 | + +The eight that bind only above 1/10: `asec_demographic_source._MAX_PERSONS` and +`current_survey_geography.MAX_HOUSEHOLDS` (both moved here), +`current_child_property_income_source.MAX_ROWS`, +`current_survey_household_roles.MAX_PERSONS`, `graph_survey_age_artifact.MAX_BYTES`, +`graph_survey_calibration.MAX_BYTES` and `.MAX_ROWS`, and +`puf55_survey_recipients.RAW_BYTES_MAX_BYTES` (defined in `microcosm-graph`, +outside this lane's scope), plus `puf_diagnostic_consumer.CURRENT_SURVEY_MAX_BYTES`. + +### 3a. `survey_origin_budget.MAX_GROUPS`, established + +The transport lane recorded it "not established". It is established, and it +binds. `_initial` builds groups with + +```python +instructions = graph.allocation_instructions( + view.selection_plan, view.receipt["origins"]["households"] +) +_require(0 < len(instructions) <= MAX_GROUPS, "GROUP_COUNT_BOUND") +``` + +and `allocation_instructions` opens with +`_require(len(household_origins) == len(plan.selected), "ORIGIN_COUNT")`. One +instruction per selected household, so the group count *is* the selected +household count: **1,587,376 at full source against a 1,000,000 bound.** + +Where it runs, stated rather than implied: it is reachable by import from +`graph_atomic_survey_financial`, but **no node in the nineteen-node graph +executes it** — the recovered run's `graph.json` lists all nineteen and none is a +budget node — and none in the completion host does either. +`survey_age_calibration` and `graph_survey_budget` execute it, over the same +full-source selection, so the count and the verdict stand. + +## 4. What moved + +| constant | was | now | counts | full source | × | +|---|---:|---:|---|---:|---:| +| `acs_pums.MAX_EXACT_HOUSEHOLDS` | 1,000,000 | **7,000,000** | exact ACS household keys | 1,531,614 | 4.57 | +| `acs_pums.MAX_EXACT_PERSON_ROWS` | 1,000,000 | **14,000,000** | `NP` over selected ACS households | 3,422,888 | 4.09 | +| `acs_person_coverage_columns.MAX_SELECTED_ROWS` | 1,000,000 | **14,000,000** | requested ACS person keys | 3,422,888 | 4.09 | +| `survey_observed_age.MAX_ROWS` | 2,000,000 | **14,000,000** | one channel's person rows | 3,422,888 | 4.09 | +| `survey_origin_budget.MAX_GROUPS` | 1,000,000 | **7,000,000** | allocation instructions = selected households | 1,587,376 | 4.41 | +| `current_survey_geography.MAX_HOUSEHOLDS` | 524,288 | **7,000,000** | selected households | 1,587,376 | 4.41 | +| `asec_demographic_source._MAX_PERSONS` | 600,000 | **14,000,000** | retained ACS persons (the larger of its two rosters) | 3,422,888 | 4.09 | + +Every refusal keeps its code, its exception type and its expression. Only the +number moves. + +| constant | refusal | raises | +|---|---|---| +| `MAX_EXACT_HOUSEHOLDS` | `"ACS exact selection requires bounded unique raw native keys."` | `ValueError` | +| `MAX_EXACT_PERSON_ROWS` | `"ACS selected complete roster exceeds native person budget."` | `ValueError` | +| `MAX_SELECTED_ROWS` | `"ACS coverage selected person count is outside the bound"`; `"SELECTED_ROWS"`, `"NATIVE_ROWS"`, `"NATIVE_ROW_BUDGET"` downstream | `ValueError`, `ACSCoverageAuthenticationError`, `ACSNativeCoverageBindingError` | +| `survey_observed_age.MAX_ROWS` | `"SURVEY_OBSERVED_AGE_ROW_BOUND"` | `ValueError` | +| `MAX_GROUPS` | `"GROUP_COUNT_BOUND"` | `SurveyOriginBudgetError` | +| `current_survey_geography.MAX_HOUSEHOLDS` | `"CURRENT_SURVEY_GEOGRAPHY_HOUSEHOLD_COUNT"`, `…_PROJECTION_STORAGE` | `ValueError` | +| `asec_demographic_source._MAX_PERSONS` | `"ACS_ROWS"`, `"MEMBERSHIP_ROWS"`, `"CLASSIFY_ROWS"`, `"DEMOGRAPHIC_ROWS"` | `ValueError` | + +The last two were not in the brief and were moved only after reading what stood +behind them: + +- **`current_survey_geography.MAX_HOUSEHOLDS` bound hardest of the seven** — + 524,288 against 1,587,376, refusing at 33% of source. It reads as a byte budget + (`64 * 1024**2 // 128`) but **streams**: `_projection_digest` feeds one bounded + row encoding at a time into a `hashlib.sha256`, the module materialises nothing + per household, and its only bytes are a 64 KiB summary receipt already bounded + by `MAX_RECEIPT_BYTES`. The borrowed 64 MiB never described anything here. +- **`asec_demographic_source._MAX_PERSONS` is one constant over two rosters**, so + it takes the larger. It is not a fixed-width encoding: `rows` is a JSON header + integer, the only `struct.pack` is `".read_bytes())`, the same call the check makes, via `regenerate_accepted_pin.py --write` | +| the other 123 inventory contracts | — | **unchanged** | same generator, over all of them | +| `_ACCEPTED["acs_inputs.py"]`, `["acs_housing_universe_source.py"]`, `["acs_person_coverage_authentication.py"]` | — | **unchanged** | same | +| all ten `implementation_manifest(stage)` | — | **built** | `graph_implementation.implementation_manifest(stage)` | + +**The first row is not this lane's, and it is a blocker on the base branch.** On +`native-scale-transport`, `b6081efcb` ("Hash a spill segment that was already +there") added `path.read_bytes()` inside `_spill_roster`. `read_bytes` is in +`graph_implementation._RESOURCE_CALLS`, so it moves that module's +`resource_accesses_sha256` — and the pin was not regenerated. +`implementation_manifest()` therefore raised `"Unclassified US +dependency/resource contract"` for **every stage containing that module**, +including `authenticated_survey_population_v1`, the stage the nineteen-node path +runs. + +Proven pre-existing two ways. The contract was recomputed from the blobs of +`origin/native-scale-transport`, `a64f7b733`, `b6081efcb` and its four +predecessors: it matches at `a51957c5c` and earlier and mismatches from +`b6081efcb` on. And the affected tests were run against a base worktree at +`a64f7b733` — **7 failed there with that exact `ValueError`, and all 7 pass +here.** This branch does not touch that file. + +The transport lane's report §4 was right that `Path.write_bytes` is invisible to +`resource_accesses_sha256` by construction. This is the converse: a later commit +on the same branch added a *read*, which is visible. + +What moves and is not a committed pin: each US stage's `implementation_hash` is +over its whole module roster, so editing any inventoried module moves it and with +it every node key and store address. #935 recorded that this holds for any change +to these files including a comment. From 47dbe457f05e514242ed758ce3f2db763efeee8d Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:17:44 -0400 Subject: [PATCH 19/22] Say plainly why the census reports fourteen and the base branch has nineteen The census agents read a moving tree: five constants were lifted while it ran and the agents correctly reported the post-lift values, recording both numbers per row. At base nineteen bind; seven moved here; twelve remain for the transport argument. Also corrects 41 verdicts to 42. Co-Authored-By: Claude Opus 5 --- docs/us-native-row-ceilings.md | 26 ++++++++++++++++++-------- experiments/native-row-ceilings/out.md | 12 +++++++++++- 2 files changed, 29 insertions(+), 9 deletions(-) diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md index 1af350114..8f354b787 100644 --- a/docs/us-native-row-ceilings.md +++ b/docs/us-native-row-ceilings.md @@ -189,10 +189,17 @@ meets; it does not, and this note corrects that.** ## 5. What still binds, and whose argument it is -The census in §6 found fourteen bounds a full-source build meets. Seven are the -row counts §3 moved. The rest belong to the transport lane's argument — a +**Nineteen bounds bind at full source at the base branch's values.** §3 moved +seven of them. The other twelve belong to the transport lane's argument — a segmented stream under one explicit total — not to this one. +(The census in §6 reports fourteen rather than nineteen because its agents read a +moving tree: five of the seven were lifted while it ran, and they correctly +report the post-lift values at which those five no longer bind. Each of those +rows records both numbers. The two the census still shows as binding — +`current_survey_geography.MAX_HOUSEHOLDS` and +`asec_demographic_source._MAX_PERSONS` — were moved afterwards, on its evidence.) + ### 5a. The test the rule applies Three of the remaining bounds *look* like row ceilings, and one of them was @@ -324,12 +331,15 @@ refusal code and exception type, what it protects, the count it meets at 1/10 and at full source, and whether it binds. Every verdict that claimed a full-source build meets the bound, or that the bound guards an encoding width or an upstream file's real size, then went through an adversarial pass that read the -code again and tried to refute it — **41 verdicts**, and no headline verdict was -overturned. - -The result: **14 bounds bind at full source, 6 of them at 1/10.** Seven of the -fourteen are the row counts §3 moved. The other seven are byte transports and -byte-derived row pre-checks; §5a and §5b say which and why. +code again and tried to refute it — **42 verdicts**, and no headline binding call +was overturned. + +The result, read as the agents read it: **14 bounds bind at full source, 6 of +them at 1/10.** Five of the seven §3 moved are absent from that fourteen because +they were lifted mid-census and the agents correctly reported the post-lift +values; each of those rows carries both numbers. **At the base branch's values +the figure is nineteen**, of which §3 moved seven, leaving twelve for the +transport argument. §5a and §5b say which and why. ## 7. Pins diff --git a/experiments/native-row-ceilings/out.md b/experiments/native-row-ceilings/out.md index 590edc1e7..8fdd6f1b4 100644 --- a/experiments/native-row-ceilings/out.md +++ b/experiments/native-row-ceilings/out.md @@ -108,11 +108,21 @@ verdicts**. Many refined a classification or completed an enforcement list. | | | |---|---:| | bounds censused | 146 | -| bind at full source | **14** | +| bind at full source, **at the base branch's values** | **19** | +| bind at full source, as `census.json` reports them | 14 | | bind at 1/10 | **6** | | verdicts adversarially verified | 42 | | headline binding calls overturned | 0 | +**Why the two figures differ, stated rather than reconciled quietly.** The census +agents read a moving tree: five of the seven constants §4 moved were lifted while +the census ran, and the agents correctly reported the *post-lift* values at which +those five no longer bind — each of those rows records both numbers and the +commit that changed it. The two that `census.json` still shows as binding, +`current_survey_geography.MAX_HOUSEHOLDS` and `asec_demographic_source._MAX_PERSONS`, +were moved afterwards on its evidence. So: nineteen bind at base, seven moved +here, twelve remain for the transport argument. + **Six bounds bind at 1/10.** That contradicts the transport lane's §5 conclusion that "a 1/10 build meets no ceiling this lane did not lift". All six are byte transports: From f40294e91f67f3379dd12d8e82c1d271b1fe5ed2 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:26:27 -0400 Subject: [PATCH 20/22] Replace an estimate with a measurement, and two wordings with what the code does I wrote in the design note that the roles projection costs "roughly 200 bytes of JSON per person" and admits "a few hundred thousand". That was reasoned, not computed, and this lane's own standard forbids it. Measured through the module's own encoder: 320.08 bytes per person row, so the 64 MiB artifact cap admits 209,661 persons -- 5.88% of source, against a row bound that admits 58.8%. The byte cap refuses ten times earlier, which is exactly why MAX_PERSONS is not this rule's to move. Also two precision fixes. The financial graph DOES call into survey_origin_budget -- for _config_payload and the _live() producer seal -- so "no node executes it" becomes "no node executes freeze_survey_origin_budget, where GROUP_COUNT_BOUND is checked, and neither of those calls reaches _initial". And MAX_SELECTED_ROWS at 14,000,000 now sits above MAX_ROWS at 6,000,000 in the same module, which is not an inconsistency -- one bounds caller-supplied keys, the other the archive's own rows -- but does mean the effective ceiling there is 6,000,000 by containment. Both said in the note rather than left for a reviewer to find. Co-Authored-By: Claude Opus 5 --- docs/us-native-row-ceilings.md | 34 +++++-- experiments/native-row-ceilings/out.md | 14 +-- .../roles-projection-size.json | 19 ++++ .../roles_projection_size.py | 95 +++++++++++++++++++ pyproject.toml | 1 + 5 files changed, 148 insertions(+), 15 deletions(-) create mode 100644 experiments/native-row-ceilings/roles-projection-size.json create mode 100644 experiments/native-row-ceilings/roles_projection_size.py diff --git a/docs/us-native-row-ceilings.md b/docs/us-native-row-ceilings.md index 8f354b787..8d394e012 100644 --- a/docs/us-native-row-ceilings.md +++ b/docs/us-native-row-ceilings.md @@ -115,6 +115,15 @@ persons. A bound met per channel is met by 3,422,888, not by 3,565,013. Each refusal keeps its code, its exception type and its expression; only the number moves. +**One consequence, named so a reader does not have to find it.** +`MAX_SELECTED_ROWS` at 14,000,000 now sits *above* `MAX_ROWS` at 6,000,000 in the +same module. That is not an inconsistency: `MAX_SELECTED_ROWS` bounds +*caller-supplied* person keys, which can exceed what an archive holds, while +`MAX_ROWS` bounds the archive's own rows. But it does mean the effective ceiling +on that path is 6,000,000, by containment. Setting `MAX_SELECTED_ROWS = MAX_ROWS` +would be tighter *and* provable rather than chosen — and it would make the rule +two rules, which is why this note keeps one. + | constant | refusal | raises | |---|---|---| | `MAX_EXACT_HOUSEHOLDS` | `"ACS exact selection requires bounded unique raw native keys."` | `ValueError` | @@ -142,12 +151,14 @@ and `allocation_instructions` opens with instruction per selected household, so the group count *is* the selected household count: 1,587,376 at full source against a 1,000,000 bound. -Two things about where it runs, stated rather than implied. It is reachable by -import from `graph_atomic_survey_financial`, but **no node in the nineteen-node -graph executes it** — the recovered run's `graph.json` lists all nineteen and -none is a budget node — and none in the completion host does either. -`survey_age_calibration` and `graph_survey_budget` execute it, over the same -full-source selection, so the count and the verdict are unchanged. +Where it runs, stated precisely rather than implied. **No node in the +nineteen-node graph executes `freeze_survey_origin_budget`**, which is where +`GROUP_COUNT_BOUND` is checked — the recovered run's `graph.json` lists all +nineteen and none is a budget node, and the completion host has none either. +`graph_atomic_survey_financial` does call into the module, but only for +`_config_payload` and the `_live()` producer seal, neither of which reaches +`_initial`. `survey_age_calibration` and `graph_survey_budget` are what execute +it, over the same full-source selection, so the count and the verdict stand. ## 4. Why a bound stays a bound, and why some must not move at all @@ -218,9 +229,14 @@ bound in front of a streaming digest has nothing behind it, and the rule applies | `current_survey_household_roles.MAX_PERSONS` | 2,097,152 | `_projection_bytes` is `table.reset_index().to_json(orient="table").encode()` — the whole per-person table in one string — under `graph_current_survey_household_roles.MAX_ARTIFACT_BYTES` = 64 MiB | materialises — transport's | | `current_child_property_income_source.MAX_ROWS` | 600,000 | `_json` is one `json.dumps(value)` under `MAX_PROJECTION_BYTES` = 64 MiB | materialises — transport's | -The middle row is the clearest case for why the test matters: at roughly 200 -bytes of JSON per person, that 64 MiB cap admits a few hundred thousand persons, -so raising the 2,097,152 row bound would move nothing at all. +The middle row is the clearest case for why the test matters, and it is measured +rather than reasoned: through the module's own encoder, a person row costs +**320.08 bytes**, so that 64 MiB cap admits **209,661 persons — 5.88% of source** +against a row bound that admits 58.8%. The byte cap refuses **ten times earlier**, +so raising `MAX_PERSONS` alone would move nothing at all. +`experiments/native-row-ceilings/roles_projection_size.py` measures it; the table +it encodes is invented but faithfully shaped, and the per-row cost is a +difference between two row counts so the fixed schema preamble cancels. A fourth bound needed the same test plus one more question, and passed both. `asec_demographic_source._MAX_PERSONS` (600,000) is enforced at five sites that diff --git a/experiments/native-row-ceilings/out.md b/experiments/native-row-ceilings/out.md index 8fdd6f1b4..abebe5a85 100644 --- a/experiments/native-row-ceilings/out.md +++ b/experiments/native-row-ceilings/out.md @@ -161,12 +161,14 @@ and `allocation_instructions` opens with instruction per selected household, so the group count *is* the selected household count: **1,587,376 at full source against a 1,000,000 bound.** -Where it runs, stated rather than implied: it is reachable by import from -`graph_atomic_survey_financial`, but **no node in the nineteen-node graph -executes it** — the recovered run's `graph.json` lists all nineteen and none is a -budget node — and none in the completion host does either. -`survey_age_calibration` and `graph_survey_budget` execute it, over the same -full-source selection, so the count and the verdict stand. +Where it runs, stated precisely. **No node in the nineteen-node graph executes +`freeze_survey_origin_budget`**, which is where `GROUP_COUNT_BOUND` is checked — +the recovered run's `graph.json` lists all nineteen and none is a budget node, +and the completion host has none either. `graph_atomic_survey_financial` does +call into the module, but only for `_config_payload` and the `_live()` producer +seal, neither of which reaches `_initial`. `survey_age_calibration` and +`graph_survey_budget` are what execute it, over the same full-source selection, +so the count and the verdict stand. ## 4. What moved diff --git a/experiments/native-row-ceilings/roles-projection-size.json b/experiments/native-row-ceilings/roles-projection-size.json new file mode 100644 index 000000000..288aceb7d --- /dev/null +++ b/experiments/native-row-ceilings/roles-projection-size.json @@ -0,0 +1,19 @@ +{ + "scope": "Encoder-measured size law for the current-survey household-roles projection, on a faithfully shaped invented table. Not a build, not a certification, not release eligible.", + "release_eligible": false, + "encoder": "table.reset_index().to_json(orient=\"table\", index=False).encode()", + "max_artifact_bytes": 67108864, + "max_persons_row_bound": 2097152, + "projection_bytes": { + "3": 1516, + "1203": 385615 + }, + "marginal_bytes_per_person_row": 320.08, + "persons_the_byte_cap_admits": 209661, + "persons_the_row_bound_admits": 2097152, + "byte_cap_is_tighter_by": 10.0, + "full_source_stacked_persons": 3565013, + "fraction_of_source_the_byte_cap_admits": 0.058810725234382036, + "fraction_of_source_the_row_bound_admits": 0.5882592854500109, + "conclusion": "The byte cap refuses first and by an order of magnitude, so raising MAX_PERSONS alone would move nothing. This bound takes the segmented-transport argument, not the row-ceiling rule." +} diff --git a/experiments/native-row-ceilings/roles_projection_size.py b/experiments/native-row-ceilings/roles_projection_size.py new file mode 100644 index 000000000..f39f37b7b --- /dev/null +++ b/experiments/native-row-ceilings/roles_projection_size.py @@ -0,0 +1,95 @@ +"""Measure why the roles row bound cannot usefully be raised on its own. + +``current_survey_household_roles._projection_bytes`` is +``table.reset_index().to_json(orient="table", index=False).encode()`` -- the +whole per-person table materialised as one JSON string -- and the artifact that +carries it is bounded by ``graph_current_survey_household_roles.MAX_ARTIFACT_BYTES`` +(64 MiB). So the module's effective ceiling is that byte cap, not its +``MAX_PERSONS`` row bound, and this measures the gap through the module's own +encoder. + +The table is invented but faithfully shaped: the real declared columns, the real +dtypes, and role/universe tokens of the lengths the module emits. The marginal +cost per row is taken as a difference between two row counts so the fixed +``orient="table"`` schema preamble cancels. + +Nothing outside the output path is written. Not a build, not a certification, +not release eligible. + + python roles_projection_size.py +""" + +from __future__ import annotations + +import json +import pathlib +import sys + +import pandas as pd + +ROOT = pathlib.Path(__file__).resolve().parents[2] +sys.path[:0] = [str(path) for path in sorted((ROOT / "packages").glob("*/src"))] + +from microcosm.build.us_runtime import current_survey_household_roles as roles +from microcosm.build.us_runtime import graph_current_survey_household_roles as graph + +for _module in (roles, graph): + if not pathlib.Path(_module.__file__).resolve().is_relative_to(ROOT): + raise SystemExit(f"{_module.__name__} resolved outside {ROOT}") + +STACKED_PERSONS = 3_565_013 # measured; see roster-census.json + + +def _table(rows: int) -> pd.DataFrame: + pattern = { + roles.CANONICAL_COLUMN: [True, False, True], + roles.SURVEY_COLUMN: ["acs", "asec", "acs"], + roles.NATIVE_ID_COLUMN: [1, 2, 3], + roles.CODE_COLUMN: [20, 25, 37], + roles.CODE_KNOWN_COLUMN: [True, True, False], + roles.ROLE_STATE_COLUMN: ["observed_reference_person"] * 3, + roles.UNIVERSE_COLUMN: ["housing_unit"] * 3, + } + frame = pd.concat([pd.DataFrame(pattern)] * ((rows + 2) // 3)).head(rows) + frame.index = pd.Index(range(rows), name="person_id") + return frame + + +def main() -> int: + out = pathlib.Path(sys.argv[1]) + small, large = 3, 1_203 + bytes_small = len(roles._projection_bytes(_table(small))) + bytes_large = len(roles._projection_bytes(_table(large))) + per_row = (bytes_large - bytes_small) / (large - small) + admitted = int(graph.MAX_ARTIFACT_BYTES // per_row) + record = { + "scope": ( + "Encoder-measured size law for the current-survey household-roles " + "projection, on a faithfully shaped invented table. Not a build, not a " + "certification, not release eligible." + ), + "release_eligible": False, + "encoder": 'table.reset_index().to_json(orient="table", index=False).encode()', + "max_artifact_bytes": graph.MAX_ARTIFACT_BYTES, + "max_persons_row_bound": roles.MAX_PERSONS, + "projection_bytes": {str(small): bytes_small, str(large): bytes_large}, + "marginal_bytes_per_person_row": round(per_row, 2), + "persons_the_byte_cap_admits": admitted, + "persons_the_row_bound_admits": roles.MAX_PERSONS, + "byte_cap_is_tighter_by": round(roles.MAX_PERSONS / admitted, 2), + "full_source_stacked_persons": STACKED_PERSONS, + "fraction_of_source_the_byte_cap_admits": admitted / STACKED_PERSONS, + "fraction_of_source_the_row_bound_admits": roles.MAX_PERSONS / STACKED_PERSONS, + "conclusion": ( + "The byte cap refuses first and by an order of magnitude, so raising " + "MAX_PERSONS alone would move nothing. This bound takes the " + "segmented-transport argument, not the row-ceiling rule." + ), + } + out.write_text(json.dumps(record, indent=1) + "\n") + print(json.dumps(record, indent=1)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/pyproject.toml b/pyproject.toml index 13510f150..2ddce6695 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -44,6 +44,7 @@ ignore = ["E501"] "experiments/native-row-ceilings/origin_budget_size.py" = ["E402"] "experiments/native-row-ceilings/selected_body_budget.py" = ["E402"] "experiments/native-row-ceilings/consumer_gap.py" = ["E402"] +"experiments/native-row-ceilings/roles_projection_size.py" = ["E402"] [tool.ruff.lint.isort] # microcosm is a PEP 420 namespace package (no top-level __init__.py), so From 431f4455bfb47c3f48008ec6cc49513eee645f7b Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:26:57 -0400 Subject: [PATCH 21/22] Add the verbatim gates and the questions for Max to the lane report 399 passed over the seven touched test files at a fixed commit with the tree untouched for the run; ci_test_groups --verify ok; the stdlib matrix contract's 15 tests OK; spec-engine coverage 42156/42156 and 41/41; ruff clean and 21 files already formatted; all three pin generators idempotent. Records honestly that an earlier full-file run reported 8 failed while this session was editing the tree underneath it, and that it measured nothing -- the same tests pass at a fixed commit. And records the base-worktree run that proves the inherited pin break is the base's: 7 failed there with the exact Unclassified-contract ValueError, all 7 pass here. Co-Authored-By: Claude Opus 5 --- experiments/native-row-ceilings/out.md | 163 +++++++++++++++++++++++++ 1 file changed, 163 insertions(+) diff --git a/experiments/native-row-ceilings/out.md b/experiments/native-row-ceilings/out.md index abebe5a85..f1e309d90 100644 --- a/experiments/native-row-ceilings/out.md +++ b/experiments/native-row-ceilings/out.md @@ -342,3 +342,166 @@ What moves and is not a committed pin: each US stage's `implementation_hash` is over its whole module roster, so editing any inventoried module moves it and with it every node key and store address. #935 recorded that this holds for any change to these files including a comment. + +## 8. Gates, verbatim + +All at `f40294e91`'s tree, with the tree untouched for the duration of the +pytest run. Every command's own last line, quoted rather than summarised. + +**The touched test files, as CI runs them — explicit flat file list, no `-k`:** + +``` +$ .venv/bin/python -m pytest \ + packages/microcosm-build/tests/test_us_native_row_ceilings.py \ + packages/microcosm-build/tests/test_us_acs_pums.py \ + packages/microcosm-build/tests/test_us_acs_person_coverage_columns.py \ + packages/microcosm-build/tests/test_us_survey_observed_age.py \ + packages/microcosm-build/tests/test_us_current_survey_geography.py \ + packages/microcosm-build/tests/test_us_asec_demographic_source.py \ + packages/microcosm-build/tests/test_us_survey_origin_budget.py -rf + +399 passed in 481.27s (0:08:01) +``` + +**The CI group gates:** + +``` +$ .venv/bin/python tools/ci_test_groups.py --verify +verification=ok +exit=0 + +$ python3 -I -B -S packages/microcosm-build/tests/test_ci_test_groups.py +Ran 15 tests in 0.396s + +OK + +$ .venv/bin/python tools/spec_engine_coverage.py --check +spec-engine coverage: 42156/42156 configuration fields; 41/41 inventory checks +exit=0 +``` + +`test_us_native_row_ceilings.py` lands in `rest` shard 3/6 (fast), `us-not` +shard 1/1 (engine) and `wheels`, and never under `[defaulted]`. The other six +touched files keep their existing groups: `us-am` for the four `a`–`m` files, +`us-qs` for the two `s` files. + +**Lint and format, over every `.py` this branch touches:** + +``` +$ git diff --name-only origin/native-scale-transport...HEAD | grep '\.py$' \ + | xargs .venv/bin/python -m ruff check +All checks passed! +exit=0 + +$ git diff --name-only origin/native-scale-transport...HEAD | grep '\.py$' \ + | xargs .venv/bin/python -m ruff format --check +21 files already formatted +exit=0 +``` + +**The pin generators, re-run at this head:** + +``` +$ .venv/bin/python experiments/native-row-ceilings/repin.py .../repin.json +inventory contracts checked: 124; moved: 0 +acs_native_coverage_binding._ACCEPTED: + acs_pums.py: unchanged + acs_inputs.py: unchanged + acs_housing_universe_source.py: unchanged + acs_person_coverage_authentication.py: unchanged +stage manifests built: 10 + +$ .venv/bin/python experiments/native-row-ceilings/regenerate_inventory_contract.py +no contract moved + +$ .venv/bin/python experiments/native-row-ceilings/regenerate_accepted_pin.py +no pin moved +``` + +**One earlier run reported failures, and it was not a valid measurement.** A +full-file run of `test_us_survey_origin_budget.py` reported `8 failed, 44 +passed` while this session was editing the working tree underneath it. The same +tests pass at a fixed commit — the subset run returned `8 passed, 44 deselected +in 196.03s`, and the run above is green over the whole file. It is recorded here +because it was run and reported, not because it measured anything. + +**Against the base branch, for the inherited pin.** Seven of those tests were run +against a worktree at `a64f7b733` with `PYTHONPATH` pointing at its own packages: + +``` +7 failed, 44 deselected in 59.22s +E ValueError: Unclassified US dependency/resource contract: + microcosm.build/us_runtime/survey_population_preparation.py. +``` + +All seven pass on this branch. That is the evidence in §7 that the break is the +base's and not this lane's. + +## 9. For Max + +**1. The inherited re-pin — keep it here, or move it to #945?** +`b6081efcb` on `native-scale-transport` added `path.read_bytes()` to +`_spill_roster` without regenerating the pin, and `implementation_manifest()` +raises for the stage the nineteen-node path runs. Seven tests fail on a base +worktree with that exact error; all seven pass here. + +- **(a) Keep it in this PR.** One generated line; it makes the base functional + and lets this lane build its own stage manifests. Costs: this PR touches a + file that lane owns. +- **(b) Move it to #945 and rebase this branch.** Cleaner ownership. Costs: this + branch cannot build a stage manifest until #945 carries the fix, so its pin + evidence is unverifiable in the meantime. +- **(c) Keep it here *and* tell #945's owner**, so the fix is not silently + inherited and re-derived twice. + +I shipped (a) and would pick (c). I did not pick (b) because required step 4 of +this lane's brief is to re-derive every pin through its generator, and the +generator raises on an unfixed base. + +**2. The twelve byte transports that still bind — one lane or several?** +Six of them bind at **1/10**, and `acs_person_coverage_authentication.MAX_BODY_BYTES` +binds at **0.38% of source**. They are all the same shape and the argument for +them already exists, written by the transport lane. + +- **(a) One follow-up lane, one argument**, carrying the segmented transport into + all of them — the same reasoning that produced this lane. +- **(b) One lane per module, in binding order**: the ACS body budget (0.38%), the + preparation consumer (6.10%), the origin budget (5.53%), the roles artifact, + the calibration pair, the age artifact. +- **(c) Only the ones on the nineteen-node path now**, leaving the calibration and + age-artifact ones to whoever meets them. + +My reading is (a), for the reason you already gave on this family. But the ACS +body budget is different in kind from the rest — it refuses at 1/265, so it +gates *any* run above the 1/1000 pilot, and it may deserve to go first whatever +the shape of the rest. + +**3. Seven moved, not the four the brief named. Keep them together?** +The brief named four; `MAX_GROUPS` it asked me to establish, and establishing it +showed it binds. The census then found two more that the rule decides with no +judgment left — `current_survey_geography.MAX_HOUSEHOLDS`, which bound hardest of +all seven at 33% of source, and `asec_demographic_source._MAX_PERSONS`. + +- **(a) Keep all seven in this PR.** One rule applied wherever it decides. +- **(b) Split the two the brief did not name into their own PR.** + +I shipped (a): splitting them is the "one build at a time" you ruled out, and +both were moved on the same read-the-code test as the rest, recorded in §4. + +**4. The multiple — and one place a tighter rule exists.** +I used four times the measured full-source count, rounded up to the next whole +million, because that is the transport lane's own headroom (3.9×) and it keeps +one law over both families. + +- **(a) Keep 4×.** +- **(b) A larger multiple** — 10× would be about three decades of ACS growth + rather than one, at no runtime cost, since these are ceilings and not + allocations. +- **(c) Where a containing bound already exists, use it instead.** + `MAX_SELECTED_ROWS` bounds requested person keys and `MAX_ROWS` bounds the + archive those keys are drawn from, so `MAX_SELECTED_ROWS = MAX_ROWS` would be + tighter than 14,000,000 *and* provable by containment rather than chosen. It + would make the rule two rules, which is why I did not. + +I shipped (a). (c) is the only one I think is genuinely arguable, and only for +that one pair. From c5ac78d2decd18fef5d636ab6bb9b589d6ec7e7b Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Thu, 17 Sep 2026 18:27:47 -0400 Subject: [PATCH 22/22] Record the pull request, the file map, and what a reader should not take PR #949, draft against native-scale-transport, MERGEABLE. CI does not run on it by design: test.yml triggers on pull_request: branches: [main], and gh pr checks reports none, so section 8's local gates are the only ones this branch has. Section 11 says what the report does not claim -- including that census.json's per-row prose is the agents', checked but not rewritten, so a row may carry a superseded classification beside its verified binding call; and that the roles projection figure is on a faithfully shaped invented table, not a real artifact. Co-Authored-By: Claude Opus 5 --- PROGRESS-native-row-ceilings.md | 11 ++++- experiments/native-row-ceilings/out.md | 58 ++++++++++++++++++++++++++ 2 files changed, 67 insertions(+), 2 deletions(-) diff --git a/PROGRESS-native-row-ceilings.md b/PROGRESS-native-row-ceilings.md index cf5226c74..fe0985017 100644 --- a/PROGRESS-native-row-ceilings.md +++ b/PROGRESS-native-row-ceilings.md @@ -81,8 +81,15 @@ the reconciliation. All three are byte transports and take the transport lane's argument, not this one. All three are pinned in tests. -## Next +## State + +Complete. **[PR #949](https://github.com/PolicyEngine/microcosm/pull/949)**, +draft against `native-scale-transport`, MERGEABLE. CI does not run on it by +design (`test.yml` triggers on `pull_request: branches: [main]`); the local +gates in the report's section 8 are the only ones it has, and all are green: +399 passed over the seven touched test files, `verification=ok`, spec-engine +42156/42156 and 41/41, ruff clean, all three pin generators idempotent. -- Final clean test run, then the draft PR against `native-scale-transport`. +## Next - Open for Max: whether the inherited re-pin stays here or moves to #945; and whether the byte transports above are one follow-up lane or several. diff --git a/experiments/native-row-ceilings/out.md b/experiments/native-row-ceilings/out.md index f1e309d90..d39c8444a 100644 --- a/experiments/native-row-ceilings/out.md +++ b/experiments/native-row-ceilings/out.md @@ -505,3 +505,61 @@ one law over both families. I shipped (a). (c) is the only one I think is genuinely arguable, and only for that one pair. + +## 10. The pull request + +| | | +|---|---| +| PR | **[PolicyEngine/microcosm#949](https://github.com/PolicyEngine/microcosm/pull/949)** — draft, and it stays draft | +| Title | Lift the seven row-count ceilings a full-source native build meets, under one rule | +| Base | `native-scale-transport` (PR #945's branch), at `a64f7b733` | +| Head | what `git rev-parse native-row-ceilings` returns; this table does not quote a sha it cannot have written | +| Mergeable | `MERGEABLE`, verified at the head this report was written against | +| `packages/microcosm-graph` hunks | **zero** | +| CI | **does not run on this PR by design.** `.github/workflows/test.yml` triggers on `pull_request: branches: [main]`, so only a PR targeting `main` reaches it, and `gh pr checks 949` reports none. §8 is the only gate this branch has, and it was run locally. | + +Files, against the base: + +| file | what | +|---|---| +| `.../us_runtime/acs_pums.py` | the two exact-selection ceilings | +| `.../us_runtime/acs_person_coverage_columns.py` | the requested-roster ceiling; `MAX_ROWS` deliberately unchanged, with the reason in a comment | +| `.../us_runtime/survey_observed_age.py` | the per-channel row ceiling | +| `.../us_runtime/survey_origin_budget.py` | `MAX_GROUPS`; `MAX_PAYLOAD_BYTES` deliberately unchanged, with the reason in a comment | +| `.../us_runtime/current_survey_geography.py` | the household ceiling, and why its byte-derived form never described anything here | +| `.../us_runtime/asec_demographic_source.py` | the person ceiling over its two rosters | +| `.../us_runtime/acs_native_coverage_binding.py` | the one moved pin, generated | +| `.../us_runtime/graph_implementation_inventory.json` | the **inherited** re-pin, generated | +| `packages/microcosm-build/tests/test_us_native_row_ceilings.py` | the rule as an executable table, the bounds that must not move, and the three consumer ceilings pinned | +| six existing `test_us_*.py` | one boundary test per moved bound | +| `docs/us-native-row-ceilings.md` | the design authority | +| `experiments/native-row-ceilings/` | five measurement tools, three pin tools, and their receipts | +| `changelog.d/native-row-ceilings.changed.md` | the towncrier fragment | +| `pyproject.toml` | four per-file `E402` ignores, for tools that must set `sys.path` before importing what they measure | +| `PROGRESS-native-row-ceilings.md` | the lane journal | + +## 11. What a reader should not take from this report + +- **Nothing here is a build, a certification or a release artifact.** Every + receipt carries `"release_eligible": false` and a scope line. No gated data was + read: the inputs are one recovered development artifact and one captured + **public** ACS PUMS archive, both read by path, neither written, moved or + linked. +- **The census agents read a moving tree.** Five constants were lifted while the + census ran; those rows record both values and the commit that changed them, and + §3 reconciles the two totals rather than quoting one. +- **`census.json`'s per-row prose is the agents', checked but not rewritten.** + Every verdict that claimed a bound binds at full source, or guards an encoding + width or an upstream file's size, was adversarially re-read — 42 of them — and + several classifications were corrected in that pass. The rows themselves were + not edited afterwards, so a row may still carry a superseded classification + beside its verified `binds_at_full_source`. The verified field is the one this + report relies on, and every number this report or the design note states was + re-read at this head before being written down. +- **The roles projection figure is on an invented table.** It is faithfully + shaped — the module's own declared columns and dtypes, its own encoder — and the + per-row cost is a difference between two row counts so the schema preamble + cancels. The 320.08 bytes is not a measurement of a real roles artifact. +- **The 12,911 and 87,838 figures are of a bound, not of a run.** They say what + the ceiling admits, computed from a measured per-row cost. No build was run at + any fraction in this lane.