Skip to content

US native: persistent content-addressed memo for ACS source derivations (acceleration C) - #1002

Draft
MaxGhenis wants to merge 517 commits into
mainfrom
native-source-auth-memo-20260923
Draft

MaxGhenis wants to merge 517 commits into
mainfrom
native-source-auth-memo-20260923

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

ACS source authentication repeatedly decompresses and parses the same pinned
archives during native preparation and later runs. This adds an opt-in,
persistent memo for six deterministic ACS derivations, keyed by exact source
bytes, parameters, runtime and owner code. Issuance, source pin checks, live
object validation and release gates still run.

The memo authenticates its index with a private key and verifies blob digests
before decoding. Invalid entries recompute; owner refusals are never cached.
Frame results preserve numeric bits, string storage, object scalar types and
missing markers. Unsupported shapes bypass storage. Live pandas parser
functions, reader dispatch and defaults are checked before a frame-cache hit,
so replacing the parser cannot conceal a current refusal behind an old result.
The graph inventory includes the memo dependency in every affected ACS stage,
and accepted owner digests are refreshed for these source changes.

Validation uses invented archives and frames. See
the experiment report for the
exact commands and observed results, including off/cold/warm equality,
trace-proven parser skips, source/parameter invalidation, parser replacement
refusals and cache tampering controls.

The final combined regression run passed 206 tests; all 14 dependency checks
also passed. Earlier targeted ACS loader, literal coverage and codec checks
passed 201 cases (overlapping codec coverage). Ruff lint/format and diff checks
passed. These are local results from the resumed working tree.

The actual-archive owner benchmark and 1/1000 graph measurement remain pending.
The resumed session could not inspect processes (ps: Operation not permitted),
so it did not satisfy the required no-competing-native-run check. No actual-data
speedup is claimed. Run the benchmark only with at least 40 GiB available and
no other native run active; compare all output digests and counts before adding
timings. These fixture checks do not certify data artifacts. Keep this PR draft.

Tracks #956, acceleration C.

🤖 Generated with Claude Code

MaxGhenis and others added 30 commits September 13, 2026 21:45
The fixture registered its two recipient kernels into the issued financial
run's own registry, which the run seals at issuance and re-checks in
check_atomic_survey_financial_run, so every test errored with
FINANCIAL_KERNEL_REGISTRY_CHANGED. Production (graph_survey_puf55)
composes a private KernelRegistry over the run's kernels; the fixture now
does the same. The last-owner mutation test also accepts the
node-population refusal that now fires first, as its siblings do.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…time

extended_status_host recorded its student _MEMBER_PINS patch on the outer
MonkeyPatch context after the generator had already patched the same
attribute, so on exit it restored the generator-era pins over the value
the module's first fixture installed, and that fixture's retained student
receipts failed STUDENT_CONTENT_CHANGED at its own teardown (attributed by
pytest to the last test, tax42). The patch now lives in its own context
nested inside the generator's lifetime, so undo order is LIFO.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One rest job and both wheels 2/4 jobs hit GitHub's six-hour limit at 76-79%
of their shards (the alphabetically late survey-graph modules cluster there);
the partition is by file count, so a few heavy modules dominate a job. Six
and eight shards spread the heavy modules 13 and 9-10 per job instead of up
to 23, and every sharded pytest run now prints --durations=25 so the next
imbalance is visible in the job log. tools/ci_test_groups.py --verify and
its regression test pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…erified-lanes-20260910

# Conflicts:
#	PROGRESS.md
…erified-lanes-20260910

# Conflicts:
#	packages/microcosm-build/src/microcosm/build/uk_runtime/__init__.py
…erified-lanes-20260910

# Conflicts:
#	packages/microcosm-build/tests/test_spec_engine_country_bundles.py
Main's #896 resolves the local git code pin before any I/O in
run_uk_calibration, so four tests that previously refused or read back
before the pin now reach git rev-parse. In the installed-wheels CI job
the package lives under /tmp/wheels-venv, where no repository exists,
and test_readback_failure_updates_build_record_and_failed_attempt
failed with "Could not resolve the local git code pin". Give those four
tests the existing invented_code_pin fixture, as their siblings already
have; the unresolved-pin refusal test keeps its own failing pin.

Verified: the file passes 23/23 in the repository and 23/23 with a git
that exits 128 on PATH, which reproduces the wheels job condition.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Ship the native property and tax graph without the child-property
completion, keep the native CPU cap and remove the repeated-verification
cost instead of buying it, and make a from-scratch build plus
certification of the default the merge gate for this integration and a
standing check on main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Records the lane's base, goal, evidence base and the two local constraints:
root out.md belongs to another lane, and the v5 cold run holds the machine.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
_update_series boxed every value of a plain float, integer or boolean column
into a Python object and framed them one at a time; on pandas 3.0.3 a plain
column is a NumpyExtensionArray with no _data/_mask, so the fast masked path
never applied to it. _object_stream now builds the identical framed bytes --
the same 'object' dtype header, the same shape header, the same per-value
length prefixes and payloads -- with vectorised numpy, and returns None for
every column kind it cannot reproduce exactly, including a Categorical whose
values are an integer-coded view.

No digest value moves. test_graph_executor_series_stream compares the live
helper against a verbatim copy of the pre-change body at the byte level over
every column kind, the float specials, non-canonical NaN payloads, both int64
endpoints and a 200-frame random sweep. Measured 0.47 s -> 0.09 s on a
6,928 x 240 frame.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
docs/us-native-verification-once.md states, per mechanism, the guarantee
today, the change, why the guarantee survives and the new cost class -- and
the two residuals the epoch's unconditional final re-validation exists to
close.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
_records read the 2.4 GB of ACS person CSV inside csv_pus.zip one Python byte
at a time; with peek and columns.lines it was 57-65% of the admission phase in
every measured native run. It now finds terminators with bytes.find, counts
quote parity with bytes.count, and keeps each terminator cursor until the scan
passes it -- so an absent carriage return costs one search per 1 MiB block
rather than one per record, which is where the bulk of the win is.

The two ceilings are charged by _fence. A record no longer than the smaller of
the live ceilings cannot have violated either, because the token counter resets
at every record start; anything longer is replayed byte by byte, so the refusal
code and the byte it fires on stay the fence's own -- including the LF that
closes a CRLF record, which the fence charges against the record ceiling only.

test_us_acs_record_fence_scan compares the shipped fence against a verbatim
copy of the byte loop: every boundary case, every block-boundary case, every
ceiling on both sides, 4,500 randomised strings, 300 row-shaped inputs, and all
9,331 strings up to five bytes over the branching alphabet under each of five
ceiling settings -- 46,655 exhaustive comparisons. Measured on the real staged
archive: 10.8 MB/s -> 522.8 MB/s, digests identical.

Re-pins acs_native_coverage_binding._ACCEPTED for the changed module:
  475aa795c8a5b49a0dd3405a0877866dddd03012a1f2b9743447fab1fe85bcff
  -> 9ec68721a4cf480ef412c51ab354db9000eb7a7989e35e6e09b1574d88e8e49f
  (shasum -a 256 .../acs_person_coverage_authentication.py, after formatting)
The module's graph_implementation_inventory contract is unchanged: the rewrite
adds no _RESOURCE_CALLS name and no non-stdlib import.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
executor.py re-derived source_content_key for every source a cold-executed
node declared, and keys._directory_identity read_bytes()es every file of the
tree: O(executed source-declaring nodes x source bytes), which on the native
graph is the 3.48 GiB source tree re-read per node.

run_graph now carries a _SourceIdentities cache in its own locals, populated by
the run-start pass as it reads. A cached key is reused only while the path's
stat signature -- (dev, ino, size, mtime_ns, ctime_ns) for a file, and for a
directory its own identity plus the relative name, type and identity of every
entry _directory_identity would walk -- is identical to the signature taken
both immediately before and immediately after the read that produced the key.
A derivation whose two signatures disagree is not cached at all.
source_content_key itself stays pure and uncached: three key and codec tests
call it twice on one path with the bytes changed in between.

Before the manifest is built, every source is re-derived in full with the cache
bypassed, written inline rather than through _source_paths_and_keys because a
build test counts calls to that code object. That closes the two cases a stat
signature cannot decide, and gives the executor a guarantee it never had: a
source changed during a node that declares none is now caught.

RunManifest gains source_identities, attached exactly like populations and
mass_ledgers -- repr=False, compare=False, outside content_addressed, outside
to_json, outside every node receipt and cache record.

test_graph_executor_source_identity pins the contract that had no test at all:
a source rewritten by its own node refuses and writes nothing to the store, a
file added to or removed from a directory source refuses, a same-length rewrite
with mtime restored refuses, a change during a source-free node refuses at run
end, and an unchanged source is read twice per run however many nodes declare
it. The graph suite falls from 108 s to 72 s.

Also records the real-archive record-fence parity receipt: both ACS person
members, 2.4 GB, 3,422,890 records, identical digests, 226.9 s -> 4.6 s.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Every borrow of AuthenticatedSurveyPopulationPreparation re-ran the whole
authentication: both source catalogues, the ACS native coverage binding (eight
full zip reads), the nested ASEC native population (seven files), the ten-file
roster re-hash and every pure seal -- once per executed node through the
population observer, and again inside every kernel that borrows it. The ASEC
native capsule did the same on every .frame access.

survey_population_preparation.verification_epoch() makes that happen once.
Inside an epoch each borrow still pays a cheap tier in full -- the live
authority and attached owner payloads, _encode(_producer()), and _file_stats
over the whole roster -- and the expensive tier is skipped only while a
signature is unchanged: the stat identity of every path the four foreign
verifications and _source_files re-read, the identity and length of every
borrowed payload, and a read-free witness of every live Frame's storage
(buffer address, shape, strides, dtype and writeable flag per column, index and
weight). A signature that moved is a memo miss, not a refusal: the complete
validation runs and raises exactly what it would have raised.

Leaving the epoch, at every nesting level, re-validates every memoised capsule
in full with the memo bypassed, so a change no signature can see -- an in-place
write into a live buffer, or a value written into a frozen plan row -- still
makes the run refuse before the caller receives anything. Outside an epoch
nothing is memoised and every existing borrow, refusal and trace-based mutation
test keeps its exact meaning; _validate and _validate_state are untouched.

The three foreign owners gain only a _verified_source_stats() helper naming the
paths their own verification re-reads. Their inventory contracts do not move.
The two capsule modules' unbound_uses digests do, because the new functions use
already-imported names in new scopes; both are re-derived through
graph_implementation._dependency_contract and recorded in PROGRESS.md.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
run_atomic_survey_population and run_atomic_survey_financial now wrap their
whole body in survey_population_preparation.verification_epoch(). Every native
capsule the run borrows is authenticated in full once, reused while its sources
and live storage are provably unchanged, and re-authenticated in full before
the run returns. The financial run's prefix opens a nested epoch; each closes
with its own full re-validation.

Both commits are an indentation plus one 'with' line. Proven mechanically: the
statement lists of both function bodies parse to identical ASTs before and
after (50 statements and 102 statements), the argument lists are unchanged, the
docstrings stay outside the wrapper, and both files were ruff-format clean at
the base so every hunk lies inside the wrapped function. Keeping the body in
its own function matters: a build test asserts
caller.f_code is runner.run_atomic_survey_population.__code__.

The epoch's exit closes the nested capsule's epoch first, so this owner's final
validation reaches the nested population's complete file check rather than its
memo, and collects both refusals so neither can be lost to the other in a
finally block. A foreign owner's error raised by a deferred borrow is
translated exactly as _checked translates it, instead of escaping raw.

test_us_native_verify_once_epoch: 17 tests. Validated once and reused (one
_source_files pass for six borrows, five memo hits); the producer encoding
never memoised; every on-disk mutation -- append, truncate, same-length rewrite
with mtime restored, a file added to a source directory, a touch -- refusing at
the borrow that follows it AND again when the epoch declines to close over it;
a signature miss outside the roster re-running without refusing; a change no
signature can see refusing at the close; a failing body finalizing nothing; and
nested epochs finalizing at every level.

Also adopts the change in docs/validation-cost-and-borrow-boundaries.md, which
recorded this optimization as proposed rather than adopted, and corrects the
codecs docstring that called the executor's check post-run.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Opus <noreply@anthropic.com>
MaxGhenis and others added 30 commits September 20, 2026 09:47
Reduce the two published CPS ASEC disability income slots to one
non-workers-compensation leaf by borrowing the retirement-detail owner's
already-qualified DIS literals and the retired eCPS arithmetic. No raw
member is read, no source authority is issued and nothing is completed.

The archived parameter dictionary and
derive_us_disability_benefits_from_asec are reused rather than restated,
and the workers' compensation exclusion is bound to the printed meaning
of code 1 rather than to the bare integer. A slot resolves only through
a closed vocabulary; NIU, missing, under-15 and contradictory slots stay
unknown instead of becoming an observed zero, and a "yes" receipt with
no populated source stays unknown too. The archived arithmetic's reading
is recorded beside the qualified leaf and never adopted.

Conditional allocation provenance is evaluated against each flag's own
printed universe and never qualifies a receipt; topcode flags describe
only the dollars the leaf admits. The thin attachment copies one
qualified row to every clone of its original source person by identity
join, takes no draw, mutates no receiving cell, and leaves rows on an
arm this source never observes explicitly unknown rather than zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Independent review found four defects in the adapter and four claims the
code did not support.

A workers' compensation slot whose amount literal was missing was being
admitted as a known zero, so the module contradicted its own rule that a
missing cell stays unknown. The exclusion now applies only to the two
statuses a coherent readable slot can carry.

The implementation fence bound the detail owner's seal but not its
qualifier or literal projection, so a swapped qualifier could have handed
over doctored amounts, and it listed constants by hand, so REPORTING_AGE
and others could be retuned after import without refusal. It now collects
every module-level constant and binds the borrowed callables this module
actually depends on, including the routing owner's receipt, literal and
code-frame helpers.

The routing owner's family allocation reading is universe-blind by its
own charter, so a flag read outside its printed universe was standing as
this family's allocation answer beside a per-flag label that said the
opposite. The two scopes are now reported under separate names.

Tests cover each fix, prove the archived function itself is called with
its archived parameters rather than merely reproduced, mutate the
borrowed table for real between the two seals, and check that every
attached cell is the source person's own qualified cell. The module
docstring and the doc drop the overbroad "reads no raw member" and
"produced a zero for every row" claims.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A second independent review confirmed the first round's four fixes and
found three more, all of which predate those fixes.

The constant sweep stored another module's live mapping by reference, so
retuning the archived parameter dictionary in place mutated the snapshot
with it and passed its own equality check. Constants are now detached
before they are sealed.

The implementation digest was taken at call time, so a file edited after
import left the loaded functions passing the fence while the recorded
`implementation_sha256` named bytes that never ran. The digest is now
pinned at import and compared, so such an edit refuses.

The fence bound the detail owner's qualifier but not its member capture
or amount comparison. The owner cross-checks the retained DIS amounts
against the money owner, but `DIS_YN`, `DIS_SC1/2` and the allocation and
topcode literals reach this leaf from the capture alone, so a replaced
capture could have turned a workers' compensation slot into an admitted
one. The capture and comparison path, the routing owner's capture and
digest helpers, and the archived function's own input guard are now bound
too. The doc says plainly that this list is enumerated, not a transitive
closure, and names the capture boundary a consumer should know about.

Topcoding reads as "some admitted slot is topcoded", so one readable flag
of 1 now settles it rather than being erased by a second admitted slot
whose flag cannot be read.

The evidence no longer claims `raw_member_read: False` from the public
qualifier, which does cause a read through its owner; it records
`adds_raw_member_reader` and names the capture owner instead. The
previous commit message's "the borrowed callables this module actually
depends on" overstated an enumerated list, and the first commit's "No raw
member is read" was true only of the pure entry points.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…laims

A third independent review confirmed the second round's fixes and found
one more defect: the receiver supplies its person table through its own
callable, which runs after the attachment's entry check, so a receiver
that retuned a module constant between the two could have had its own
name used for the canonical leaf. The attachment now rechecks the
implementation before returning.

The implementation pin attests this file as read at the end of its own
import; it cannot attest the window between the loader's read and that
one, and the evidence scope and the doc now say so rather than claiming
"the bytes as imported".

Three descriptions were more general than the code: the owner
cross-checks only the two retained DIS amounts against the money owner,
not the receipt, source and flag literals; the archived arithmetic read
whatever a never-asked row's literals held, usually but not always zero;
and a receipt answer alone does not settle a slot, since a "no" or NIU
answer beside a nonzero amount or a populated source code is a
contradiction rather than a nonreceipt.

New tests close the gaps the review named: the published source-code
table is pinned by value instead of read back from the owner, the age-15
boundary and receipt-universe drift are exercised, the second slot
excludes on the same terms as the first, a contradictory out-of-universe
row is shown carrying a positive archived reading that is still not
adopted, a definite allocation is shown outranking an unresolved flag
where the owner's universe-blind reading does not, clone transport is
shown following identity rather than row order, and a hostile receiver
is shown being refused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A fourth independent review found that binding a class object leaves its
initializer swappable: the identity survives, so the implementation fence
passed while the constructor this module calls after its last check could
be replaced by whoever supplied the receiver, whose own table() runs
inside the attachment. The fence now seals each returned type's own
methods, the attachment is constructed before the terminal checks, and
both results are required to have kept the payload they were handed.

The archived replay runs over the four fields as the owner parsed them,
so it is evaluable only where each read inside its own printed domain,
while the retired function accepted any finite number. A source code
outside the published table is therefore a row the old pipeline would
have summed and this replay reports as unevaluable. The docstring, the
doc and a new test say so; the previous "whatever it read is recorded"
wording did not.

Two further descriptions are narrowed to what the code establishes: the
routing owner answers an unresolved flag literal before a nonzero flag
outside its universe, so its universe-blind reading is not unconditionally
a publisher allocation; and "usually zero" claimed a frequency no static
code or invented fixture establishes. The third commit's claim that the
receiver fence was closed was true only of the path its test exercised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`native_survey_handoff.REQUIRED_RELEASE_EVIDENCE` lists
`national_and_cd_target_fit`, and nothing produced it. The canonical
two-artifact yardstick, `tools/score_us_release_head_to_head.py`, emitted
`family`, `entity`, `value_basis` and `period` per fiscal row and no
geographic identifier at all, so a national-versus-congressional-district
reading of a head-to-head could not be computed from the scorecard even
after a real run — while the native measurement kernel that will produce
the candidate is already geography-scoped by construction
(`graph_fiscal_measurement._LEVELS`, `FiscalGeographyScope`).

Add `us_runtime/target_geography_view`: a pure classification of one
compiled target row into the national/state/congressional_district view its
declared evidence names. It prefers explicit declared level evidence, in the
order the compiler produces it (the calibration hierarchy's geography tier,
then `ledger_geography_level`, then `geography_scope`), renames ledger's
`country` to `national` exactly as `fiscal_targets.py:2816-2826` already
does, and reuses the shared
`microcosm.data.us_critical_targets.is_congressional_district_target` rather
than restating CD membership. A row no declared evidence places is reported
`unresolved` and kept out of the three views: labelling it national would
make the national view look better than the artifact is. A disagreement
between the shared CD classifier and an explicit declared level is counted
and reported, never silently resolved.

Wire it through the scorecard: per-row `geography_level`/`geography_id` and
their provenance, a `by_geography_level` rollup and a
`geography_view_resolution` receipt on each artifact, a comparative
`by_geography_level` block on the head-to-head, and Markdown sections for
both. Per-view contributions are shares of the same
`relative_error_loss` aggregate, so they sum to it exactly; the tests pin
that, the counter reconciliation, the refusal routes, and equality with the
frozen kernel's scope vocabulary. Schema version 3 -> 4.

No gate, threshold, tolerance, band or verdict is added. The scorecard stays
evidence for the owner's flip decision; this only decomposes that evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`experiments/replacement_scorecard/incumbent_48b9d479.json` is an
incumbent-only run with `comparison: null`, which is the shape published
before a candidate exists. Pin that path: the per-artifact view rollup and
its unresolved count must render there too, and the comparative section must
not, or the `national_and_cd_target_fit` evidence would only appear once a
candidate is available.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Root reviewed e45793e and did not adopt it:
`target_geography_view.us_target_geography_view` chose a geography level
from one piece of evidence and then searched an ordered metadata key list
(`ledger_geography_id`, `congressional_district_geoid`, `state_fips`) that
consulted neither the level nor the declaration that named it. A row whose
hierarchy sat below the advertised views would have reported that sub-state
id as a state's, and a district row without its own geoid would have
reported the parent `state_fips` the compiler derives from that geoid
(`fiscal_targets.py:2825-2830`).

That search was unreachable for every spec the US compiler can currently
emit: each reference carries a hierarchy seed, `_calibration_hierarchy`
refuses an empty or non-single-valued fact geography
(`ledger_targets.py:934-952`), and `HierarchyNode` refuses an empty id, so
every compiled row resolves through `hierarchy_geography` with a bound id.
No scorecard's counts change. This is a contract repair that closes the
hole before a producer reaches it, and the commit claims nothing more.

An identifier is now read only from a declaration whose own declared level
is the row's resolved level -- hierarchy geography, then the ledger copy,
then the key that level owns -- and a row with none keeps an empty
identifier with `geography_id_source` `none` rather than borrowing one.
Every declaration at the level stays visible in `declared_geography_ids`.

Two encodings exist and are kept apart rather than reconciled. The prefixed
census GEOID is one ledger field read twice (`hierarchy.geography.id` and
`metadata['ledger_geography_id']`), so those two are comparable and
`geography_id_declarations_conflict` asserts they agree -- a defensive
invariant expected to read zero, labelled as such in the artifact note. The
bare `state_fips` / `congressional_district_geoid` is a prefix-stripped
restatement, never compared against a prefixed id and never a conflict; it
is also not interchangeable, because both CD vintage prefixes strip to the
same four digits. `"0400000US"` is an unnamed literal in four production
modules with no shared constant, so no conversion is attempted here.

Because of that, the rollup counts distinct *canonical* ids and reports
`rows_with_bare_geography_id` and `rows_without_geography_id` beside them,
with per-view `geography_id_source_counts`: mixing the two spellings in one
set would count one area twice, which is the same class of wrong count the
level binding exists to prevent. Markdown states the new counts in both the
per-artifact and comparison sections. Schema version 4 -> 5.

No gate, threshold, tolerance, band or verdict is added, and the scoring
path, the loss definition and the per-target error rows are untouched.

TESTS NOT EXECUTED. Free disk measured 8.54 GiB falling to 7.76 GiB during
this work, below the operator's 10 GiB floor, and root notices -180/-184
hold new test runs while it is; root's continuity instruction for this job
is code and static review only. Eight new tests and two fixture updates are
authored and reviewed but unrun; the bounded no-engine harness is staged at
893/native-comparison-geography-id-binding-20260920/run_tests.py with the
exact invocation recorded in the lane report. Static checks that were run:
`ast.parse` on all three edited modules, repo-wide `ruff check` (passes),
`ruff format --check` on the touched files (clean), and
`tools/ci_test_groups.py --verify` (verification=ok).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The inline comment on `_CANONICAL_GEOGRAPHY_ID_SOURCES` called a difference
between the two copies "a real conflict" without saying why. Name the
upstream refusal (`ledger_targets.py:934-952`) that makes it an invariant
violation rather than an encoding difference, matching the property
docstring and the artifact note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`_declared_geography_ids` returns `(source, identifier)` pairs, and the
binding unpacked the first pair as `geography_id, geography_id_source`, so
every resolved row reported its identifier as the source name and its
source as the identifier: a California state row came out with
`geography_id="state_fips"` and `geography_id_source="06"`. The empty
branch was reversed the same way, so an unbound row reported
`geography_id="none"` rather than `""` -- which would also have put the
string "none" into a view's distinct-area set.

Caught by the bounded no-engine suite, which had been held while free disk
sat below the 10 GiB floor and ran once that cleared (21.45 GiB free): nine
tests failed on the swap, including the pre-existing
`test_view_resolution_routes_and_refusals`, whose geography_id assertions
predate this branch's identifier work.

After the fix the whole file is green under the same envelope: 54 tests, 0
failures, 0 errors, 4 pre-existing `requires_us` skips, 24.45 sampled CPU
seconds, 39.71 wall seconds, 460,275,712 bytes peak RSS, and zero guard
denials for country imports, network, child processes, external data and
external writes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Source-qualified SS report completion for selected original ASEC and ACS
people: full-original ASEC DESIGN donors, real categorical fit/probability
artifacts, reason-constrained completion, and a no-fit empty-recipient path.
No canonical beneficiary writes and no release qualification.

The probability and normal-branch report nodes join the donor version: the
graph compiler makes a FILTER depend on every member of its base version, so
a fit consumer on the full-source version cycled through the donor FILTER
(graph-successor3). Successor4 passed 6/6 under the admitted harness; the
full graph and source files pass 32/32.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The keep-all and placement nodes named every entity ID and person membership
column in their Slices. Those columns arrive in the executor's structural
view and have no owner in the compiled declaration, so the first genuine
Stage B run refused at compile: placement_receiving read person.person_id,
which no node owns in the canonical state version. Declare only the
qualified person values and candidate outputs, and omit an entity with
neither. The host compile test now owns no structural column in its fixture
CREATE, matching the real host, and asserts the slices name none.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The genuine Stage B successor3 run reached the enrichment compile and
refused with CURRENT_SURVEY_AMOUNTS_PREFIX_DECLARATION_OR_KEY: 113 parent
nodes kept their declarations and kernel hashes but changed keys, starting
at survey_puf55.receiving. compile_graph makes a structural node depend on
every non-structural member of its base version, and that FILTER's base is
the financial version, so the arm-zero recipient projection, matrices and
fixed-input nodes declared on the financial version joined the authenticated
parent's predecessor closure.

The original arm now opens survey_puf55.original_source_version, a keep-all
FILTER off the financial version, and declares its three source nodes there.
Arm-one declarations are byte-identical. The host reconstructs the version
through the binding, and a new test compiles a parent FILTER with and
without the original arm and asserts its predecessors do not change (and
that the old binding did change them).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The genuine Stage B successor4 run passed the prefix invariant and executed
the enrichment graph to the terminal observation, where the original
placement cut refused PUF55_ORIGINAL_PLACEMENT_ARM_ONE_OUTPUT_OWNER. A replay
diagnostic showed every arm-one output column at the terminal owned by
survey_geography.canonical_state_version: population.patch gives a
structural version ownership of every carried cell, so the receiving owner
can never equal survey_puf55.attach once a FILTER sits between arm one and
the terminal.

candidate_outputs now accepts the arm-one attach node or the receiving
version as the owner and, when the version re-owned the cells, requires the
carried values to equal the arm-one values (ARM_ONE_OUTPUT_CARRIED). Any
other owner still refuses. A new test covers the accepted re-owned shape,
a changed carried value, and a later owner.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
dataclasses.replace shares the receiving frame between the reowned inputs,
so the value-mutation case must run after the later-owner case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…acement artifact

The genuine Stage B successor6 run executed the original-arm enrichment cold,
ran the required replay, and refused PUF55_ORIGINAL_PLACEMENT_PLACEMENT_ARTIFACTS
in the host's independent reconstruction: the persisted placement document
differed from the recomputed one in exactly one field, the third
input_population_stamps entry (the receiving terminal). _stamp is an
in-process mutation seal over replay-variant state and is already verified
before and after the document is built; persisting it made the artifact
non-reproducible under replay. The document now records the three input
population versions instead. A new test permutes the receiving table columns
(changing the seal, not the values) and asserts byte-identical artifacts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
placement_result returns a columns/document pair, not a KernelResult.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Exclusive partitions of the 9af56aa attribution run's sampled stacks
(aggregate report 4e7be8ad...): ACS catalogue plus native coverage issuance
is 936.9 of 1,681.6 outside-graph cold CPU-s, ASEC catalogue plus native
population 351.4, geography/predictor requalification 209.2. Deterministic
source-byte functions (ACS _reconstruct/_inventory/_collect/PUMS tables and
ASEC coverage reconstruction) are listed with inclusive CPU as memo
candidates. Statistical estimates, not exact function CPU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds us_runtime/source_memo.py (opt-in, HMAC-signed index over
content-addressed blobs, keyed by exact input SHA-256s, runtime and owner
code identity) and wires it into the ACS housing full projection and
selection, the coverage member inventories and the catalogue collection.
Tests and receipts follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…keys

Owner code identities now add source_memo.live_code over the modules each
derivation executes (file bytes, loaded-code check, bound functions,
immutable constants). Keys and indexes are strict UTF-8 JSON; the housing
selection is keyed by the exact tuple's digest; blobs are verified by
streaming. Repins the inventory contracts and the native _ACCEPTED digests
and classifies source_memo.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
test_us_source_memo.py covers activation and isolation, content-addressed
hits, key sensitivity to bytes, code and parameters, unrecorded refusals and
inexact values, bypass of unformable identities, every tamper class (blob
digest/size/missing, index MAC/fields/form, garbage, symlinks, foreign key,
misfiled entry, decode failure), live_code binding, and owner-level identity:
the housing source, the ACS catalogue and the whole ACS+ASEC preparation
payload are byte-identical with the memo off, cold and warm, and a warm run
executes none of _archive, _select, _inventory_uncached or _collect.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A virtual environment reaches python through a symlink, which file_sha256
refuses to follow, so the runtime identity could not be formed for a
statically linked interpreter and every memoized call was bypassed. Resolve
that one path before hashing it; the test pins the digest to the real file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The memo codec now round-trips object-dtype columns (scalars, missing
markers, string storage) and preserves floating-point bits; it rejects
unsupported dtypes and metadata instead of coercing them. A memo entry
binds the live parser identity, so a changed parser invalidates and
refuses stale entries. Three reviewed dependency contracts were
refreshed, four more ACS stages use the memo, and the changed owner-file
hashes are updated in the implementation inventory.

New tests cover the codec, frame identity and parser binding; with the
existing memo, ACS PUMS, person-coverage and store-dependency suites,
391 tests pass. The actual-archive benchmark is still pending, so no
speedup is claimed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant