Smokescreen blinding: per-catalogue custody, blind drawn once, parts sealed at birth - #253
cailmdaley wants to merge 196 commits into
Conversation
971767b to
11b8895
Compare
…ll image from lock (#266) * felt: fork-implementation — #243 green, #253 rework under fix; fork PRs reviewed Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KpaRHk3QwN13myduQ3hJyf * felt: fork-implementation — constitution absorbs #241 re-rulings (one-file layout, per-part blinding at birth) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KpaRHk3QwN13myduQ3hJyf * deps: declare cosmo_numba + numba, adopt committed uv.lock, install image from lock Make sp_validation's environment reproducible so an image build can never re-resolve numpy past numba's ceiling — the drift that silently upgraded numpy to 2.5 and broke numba/ngmix. - Declare cosmo-numba (aguinot/cosmo-numba@main; not on PyPI) — the numba B-mode kernels b_modes.py imports. main carries numpy-2 FFT support via its rocket-fft dep and declares its own deps, so numba's numpy window reaches the resolver. Also pin numba directly: the one load-bearing constraint, made visible and resilient to cosmo-numba's dep metadata (which has emptied out between refs). - Declare the other imported-but-undeclared deps the audit found: matplotlib, pandas, pyyaml (core, src/); fitsio (glass extra); a new `workflow` extra for snakemake + mpi4py. cv_runner and unions_wl left undeclared (no resolvable source) with a NOTE. - requires-python and ruff target -> 3.12 (the container's Python; cosmo-numba's floor). - Commit uv.lock (un-ignored) as the SSOT; scope it to Linux via [tool.uv]. numpy resolves to 2.4.6, inside numba 0.66's <2.5 window. - Dockerfile installs from the lock: `uv sync --frozen --inexact` into the base image's /app/.venv (--inexact keeps the ShapePipe stack). Drops the ad-hoc snakemake layer and the cs_util `--upgrade` workaround — the lock pins cs_util 0.2.2 (with cs_util.size), decoupling us from the base image's cs_util. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup * felt: uv-lock-cosmo-numba — @main resolution, install model, cosmo-numba gotcha finding Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup * test: don't load base image's stale pytest-pydocstyle/pycodestyle plugins uv sync installs a newer pytest than the base ShapePipe image's pytest-pydocstyle/pytest-pycodestyle (>=2.4) expect; their pytest_collect_file hooks use the removed `path` arg and crash collection. sp_validation lints with ruff, so disable both via addopts (`-p no:`). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Add sp_validation.sacc_io: the writer/reader layer for the two-file SACC
layout that becomes the package's standard data-product format.
{version}.sacc analysis vector — NZ tracers, coarse xi+/-, pseudo-Cl
(EE/BB/EB) with a shared BandpowerWindow, COSEBIs,
pure E/B, rho/tau PSF diagnostics; one FullCovariance
assembled block-diagonally (zero cross-blocks).
{version}_xi_fine COSEBIs/pure-EB integration input — same NZ tracers,
fine-grid xi+/-, DiagonalCovariance from TreeCorr
varxip/varxim.
Covariance order is point-insertion order (SACC preserves it bitwise
through FITS). Writers insert in the canonical order — xi+ then xi-, Cl
(ee, bb, eb), COSEBIs (all En then all Bn), pure E/B in _EB_KEYS order
(xip_E, xim_E, xip_B, xim_B, xip_amb, xim_amb, matching
b_modes.calculate_eb_statistics), rho, then tau — but readers never assume
global order: every getter resolves indices through s.indices(dtype,
tracers, **tags). assemble_covariance validates that blocks are contiguous,
ascending and tile the data vector exactly, failing loud otherwise.
Custom data types (pure E/B, rho, tau) all parse under
sacc.parse_data_type_name. Tag filters are plain kwargs; the tags={...}
form silently selects nothing and is never used.
Test suite (test_sacc_io.py, all synthetic and fast): per-writer
round-trips (arrays/tags/windows/NZ bitwise), covariance block alignment
and zero cross-blocks, assemble_covariance failure modes, DiagonalCovariance
round-trip, extract() sub-covariance alignment, a tomographic multi-pair
case, reader/writer mirroring on a mixed file, and the end-to-end two-file
layout. 20 passed.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
Fresh-eyes review caught a correctness bug: readers re-sorted selections by
theta/ell/n, but covariance blocks and bandpower windows stay in insertion
order. On a non-ascending grid the reader output silently desynchronised
from its covariance, and get_pseudo_cl returned sorted cl arrays against
unsorted window columns — internally inconsistent within one return tuple.
Fix by construction, not by sort:
- Writers validate their grids. add_xi/add_pure_eb/add_rho/add_tau require
strictly ascending theta; add_pseudo_cl requires strictly ascending
ell_eff (add_cosebis is inherently safe — it enumerates the mode index).
Out-of-order grids raise a loud ValueError naming the argument.
- Readers drop the sort entirely and return in s.indices (insertion) order,
so every getter is covariance- and window-aligned for ANY file, and
ascending for canonical files. The _sorted_* helpers are replaced by
plain insertion-order accessors (_mean/_tag).
Also:
- _pair normalises (i, j) -> sorted, so get_xi(s, (1, 0)) addresses the same
symmetric shear-shear pair as (0, 1) instead of a silent empty read.
- Module docstring documents the tomographic ξ covariance ordering:
insertion is pair-major ([pair0 xip; pair0 xim; pair1 xip; …]), supplied
to assemble_covariance as one contiguous block matching add_xi call order;
type-major converters (DES 2pt-FITS) permute explicitly via s.indices.
- extract() docstring states tracers takes SACC names, not integer bins.
New tests (26 total, was 20): writers reject non-ascending theta/ell;
(1, 0) == (0, 1) round-trip; a 3-pair tomographic ξ covariance assembled as
one contiguous pair-major block with per-pair sub-blocks recovered via
extract(); get_pseudo_cl window column j <-> ell_eff[j] via window_ind tags.
Co-Authored-By: Claude Opus <noreply@anthropic.com>
The two-file split (analysis + {version}_xi_fine.sacc) was premised on a
10000-bin fine grid; the production operating point (Paper II B-modes) is
1000 bins, where a dense per-pair fine covariance block is ~32 MB and the
CosmoCov integration-binning covariance — which feeds pure-E/B and COSEBIs
error propagation — has a natural home as a BlockDiagonal block alongside
the analysis blocks. Layout test replaced with the one-file end-to-end
case (dense fine block, extract() sub-covariance alignment, zero
cross-blocks) plus a varxi-diagonal fallback test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019R5eiy11Lihkgn4MKufXSp
…-forward) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019R5eiy11Lihkgn4MKufXSp
…date_statistic PRD #241 §4 (Mocks vs data): save() requires type='data'|'mock' and stamps it into metadata; load() raises on type='data' files lacking the concealed=True blinding stamp, so skipping the blind can never silently expose real data. Mocks load freely. allow_unblinded=True is the loud escape hatch reserved for the blinding/unblinding tooling. merge() wraps sacc.concatenate_data_sets thinly: per-statistic files combine in order, shared tracers stored once, covariance block-diagonal (library-enforced all-or-none), metadata union with loud conflicts. update_statistic() is the value-only merge-back for the extract -> conceal -> merge blinding flow: matches each sub point by (data_type, tracers, tags) and overwrites the value, leaving order and covariance untouched. Tests: all four data/mock x concealed/unconcealed quadrants, the escape hatch, stamp validation, merge (points, covariance, metadata conflict, mixed-covariance failure) and update_statistic (values-only, unique-match). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
…comments - grid tag values: coarse -> "reporting", fine -> "integration" (descriptive, not relative); tag kept because both grids share data type and tracer pair, so the tag is the sole disambiguator. Swept module docstrings and tests. - Module docstring now spells out why insertion order is load-bearing: covariance row/column i refers to the i-th inserted data point. - new_sacc: comment that psf_stars rides in the same tracer list as the NZ tracers only because a Sacc has one flat tracer namespace (bookkeeping, not physics). - source_name/_pair helpers kept (4 and 8 call sites) with one-line justifications of the conventions they centralize. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
- Factor the theta-tagged insertion loop into _add_theta_series (add_rho, add_tau, add_pure_eb); add_pure_eb zips PURE_TYPES.values() against its signature order, and PURE_KEYS is derived from PURE_TYPES instead of restating it. - Factor the (theta, plus, minus) read pattern into _get_pm (get_xi, get_rho, get_tau). - add_xi hoists the optional-tag None-filtering out of the point loop; extract and merge lose their throwaway mutable dicts; get_cosebis builds its scale-cut tags in one expression. - Tests: shared _add_xi default-ξ builder and _xi_block/_cl_block/ _cosebi_block canonical index-block helpers replace ~60 lines of copy-pasted setup; test_readers_on_mixed_file builds on _multi_statistic_sacc. Behaviour unchanged; 41/41 tests green in the container. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
55cb97d to
4489dd4
Compare
Adversarial review (pass 2) findings: - Sacc.indices returns an EMPTY array (warning only) on an unmatched selection; every reader, extract(), and covariance-block selector now funnels through a shared _indices guard that raises instead — a typo'd tag or non-bitwise-identical float scale cut can no longer propagate empty arrays downstream (assemble_covariance previously died with an opaque IndexError on the same path). - get_cosebis(scale_cut=None) on a file carrying several scale cuts silently concatenated them; now raises and asks for an explicit cut. - update_statistic let two sub points claim the same target point (last-write-wins); now raises. - merge: the library keeps the FIRST input's tracer on a name clash with no equality check; shared tracers are now verified identical (z, nz) across inputs before concatenation. 7 regression tests; 48 total. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…abulary
Rebuild the per-part-at-birth blinding stack on top of feat/sacc-2-sacc-io.
sacc_io.py is now PR-2's canonical module verbatim + a single appended
gather(): the terminal assembly delegates its tracer/point/covariance
assembly to PR-2's merge() and adds only the blind-custody call
(assert_consistent_blind) and shared-stamp write.
The fail-closed load gate lives in sacc_io.load() per the PRD ("sacc_io fails
closed on load"); the blinding tooling uses allow_unblinded=True explicitly
where it must read a not-yet-blinded real vector (blind_part). save() now
carries the required type= (inherited from each part's provenance). Grid
vocabulary swept coarse->reporting, fine->integration across blinding.py,
blinding_theory.py, b_modes.py, blind_data_vector.py, and the blinding tests.
_extract_block/_set_values stay index-based (a code comment says why):
Smokescreen's ConcealDataVector aligns theory_fn output to the sub-SACC mean
by row position, so the block must be carved/written by contiguous integer
index, not by PR-2's tag-matching extract()/update_statistic().
Escrow/custody/CAMB<->CCL machinery is preserved exactly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…tion The branch previously carried a stale vendored copy of sacc_io.py / test_sacc_io.py from before PR2's review rounds. Rebuilt directly on feat/sacc-2-sacc-io (8bd3817) so the canonical module is inherited, bringing over the cosmo_val + Snakemake born-as-SACC migration work from the old tip unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…nfig Three integration-drift fixes surfaced by the reconciled base: - test_ac6_ac8 / test_ac8_dotted: the end-to-end fixtures now pass type="mock" to sio.save (PR-2 requires the provenance stamp); the parts are mocks, blind_part re-saves inheriting that type. - test_ac13: the CosmoSIS halofit config moved to cosmo_inference/cosmosis_config/templates/ on develop; point the AC13 assertion at the new path (token unchanged: mead2020_feedback). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…rename
Sweep the migration code onto PR2's canonical vocabulary and contracts:
- grid='coarse'/'fine' -> 'reporting'/'integration' everywhere (writer
calls, readers, tests), including the internal DAG intermediates:
{version}_xi_coarse* -> _xi_reporting*, _xi_fine -> _xi_integration,
the xi_coarse/xi_fine Snakemake output keys, CANONICAL part names,
cv_xi_reporting_sacc, write_xi_integration_sacc.
- save(s, path) -> save(s, path, type=...): 'data' at every production
writer (the cosmo_val pipeline measures the real UNIONS catalogues;
GLASS mocks do not flow through these writers), 'mock' for synthetic
test fixtures. assemble_sacc propagates its parts' type stamp rather
than hardcoding, so mock parts assemble into a mock analysis file.
- load(path) -> fail-closed load: pipeline-internal readbacks of
freshly written pre-blind data parts pass allow_unblinded=True
(blinding is a downstream Smokescreen step); mock fixtures load freely.
Readers raising on unmatched selections needed no call-site changes:
every get_* reads a statistic guaranteed present in the file just
written or assembled.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
4b058f7 to
c09b063
Compare
Consistency follows tagging semantics: the grid tag declares which binning a set of points lives on, so all same-length theta arrays under one tag value must be bitwise identical (sacc never validates angles, and grids diverging at floating-point level choke CosmoSIS downstream). Different lengths within a tag group pass (scale-cut subsets); grids under different tag values are unconstrained (reporting vs integration differ by design). The rho/tau/pure-EB writers now tag their points grid="reporting" by default (overridable via grid=), so they join xi's consistency group and no untagged group appears in our own files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MN9VazXKHUHQg16kiG7Ufk
…seudo-Cl grid="reporting" The theta consistency guard generalizes to both angular domains: theta and ell points are grouped separately by grid tag value, and within a tag value all same-length grids must be bitwise identical. Two ell series sharing a grid must also carry equal bandpower windows (window ells and weight matrix); series without windows skip that check. add_pseudo_cl now stamps grid="reporting" on every point by default (overridable), joining the merge guard's consistency groups; since sacc's add_ell_cl accepts no extra tags, its per-point insertion (ell + shared window + window_ind) is inlined. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MN9VazXKHUHQg16kiG7Ufk
assemble_covariance now passes sacc.add_covariance a list of per-block matrices instead of a dense zero-filled N×N array. sacc.BaseCovariance.make turns a list into a BlockDiagonalCovariance (one FITS table per block, Σ block² on disk), so cross-blocks are zero and implicit rather than materialized. Validation (contiguous/ascending indices, no gap/overlap, square blocks matching their index span) is unchanged. Every existing consumer already read through the polymorphic .dense property, so only the type assertions needed updating (FullCovariance -> BlockDiagonalCovariance). Added a test that merging two files that already carry a BlockDiagonalCovariance (assembled via assemble_covariance) stays block-diagonal through merge and a save/load round-trip. Updated the module docstring's storage-cost discussion accordingly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pure_eb.py's calculate_pure_eb/plot_pure_eb defaulted nbins_int=100; cosebis.py's calculate_cosebis already defaulted to 1000, and every production config (papers/bmodes, papers/cosmo_val fiducial/pure_eb blocks) already overrides to 1000. This aligns the function-signature default with what every caller actually uses; config plumbing is untouched, so any explicit override still wins. The papers/cosmo_val/config/config.yaml cosebis.nbins_int (currently 2000, production numerics) is deliberately left unchanged — see report. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Writers no longer force components an analysis may not have computed: add_pseudo_cl's BB/EB, add_cosebis's Bn, and add_pure_eb's B/ambiguous blocks now default to None and are simply not written when omitted (add_pure_eb still requires each +/- pair together). Structural arguments (grids, tracer/bin identifiers, the value for a component you ARE adding) remain required with no default. Composite readers (get_pseudo_cl, get_cosebis, get_pure_eb) return None / omit the key for an absent optional component instead of raising, via a new _mean_optional helper; a selection naming a missing component explicitly (s.indices, _mean, extract) still fails loud, per the existing empty-selection guard. merge() and assemble_covariance work unchanged on files with only a subset of components. Documents the optionality contract in the module docstring, and adds partial-file round-trip, merge, and explicit-selection-fails-loud tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…fault cosebis.nbins_int was 2000; every other integration-grid entry in this config (pure_eb, the fiducial block) is already 1000, and cosebis.py's own function defaults are 1000. Unify. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…acc-4-cosmo-val-sacc
Fold the integration-grid ξ± into the single terminal {version}.sacc, and
make part loading fail closed on unblinded real data.
Integration ξ± (grid='integration') is no longer its own terminal product.
The xi_highres part is now gathered by rule assemble_sacc into {version}.sacc
as tagged points, next to the reporting ξ± block. It is fiducial-only (the
10k-bin MPI run emits only the fiducial part), so it joins the fiducial
version's terminal file alone. assemble_sacc.py adds xi_integration to
CANONICAL; its own DiagonalCovariance passes straight through.
assemble_sacc.py no longer loads every part with allow_unblinded=True. The
run type (data|mock, from config, default data) gates it: mock runs load
freely, data runs fail closed unless a part carries the concealed=True stamp.
This is the seam for PR #253's blind-at-birth — a concealed data part then
assembles with allow_unblinded=False untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EaL7prmKUHwJQcDyW3LoxD
…custody The scaffolding removal, the private COSMO_INFERENCE paragraph and the non-custody docstrings are chore/cosmo-val-cleanup's; run_rho_tau passes its rule's custody token and nothing else changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y you pass
blinding.STANDARD declares each estimator cosmo_val computes and its rule
once: ξ± and Cℓ_EE shifted by the default PyCCL theory, Cℓ_BB/EB unshifted,
COSEBIs and pure-E/B derived from ξ±, ρ/τ signal-free. Any other data type
is shifted when its birth passes a function,
sacc_io.save(s, path, custody=c, theory={data_type: f}), and is refused
under a blind without one; its sealed rows pass later derivations as copies.
γt leaves until cosmo_val calculates it: GAMMA_T/GAMMA_X, add_lens,
add_gamma_t, the lens bias, source_lens and the null-lens exemption go, with
their tests. The secrecy test goes; Blind still holds only a name and path.
The README and CLAUDE.md state the calculate-then-save rule.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GJJyGbKSuEjjX9m2zdv6yi
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
theory.shear predicts ξ± and Cℓ_EE for every row of a SACC (zeros for Cℓ_BB/EB), theory.none gives zeros. blinding.conceal hands the theory to smokescreen's concealing_factor and adds the vector; a failure names only its exception type. sacc_io.save conceals a birth under a Blind and stamps its name, or stamps a derivation with its inputs' one blind. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…more An entry must declare blind: none or blind: <name>; entries reading one shear file declare one blind. The catalogue-path helpers stay here, stdlib-only for the host. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CosmologyValidation.blind(version) opens the blind its catalogue config declares; producer rules carry it as params.blind, so flipping it reruns what it touches, and assembly refuses parts stamped otherwise. A launch prints each catalogue's blind. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…riance calculate_pure_eb requires cov_path_int; the interactive jackknife path is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing in the repository calls them; ξ± and ⟨M_ap²⟩ live in cosmo_val. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
conceal moves each row by t(hidden) − t(fiducial), a custom theory blinds a custom type, theory.none stamps without shifting, derivations inherit one stamp, catalogues declare their blind, B-modes respond to the shift only through each transform; DAG tests for the launch line and a blind flip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
…art can't be shifted twice Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
…hetic helper Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
…gument Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
|
@LisaGoh and @sachaguer this is ready for feedback now. some open questions to settle and some oversights i'm sure, but i worked hard to get this as simple as possible. i think it's pretty nice. make sure to click the dropdown in the PR description to appreciate the full shifted data vector plots! |
Pure E/B takes #373's design: one fixed linear operator on the integration-grid ξ± part with covariance K C Kᵀ; the MC covariance and the chunked precompute go. The blinding stays: cv_pure_eb and calculate_pure_eb read the sealed integration part through sacc_io.xi_correlation (which carries the pair weights, so get_xi_weight is not added), and the pure-E/B part is saved derived_from that part. papers/bmodes/pure_eb_modes.py reads the SACC part; the B-mode blinding test uses the operator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r go blinding.py now holds the theory (shear, fiducial, cosmology, and no_signal for statistics without cosmological signal) and blind_of, which reads an already-resolved catalogue entry's `blind:` and refuses a missing one or one shear file read under two blinds. Its scientific imports live inside the functions, so workflow/common.py loads it by path on the host. custody.py and theory.py are deleted. Catalogue and seed resolution are develop's again (_materialize_seed_path, catalogue_entry, get_shear_catalog); CosmologyValidation.blind and common.blind_of call blinding.blind_of on the resolved entry. The per-launch [blind] banner is gone; configure() still refuses a run catalogue that declares no blind before any job is built. Patch centres are written as on develop (PR #367 owns them). Tests keep what guards blinding: ξ±/Cℓ_EE shift and BB/EB do not; pure-E/B under a blind moves E and not B (develop's operator); seal refuses a double stamp; one file under two blinds is refused; a flip reruns the part; an undeclared catalogue stops the launch. _synthetic.py, test_blinding_bmodes.py, test_custody.py and test_toy_run.py go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rsions resolve on the host CosmologyValidation resolved the selected entry's paths before the one-file-one-blind check, so a twin entry under another blind passed. The workflow's catalogue_name strips _seed<N> as well as _leak_corr. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0186WiB28hK7DK2pYaTBLL6j
|
@AstroMike also happy to hear your opinion about this and whether it's flexible enough for your delta sigma example. i think the only work required to blind a new data vector is to a) get your data vector into sacc format and b) pass a function that calculates a theory curve when you save the data vector to disk. the data vector itself needn't be calculated by sp_validation as long as it's in SACC format. |
|
Thank you for this Cail! I think I'd need to try running the whole thing to fully understand the logic of the pipeline, but for now I think it's good to go! With this we should also start thinking about blinding protocol (linked to the open questions you posted): when it happens, by whom, and on which catalogues. |
This PR blinds UNIONS cosmic-shear data vectors with the UNIONS-WL Smokescreen fork. On a blinded catalogue, ξ± and pseudo-Cℓ_EE are shifted by t(hidden) − t(fiducial) before they are first written. That is the theory at a secret cosmology minus the theory at the fiducial, so the true values never reach disk. Everything computed from them inherits the shift: COSEBIs, pure-E/B, ⟨M_ap²⟩ and the assembled per-version SACC. The aim is to prevent accidents in an honest collaboration, not to resist an adversary.
What blinding does to v1.4.6.3
The figure shows SP_v1.4.6.3_leak_corr under a throwaway blind, now revealed (ΔS8 = −0.060, ΔΩm = −0.043). The shift is a smooth, coherent change of amplitude:
Cℓ_BB, B_n and the pure-mode ξ^B are unchanged. The one exception is a pure-mode B bin at 1.5′ (see Testing).
Data vectors: ξ±, Cℓ, E and B modes
The pure-E residuals look jagged because the pure-mode covariance diagonal alternates from bin to bin. The shift itself is smooth.
Design
<paths.blinds>/<name>.blind.json. It holds a seed, the envelope the hidden point is drawn in (S8 ± 0.075, Ωm ± 0.1) and the fiducial point. It is drawn once (python -m sp_validation.blinding init <name>) and never committed:*.blind.jsonis gitignored.cosmo_val/cat_config.yaml, asblind: noneorblind: <name>. An entry without one doesn't run, so nobody measures a new catalogue before someone has decided._leak_corrand_seed<N>versions take their entry's blind, and entries reading the same shear file must declare the same one.theory(params, s)that returns the prediction for every row of SACCs, as an array the length ofs.mean, zero where cosmology has no effect. Smokescreen draws the hidden point from the seed, calls the theory at both points, and the difference is added to the data.Calculate, then save. A measurement function computes its data vector and saves it in the same function, with
sacc_io.save(s, path, blind=cv.blind(version)). It returns the saved part, which is shifted if the catalogue is blinded, so raw signal never leaves the function.calculate_2pcffollows this pattern.Statistics computed from parts are saved with
derived_from=partsand carry the parts' stamp (s.metadata["blind"]). Assembly refuses a part whose stamp isn't the catalogue's declared blind.Theories.
blinding.shear, computes ξ± and Cℓ_EE with pyccl throughcs_util.cosmo.get_cosmo, on the same CAMB HMCode2020-feedback route as inference. It returns zero for Cℓ_BB and Cℓ_EB.blinding.no_signalreturns zeros, for ρ/τ.paramsis a plain dict:S8,Omega_m,Omega_b,h,n_s,m_nu,w0,wa,logT_AGNandA_IA.sis the SACC being saved. Its rows carry a data type, a tracer pair, a θ or ℓ tag and, for Cℓ, a window; its tracers carry the n(z).Unblinding means flipping the declaration to
blind: nonein a reviewed PR and rerunning. Producer rules carry the blind's name inparams.blind, so Snakemake reruns exactly what the flip touches.Other changes for users
sacc_io.savetakesblind=for a new measurement (shifted if blinded) orderived_from=for a derived statistic.sacc_io.loadreads any SACC.calculate_2pcfreturns the saved ξ± part, whichsacc_io.xi_correlationreads like aGGCorrelation. ξ± is no longer written to text files.calculateMapSqto 1e-12.scripts/xip_xim.pyandscripts/map2.pyare deleted. They computed ξ± and ⟨M_ap²⟩ outside cosmo_val, so on a blinded catalogue they would write true values. Nothing in the repo calls them, and cosmo_val computes both.Open questions
/n17data/UNIONS/WL/blinds, isn't created yet) or in a private repo such asUNIONS-WL/blinds?blind: none. If v1.6.x has been seen, blinding it protects little, and the first real blind belongs on the tomographic catalogue.A blind file doesn't record how Smokescreen draws the hidden point. If the fork changed its draw, an existing blind would give a different shift. The lock pins the fork.
Testing
Closes #252. Part of #280: the default theory's settings were matched to
cosmosis_pipeline_A_*.iniby reading, and the numerical comparison with CosmoSIS isn't done yet.— Claude (Opus) on behalf of Cail.
🤖 Generated with Claude Code