Skip to content

Smokescreen blinding: per-catalogue custody, blind drawn once, parts sealed at birth - #253

Open
cailmdaley wants to merge 196 commits into
developfrom
feat/sacc-6-blinding
Open

cailmdaley wants to merge 196 commits into
developfrom
feat/sacc-6-blinding

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Jul 10, 2026 •

Copy link
Copy Markdown
Collaborator

This PR blinds UNIONS cosmic-shear data vectors with the UNIONS-WL Smokescreen fork. On a blinded catalogue, ξ± and pseudo-Cℓ_EE are shifted by t(hidden) − t(fiducial) before they are first written. That is the theory at a secret cosmology minus the theory at the fiducial, so the true values never reach disk. Everything computed from them inherits the shift: COSEBIs, pure-E/B, ⟨M_ap²⟩ and the assembled per-version SACC. The aim is to prevent accidents in an honest collaboration, not to resist an adversary.

What blinding does to v1.4.6.3

blinded minus true, in σ

The figure shows SP_v1.4.6.3_leak_corr under a throwaway blind, now revealed (ΔS8 = −0.060, ΔΩm = −0.043). The shift is a smooth, coherent change of amplitude:

  • up to 2.7σ per bin in ξ±;
  • up to 1.9σ in Cℓ_EE;
  • in COSEBIs, only E₁–E₃.

Cℓ_BB, B_n and the pure-mode ξ^B are unchanged. The one exception is a pure-mode B bin at 1.5′ (see Testing).

Data vectors: ξ±, Cℓ, E and B modes

xi
cl
bmodes

The pure-E residuals look jagged because the pure-mode covariance diagonal alternates from bin to bin. The shift itself is smooth.

Design

  • A blind is a named secret file, <paths.blinds>/<name>.blind.json. It holds a seed, the envelope the hidden point is drawn in (S8 ± 0.075, Ωm ± 0.1) and the fiducial point. It is drawn once (python -m sp_validation.blinding init <name>) and never committed: *.blind.json is gitignored.
  • Each catalogue declares its blind in cosmo_val/cat_config.yaml, as blind: none or blind: <name>. An entry without one doesn't run, so nobody measures a new catalogue before someone has decided. _leak_corr and _seed<N> versions take their entry's blind, and entries reading the same shear file must declare the same one.
  • A theory is a function theory(params, s) that returns the prediction for every row of SACC s, as an array the length of s.mean, zero where cosmology has no effect. Smokescreen draws the hidden point from the seed, calls the theory at both points, and the difference is added to the data.

Calculate, then save. A measurement function computes its data vector and saves it in the same function, with sacc_io.save(s, path, blind=cv.blind(version)). It returns the saved part, which is shifted if the catalogue is blinded, so raw signal never leaves the function. calculate_2pcf follows this pattern.

Statistics computed from parts are saved with derived_from=parts and carry the parts' stamp (s.metadata["blind"]). Assembly refuses a part whose stamp isn't the catalogue's declared blind.

Theories.

  • The default, blinding.shear, computes ξ± and Cℓ_EE with pyccl through cs_util.cosmo.get_cosmo, on the same CAMB HMCode2020-feedback route as inference. It returns zero for Cℓ_BB and Cℓ_EB.
  • blinding.no_signal returns zeros, for ρ/τ.
  • Any other statistic passes its own:
def my_theory(params, s):
    out = np.zeros(len(s.mean))
    ee = s.indices(sacc_io.CL_EE)
    window = s.get_bandpower_windows(ee)
    out[ee] = window.weight.T @ my_cl_ee(params, window.values)  # CCL, an emulator, ...
    return out                                                    # BB, EB: zero

sacc_io.save(s, path, blind=cv.blind(version), theory=my_theory)
  • params is a plain dict: S8, Omega_m, Omega_b, h, n_s, m_nu, w0, wa, logT_AGN and A_IA.
  • s is the SACC being saved. Its rows carry a data type, a tracer pair, a θ or ℓ tag and, for Cℓ, a window; its tracers carry the n(z).
  • The default theory raises on a data type it doesn't model, so such a type is never saved unshifted.

Unblinding means flipping the declaration to blind: none in a reviewed PR and rerunning. Producer rules carry the blind's name in params.blind, so Snakemake reruns exactly what the flip touches.

Other changes for users

  • sacc_io.save takes blind= for a new measurement (shifted if blinded) or derived_from= for a derived statistic. sacc_io.load reads any SACC.
  • calculate_2pcf returns the saved ξ± part, which sacc_io.xi_correlation reads like a GGCorrelation. ξ± is no longer written to text files.
  • ⟨M_ap²⟩ is computed from the ξ± part, matching TreeCorr's calculateMapSq to 1e-12.
  • Pure E/B reads the ξ± integration part and carries its stamp; the modes and covariance are develop's operator (Pure E/B: fixed-quadrature operator with exact covariance #373) applied to it.
  • scripts/xip_xim.py and scripts/map2.py are deleted. They computed ξ± and ⟨M_ap²⟩ outside cosmo_val, so on a blinded catalogue they would write true values. Nothing in the repo calls them, and cosmo_val computes both.

Open questions

  1. Where do blinds live? On a candide path (the default, /n17data/UNIONS/WL/blinds, isn't created yet) or in a private repo such as UNIONS-WL/blinds?
  2. Who may draw and hold a blind?
  3. Has v1.6.x been seen unblinded? For now every entry is blind: none. If v1.6.x has been seen, blinding it protects little, and the first real blind belongs on the tomographic catalogue.
  4. Should cosmo_val refuse to overlay unblinded v1.4.6.3 on a blinded tomographic catalogue? That overlay shows the shift, which is a few σ per point. Otherwise we don't police plots.
  5. Envelope and fiducial: keep S8 ± 0.075 and Ωm ± 0.1 around Planck18, or centre on the inference prior?
  6. Nuisance parameters: Smokescreen shifts only S8 and Ωm. Is that enough once γt and galaxy bias enter? We expect the effect to be second order.
  7. Reveal criteria (PRD: SACC data format + Smokescreen blinding #241 §7).

A blind file doesn't record how Smokescreen draws the hidden point. If the fork changed its draw, an existing blind would give a different shift. The lock pins the fork.

Testing

  • Container suite including the slow tests: 317 passed, 1 xfailed (GLASS maps, unrelated). Host workflow tests: 15 passed, 1 skipped (the SLURM submission smoke test).
  • Synthetic SACCs:
    • ξ± and Cℓ_EE move by the theory difference; BB, EB and ρ/τ don't move.
    • A custom theory blinds a new data type, and the default theory refuses an unknown one.
    • Assembly refuses mixed stamps, and a saved part can't be shifted twice.
    • Flipping a blind reruns exactly the ξ± parts, and in the apptainer toy run the ξ± parts carry the blind's stamp.
  • B-modes on synthetic data vectors with SP_v1.4.6.3 shape noise, at each corner of the envelope: COSEBIs B_n move by at most 0.034σ.
  • End to end on SP_v1.4.6.3 (the figures above):
    • On every ξ± and Cℓ_EE part, blinded − true equals the shift to 1e-12. BB and EB are unchanged.
    • No true ξ± value, seed or hidden parameter appears in any byte of the blinded tree or its logs.
    • Rerun on this head under the same blind, the blinded tree reproduces the earlier one to 1e-12 in every ξ±, Cℓ and COSEBI part.
  • The pure-mode B bin at 1.5′ moves 1.15σ. cosmo_numba's pure-E/B integral doesn't converge on noisy ξ±, so the transform isn't additive in that bin, blinded or not. The bin is already an outlier in the unblinded data, and the fix belongs to the B-mode estimator work.

Closes #252. Part of #280: the default theory's settings were matched to cosmosis_pipeline_A_*.ini by reading, and the numerical comparison with CosmoSIS isn't done yet.

— Claude (Opus) on behalf of Cail.

🤖 Generated with Claude Code

@cailmdaley cailmdaley changed the title Smokescreen blinding: seeded data-vector blind, hash-commitment custody, blinded-COSEBIs derivation (PR 6) Smokescreen-fork blinding: seeded master-SACC blind, hash-commitment custody, CAMB↔CCL cross-check (PR 6) Jul 10, 2026
@cailmdaley
cailmdaley force-pushed the feat/sacc-6-blinding branch from 971767b to 11b8895 Compare July 10, 2026 22:52
@cailmdaley
cailmdaley changed the base branch from feat/sacc-5-firecrown-likelihood to feat/sacc-2-sacc-io July 10, 2026 22:52
@cailmdaley cailmdaley changed the title Smokescreen-fork blinding: seeded master-SACC blind, hash-commitment custody, CAMB↔CCL cross-check (PR 6) Smokescreen-fork blinding: per-part-at-birth blind, hash-commitment custody, CAMB↔CCL cross-check (PR 6) Jul 11, 2026
cailmdaley added a commit that referenced this pull request Jul 11, 2026
…ll image from lock (#266)

* felt: fork-implementation — #243 green, #253 rework under fix; fork PRs reviewed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KpaRHk3QwN13myduQ3hJyf

* felt: fork-implementation — constitution absorbs #241 re-rulings (one-file layout, per-part blinding at birth)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KpaRHk3QwN13myduQ3hJyf

* deps: declare cosmo_numba + numba, adopt committed uv.lock, install image from lock

Make sp_validation's environment reproducible so an image build can never
re-resolve numpy past numba's ceiling — the drift that silently upgraded numpy
to 2.5 and broke numba/ngmix.

- Declare cosmo-numba (aguinot/cosmo-numba@main; not on PyPI) — the numba
  B-mode kernels b_modes.py imports. main carries numpy-2 FFT support via its
  rocket-fft dep and declares its own deps, so numba's numpy window reaches the
  resolver. Also pin numba directly: the one load-bearing constraint, made
  visible and resilient to cosmo-numba's dep metadata (which has emptied out
  between refs).
- Declare the other imported-but-undeclared deps the audit found: matplotlib,
  pandas, pyyaml (core, src/); fitsio (glass extra); a new `workflow` extra for
  snakemake + mpi4py. cv_runner and unions_wl left undeclared (no resolvable
  source) with a NOTE.
- requires-python and ruff target -> 3.12 (the container's Python; cosmo-numba's
  floor).
- Commit uv.lock (un-ignored) as the SSOT; scope it to Linux via [tool.uv].
  numpy resolves to 2.4.6, inside numba 0.66's <2.5 window.
- Dockerfile installs from the lock: `uv sync --frozen --inexact` into the base
  image's /app/.venv (--inexact keeps the ShapePipe stack). Drops the ad-hoc
  snakemake layer and the cs_util `--upgrade` workaround — the lock pins
  cs_util 0.2.2 (with cs_util.size), decoupling us from the base image's cs_util.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup

* felt: uv-lock-cosmo-numba — @main resolution, install model, cosmo-numba gotcha finding

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup

* test: don't load base image's stale pytest-pydocstyle/pycodestyle plugins

uv sync installs a newer pytest than the base ShapePipe image's
pytest-pydocstyle/pytest-pycodestyle (>=2.4) expect; their pytest_collect_file
hooks use the removed `path` arg and crash collection. sp_validation lints with
ruff, so disable both via addopts (`-p no:`).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtw1qcgTrQ6YuMvxwYzNup

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
cailmdaley and others added 7 commits July 16, 2026 02:17
Add sp_validation.sacc_io: the writer/reader layer for the two-file SACC
layout that becomes the package's standard data-product format.

  {version}.sacc       analysis vector — NZ tracers, coarse xi+/-, pseudo-Cl
                       (EE/BB/EB) with a shared BandpowerWindow, COSEBIs,
                       pure E/B, rho/tau PSF diagnostics; one FullCovariance
                       assembled block-diagonally (zero cross-blocks).
  {version}_xi_fine    COSEBIs/pure-EB integration input — same NZ tracers,
                       fine-grid xi+/-, DiagonalCovariance from TreeCorr
                       varxip/varxim.

Covariance order is point-insertion order (SACC preserves it bitwise
through FITS). Writers insert in the canonical order — xi+ then xi-, Cl
(ee, bb, eb), COSEBIs (all En then all Bn), pure E/B in _EB_KEYS order
(xip_E, xim_E, xip_B, xim_B, xip_amb, xim_amb, matching
b_modes.calculate_eb_statistics), rho, then tau — but readers never assume
global order: every getter resolves indices through s.indices(dtype,
tracers, **tags). assemble_covariance validates that blocks are contiguous,
ascending and tile the data vector exactly, failing loud otherwise.

Custom data types (pure E/B, rho, tau) all parse under
sacc.parse_data_type_name. Tag filters are plain kwargs; the tags={...}
form silently selects nothing and is never used.

Test suite (test_sacc_io.py, all synthetic and fast): per-writer
round-trips (arrays/tags/windows/NZ bitwise), covariance block alignment
and zero cross-blocks, assemble_covariance failure modes, DiagonalCovariance
round-trip, extract() sub-covariance alignment, a tomographic multi-pair
case, reader/writer mirroring on a mixed file, and the end-to-end two-file
layout. 20 passed.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Fresh-eyes review caught a correctness bug: readers re-sorted selections by
theta/ell/n, but covariance blocks and bandpower windows stay in insertion
order. On a non-ascending grid the reader output silently desynchronised
from its covariance, and get_pseudo_cl returned sorted cl arrays against
unsorted window columns — internally inconsistent within one return tuple.

Fix by construction, not by sort:
  - Writers validate their grids. add_xi/add_pure_eb/add_rho/add_tau require
    strictly ascending theta; add_pseudo_cl requires strictly ascending
    ell_eff (add_cosebis is inherently safe — it enumerates the mode index).
    Out-of-order grids raise a loud ValueError naming the argument.
  - Readers drop the sort entirely and return in s.indices (insertion) order,
    so every getter is covariance- and window-aligned for ANY file, and
    ascending for canonical files. The _sorted_* helpers are replaced by
    plain insertion-order accessors (_mean/_tag).

Also:
  - _pair normalises (i, j) -> sorted, so get_xi(s, (1, 0)) addresses the same
    symmetric shear-shear pair as (0, 1) instead of a silent empty read.
  - Module docstring documents the tomographic ξ covariance ordering:
    insertion is pair-major ([pair0 xip; pair0 xim; pair1 xip; …]), supplied
    to assemble_covariance as one contiguous block matching add_xi call order;
    type-major converters (DES 2pt-FITS) permute explicitly via s.indices.
  - extract() docstring states tracers takes SACC names, not integer bins.

New tests (26 total, was 20): writers reject non-ascending theta/ell;
(1, 0) == (0, 1) round-trip; a 3-pair tomographic ξ covariance assembled as
one contiguous pair-major block with per-pair sub-blocks recovered via
extract(); get_pseudo_cl window column j <-> ell_eff[j] via window_ind tags.

Co-Authored-By: Claude Opus <noreply@anthropic.com>
The two-file split (analysis + {version}_xi_fine.sacc) was premised on a
10000-bin fine grid; the production operating point (Paper II B-modes) is
1000 bins, where a dense per-pair fine covariance block is ~32 MB and the
CosmoCov integration-binning covariance — which feeds pure-E/B and COSEBIs
error propagation — has a natural home as a BlockDiagonal block alongside
the analysis blocks. Layout test replaced with the one-file end-to-end
case (dense fine block, extract() sub-covariance alignment, zero
cross-blocks) plus a varxi-diagonal fallback test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019R5eiy11Lihkgn4MKufXSp
…date_statistic

PRD #241 §4 (Mocks vs data): save() requires type='data'|'mock' and stamps
it into metadata; load() raises on type='data' files lacking the
concealed=True blinding stamp, so skipping the blind can never silently
expose real data. Mocks load freely. allow_unblinded=True is the loud
escape hatch reserved for the blinding/unblinding tooling.

merge() wraps sacc.concatenate_data_sets thinly: per-statistic files
combine in order, shared tracers stored once, covariance block-diagonal
(library-enforced all-or-none), metadata union with loud conflicts.
update_statistic() is the value-only merge-back for the
extract -> conceal -> merge blinding flow: matches each sub point by
(data_type, tracers, tags) and overwrites the value, leaving order and
covariance untouched.

Tests: all four data/mock x concealed/unconcealed quadrants, the escape
hatch, stamp validation, merge (points, covariance, metadata conflict,
mixed-covariance failure) and update_statistic (values-only, unique-match).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
…comments

- grid tag values: coarse -> "reporting", fine -> "integration" (descriptive,
  not relative); tag kept because both grids share data type and tracer pair,
  so the tag is the sole disambiguator. Swept module docstrings and tests.
- Module docstring now spells out why insertion order is load-bearing:
  covariance row/column i refers to the i-th inserted data point.
- new_sacc: comment that psf_stars rides in the same tracer list as the NZ
  tracers only because a Sacc has one flat tracer namespace (bookkeeping,
  not physics).
- source_name/_pair helpers kept (4 and 8 call sites) with one-line
  justifications of the conventions they centralize.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
- Factor the theta-tagged insertion loop into _add_theta_series (add_rho,
  add_tau, add_pure_eb); add_pure_eb zips PURE_TYPES.values() against its
  signature order, and PURE_KEYS is derived from PURE_TYPES instead of
  restating it.
- Factor the (theta, plus, minus) read pattern into _get_pm (get_xi,
  get_rho, get_tau).
- add_xi hoists the optional-tag None-filtering out of the point loop;
  extract and merge lose their throwaway mutable dicts; get_cosebis
  builds its scale-cut tags in one expression.
- Tests: shared _add_xi default-ξ builder and _xi_block/_cl_block/
  _cosebi_block canonical index-block helpers replace ~60 lines of
  copy-pasted setup; test_readers_on_mixed_file builds on
  _multi_statistic_sacc.

Behaviour unchanged; 41/41 tests green in the container.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QWF72ofwJh6ekgnCt9Xx6C
@cailmdaley
cailmdaley force-pushed the feat/sacc-2-sacc-io branch from 55cb97d to 4489dd4 Compare July 16, 2026 00:39
cailmdaley and others added 5 commits July 16, 2026 11:05
Adversarial review (pass 2) findings:

- Sacc.indices returns an EMPTY array (warning only) on an unmatched
  selection; every reader, extract(), and covariance-block selector now
  funnels through a shared _indices guard that raises instead — a typo'd
  tag or non-bitwise-identical float scale cut can no longer propagate
  empty arrays downstream (assemble_covariance previously died with an
  opaque IndexError on the same path).
- get_cosebis(scale_cut=None) on a file carrying several scale cuts
  silently concatenated them; now raises and asks for an explicit cut.
- update_statistic let two sub points claim the same target point
  (last-write-wins); now raises.
- merge: the library keeps the FIRST input's tracer on a name clash with
  no equality check; shared tracers are now verified identical (z, nz)
  across inputs before concatenation.

7 regression tests; 48 total.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…abulary

Rebuild the per-part-at-birth blinding stack on top of feat/sacc-2-sacc-io.
sacc_io.py is now PR-2's canonical module verbatim + a single appended
gather(): the terminal assembly delegates its tracer/point/covariance
assembly to PR-2's merge() and adds only the blind-custody call
(assert_consistent_blind) and shared-stamp write.

The fail-closed load gate lives in sacc_io.load() per the PRD ("sacc_io fails
closed on load"); the blinding tooling uses allow_unblinded=True explicitly
where it must read a not-yet-blinded real vector (blind_part). save() now
carries the required type= (inherited from each part's provenance). Grid
vocabulary swept coarse->reporting, fine->integration across blinding.py,
blinding_theory.py, b_modes.py, blind_data_vector.py, and the blinding tests.

_extract_block/_set_values stay index-based (a code comment says why):
Smokescreen's ConcealDataVector aligns theory_fn output to the sub-SACC mean
by row position, so the block must be carved/written by contiguous integer
index, not by PR-2's tag-matching extract()/update_statistic().

Escrow/custody/CAMB<->CCL machinery is preserved exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…tion

The branch previously carried a stale vendored copy of sacc_io.py /
test_sacc_io.py from before PR2's review rounds. Rebuilt directly on
feat/sacc-2-sacc-io (8bd3817) so the canonical module is inherited,
bringing over the cosmo_val + Snakemake born-as-SACC migration work
from the old tip unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…nfig

Three integration-drift fixes surfaced by the reconciled base:

- test_ac6_ac8 / test_ac8_dotted: the end-to-end fixtures now pass
  type="mock" to sio.save (PR-2 requires the provenance stamp); the parts
  are mocks, blind_part re-saves inheriting that type.
- test_ac13: the CosmoSIS halofit config moved to
  cosmo_inference/cosmosis_config/templates/ on develop; point the AC13
  assertion at the new path (token unchanged: mead2020_feedback).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
…rename

Sweep the migration code onto PR2's canonical vocabulary and contracts:

- grid='coarse'/'fine' -> 'reporting'/'integration' everywhere (writer
  calls, readers, tests), including the internal DAG intermediates:
  {version}_xi_coarse* -> _xi_reporting*, _xi_fine -> _xi_integration,
  the xi_coarse/xi_fine Snakemake output keys, CANONICAL part names,
  cv_xi_reporting_sacc, write_xi_integration_sacc.
- save(s, path) -> save(s, path, type=...): 'data' at every production
  writer (the cosmo_val pipeline measures the real UNIONS catalogues;
  GLASS mocks do not flow through these writers), 'mock' for synthetic
  test fixtures. assemble_sacc propagates its parts' type stamp rather
  than hardcoding, so mock parts assemble into a mock analysis file.
- load(path) -> fail-closed load: pipeline-internal readbacks of
  freshly written pre-blind data parts pass allow_unblinded=True
  (blinding is a downstream Smokescreen step); mock fixtures load freely.

Readers raising on unmatched selections needed no call-site changes:
every get_* reads a statistic guaranteed present in the file just
written or assembled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WzUt7VbtXwr2SCHUdiQTyt
@cailmdaley
cailmdaley force-pushed the feat/sacc-6-blinding branch from 4b058f7 to c09b063 Compare July 16, 2026 09:55
cailmdaley and others added 6 commits July 18, 2026 11:46
Consistency follows tagging semantics: the grid tag declares which
binning a set of points lives on, so all same-length theta arrays under
one tag value must be bitwise identical (sacc never validates angles,
and grids diverging at floating-point level choke CosmoSIS downstream).
Different lengths within a tag group pass (scale-cut subsets); grids
under different tag values are unconstrained (reporting vs integration
differ by design).

The rho/tau/pure-EB writers now tag their points grid="reporting" by
default (overridable via grid=), so they join xi's consistency group
and no untagged group appears in our own files.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MN9VazXKHUHQg16kiG7Ufk
…seudo-Cl grid="reporting"

The theta consistency guard generalizes to both angular domains: theta
and ell points are grouped separately by grid tag value, and within a
tag value all same-length grids must be bitwise identical. Two ell
series sharing a grid must also carry equal bandpower windows (window
ells and weight matrix); series without windows skip that check.

add_pseudo_cl now stamps grid="reporting" on every point by default
(overridable), joining the merge guard's consistency groups; since
sacc's add_ell_cl accepts no extra tags, its per-point insertion
(ell + shared window + window_ind) is inlined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MN9VazXKHUHQg16kiG7Ufk
assemble_covariance now passes sacc.add_covariance a list of per-block
matrices instead of a dense zero-filled N×N array. sacc.BaseCovariance.make
turns a list into a BlockDiagonalCovariance (one FITS table per block, Σ
block² on disk), so cross-blocks are zero and implicit rather than
materialized. Validation (contiguous/ascending indices, no gap/overlap,
square blocks matching their index span) is unchanged.

Every existing consumer already read through the polymorphic .dense
property, so only the type assertions needed updating (FullCovariance ->
BlockDiagonalCovariance). Added a test that merging two files that already
carry a BlockDiagonalCovariance (assembled via assemble_covariance) stays
block-diagonal through merge and a save/load round-trip. Updated the module
docstring's storage-cost discussion accordingly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pure_eb.py's calculate_pure_eb/plot_pure_eb defaulted nbins_int=100;
cosebis.py's calculate_cosebis already defaulted to 1000, and every
production config (papers/bmodes, papers/cosmo_val fiducial/pure_eb
blocks) already overrides to 1000. This aligns the function-signature
default with what every caller actually uses; config plumbing is
untouched, so any explicit override still wins.

The papers/cosmo_val/config/config.yaml cosebis.nbins_int (currently
2000, production numerics) is deliberately left unchanged — see report.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Writers no longer force components an analysis may not have computed:
add_pseudo_cl's BB/EB, add_cosebis's Bn, and add_pure_eb's B/ambiguous
blocks now default to None and are simply not written when omitted
(add_pure_eb still requires each +/- pair together). Structural
arguments (grids, tracer/bin identifiers, the value for a component
you ARE adding) remain required with no default.

Composite readers (get_pseudo_cl, get_cosebis, get_pure_eb) return
None / omit the key for an absent optional component instead of
raising, via a new _mean_optional helper; a selection naming a missing
component explicitly (s.indices, _mean, extract) still fails loud, per
the existing empty-selection guard. merge() and assemble_covariance
work unchanged on files with only a subset of components.

Documents the optionality contract in the module docstring, and adds
partial-file round-trip, merge, and explicit-selection-fails-loud
tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…fault

cosebis.nbins_int was 2000; every other integration-grid entry in this
config (pure_eb, the fiducial block) is already 1000, and cosebis.py's
own function defaults are 1000. Unify.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Base automatically changed from feat/sacc-2-sacc-io to develop July 21, 2026 12:24
cailmdaley and others added 2 commits July 21, 2026 14:28
Fold the integration-grid ξ± into the single terminal {version}.sacc, and
make part loading fail closed on unblinded real data.

Integration ξ± (grid='integration') is no longer its own terminal product.
The xi_highres part is now gathered by rule assemble_sacc into {version}.sacc
as tagged points, next to the reporting ξ± block. It is fiducial-only (the
10k-bin MPI run emits only the fiducial part), so it joins the fiducial
version's terminal file alone. assemble_sacc.py adds xi_integration to
CANONICAL; its own DiagonalCovariance passes straight through.

assemble_sacc.py no longer loads every part with allow_unblinded=True. The
run type (data|mock, from config, default data) gates it: mock runs load
freely, data runs fail closed unless a part carries the concealed=True stamp.
This is the seam for PR #253's blind-at-birth — a concealed data part then
assembles with allow_unblinded=False untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EaL7prmKUHwJQcDyW3LoxD
cailmdaley and others added 2 commits September 29, 2026 11:20
…custody

The scaffolding removal, the private COSMO_INFERENCE paragraph and the
non-custody docstrings are chore/cosmo-val-cleanup's; run_rho_tau passes its
rule's custody token and nothing else changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cailmdaley and others added 2 commits September 29, 2026 17:07
…y you pass

blinding.STANDARD declares each estimator cosmo_val computes and its rule
once: ξ± and Cℓ_EE shifted by the default PyCCL theory, Cℓ_BB/EB unshifted,
COSEBIs and pure-E/B derived from ξ±, ρ/τ signal-free. Any other data type
is shifted when its birth passes a function,
sacc_io.save(s, path, custody=c, theory={data_type: f}), and is refused
under a blind without one; its sealed rows pass later derivations as copies.

γt leaves until cosmo_val calculates it: GAMMA_T/GAMMA_X, add_lens,
add_gamma_t, the lens bias, source_lens and the null-lens exemption go, with
their tests. The secrecy test goes; Blind still holds only a name and path.
The README and CLAUDE.md state the calculate-then-save rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GJJyGbKSuEjjX9m2zdv6yi
cailmdaley and others added 11 commits October 1, 2026 04:10
theory.shear predicts ξ± and Cℓ_EE for every row of a SACC (zeros for
Cℓ_BB/EB), theory.none gives zeros. blinding.conceal hands the theory to
smokescreen's concealing_factor and adds the vector; a failure names only
its exception type. sacc_io.save conceals a birth under a Blind and stamps
its name, or stamps a derivation with its inputs' one blind.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…more

An entry must declare blind: none or blind: <name>; entries reading one
shear file declare one blind. The catalogue-path helpers stay here,
stdlib-only for the host.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CosmologyValidation.blind(version) opens the blind its catalogue config
declares; producer rules carry it as params.blind, so flipping it reruns
what it touches, and assembly refuses parts stamped otherwise. A launch
prints each catalogue's blind.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…riance

calculate_pure_eb requires cov_path_int; the interactive jackknife path
is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing in the repository calls them; ξ± and ⟨M_ap²⟩ live in cosmo_val.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
conceal moves each row by t(hidden) − t(fiducial), a custom theory blinds a
custom type, theory.none stamps without shifting, derivations inherit one
stamp, catalogues declare their blind, B-modes respond to the shift only
through each transform; DAG tests for the launch line and a blind flip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…art can't be shifted twice

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
…hetic helper

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UzeJqdtGgoeWizre32L5fD
@cailmdaley

Copy link
Copy Markdown
Collaborator Author

@LisaGoh and @sachaguer this is ready for feedback now. some open questions to settle and some oversights i'm sure, but i worked hard to get this as simple as possible. i think it's pretty nice. make sure to click the dropdown in the PR description to appreciate the full shifted data vector plots!

cailmdaley and others added 3 commits October 2, 2026 01:54
Pure E/B takes #373's design: one fixed linear operator on the
integration-grid ξ± part with covariance K C Kᵀ; the MC covariance and the
chunked precompute go. The blinding stays: cv_pure_eb and calculate_pure_eb
read the sealed integration part through sacc_io.xi_correlation (which
carries the pair weights, so get_xi_weight is not added), and the pure-E/B
part is saved derived_from that part. papers/bmodes/pure_eb_modes.py reads
the SACC part; the B-mode blinding test uses the operator.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r go

blinding.py now holds the theory (shear, fiducial, cosmology, and
no_signal for statistics without cosmological signal) and blind_of, which
reads an already-resolved catalogue entry's `blind:` and refuses a missing
one or one shear file read under two blinds. Its scientific imports live
inside the functions, so workflow/common.py loads it by path on the host.
custody.py and theory.py are deleted.

Catalogue and seed resolution are develop's again (_materialize_seed_path,
catalogue_entry, get_shear_catalog); CosmologyValidation.blind and
common.blind_of call blinding.blind_of on the resolved entry. The per-launch
[blind] banner is gone; configure() still refuses a run catalogue that
declares no blind before any job is built. Patch centres are written as on
develop (PR #367 owns them).

Tests keep what guards blinding: ξ±/Cℓ_EE shift and BB/EB do not; pure-E/B
under a blind moves E and not B (develop's operator); seal refuses a
double stamp; one file under two blinds is refused; a flip reruns the
part; an undeclared catalogue stops the launch. _synthetic.py,
test_blinding_bmodes.py, test_custody.py and test_toy_run.py go.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rsions resolve on the host

CosmologyValidation resolved the selected entry's paths before the
one-file-one-blind check, so a twin entry under another blind passed.
The workflow's catalogue_name strips _seed<N> as well as _leak_corr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186WiB28hK7DK2pYaTBLL6j
@cailmdaley
cailmdaley marked this pull request as ready for review October 2, 2026 10:28
@cailmdaley

cailmdaley commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

@AstroMike also happy to hear your opinion about this and whether it's flexible enough for your delta sigma example. i think the only work required to blind a new data vector is to a) get your data vector into sacc format and b) pass a function that calculates a theory curve when you save the data vector to disk. the data vector itself needn't be calculated by sp_validation as long as it's in SACC format.

@LisaGoh

LisaGoh commented Oct 2, 2026

Copy link
Copy Markdown
Member

Thank you for this Cail! I think I'd need to try running the whole thing to fully understand the logic of the pipeline, but for now I think it's good to go! With this we should also start thinking about blinding protocol (linked to the open questions you posted): when it happens, by whom, and on which catalogues.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Smokescreen blinding wiring (fork protocol, three theory backends, custody)

2 participants