The per-tile final_cat is FITS, while the campaign merge that concatenates the tiles writes HDF5 (hdf5_reconcile.py, lzf since #937). This issue proposes writing the per-tile file as HDF5 as well, with one dataset per column. The merge itself keeps one compound dataset per tile.
Why:
Not a speed project. The assessment behind #937–#940 found the pipeline's I/O cost came from access patterns, not FITS itself. After #939, a one-shot write of either format to node-local disk takes under a second. The gain here is consistency, plus a format that suits how make_cat builds the catalogue.
What it touches:
- make_cat's output, plus
merge_final_cat.py and hdf5_reconcile.py, which read per-tile files.
- The completeness and params pins: the declared output changes, which makes this a campaign boundary.
- Four or so test files.
- Anyone who reads per-tile
final_cat FITS directly.
Parquet is out of scope. If the analysis side moves to dataframe tooling, the merged HDF5 can be exported to Parquet then.
Claude Opus 5.5 on behalf of Cail
The per-tile
final_catis FITS, while the campaign merge that concatenates the tiles writes HDF5 (hdf5_reconcile.py, lzf since #937). This issue proposes writing the per-tile file as HDF5 as well, with one dataset per column. The merge itself keeps one compound dataset per tile.Why:
Not a speed project. The assessment behind #937–#940 found the pipeline's I/O cost came from access patterns, not FITS itself. After #939, a one-shot write of either format to node-local disk takes under a second. The gain here is consistency, plus a format that suits how make_cat builds the catalogue.
What it touches:
merge_final_cat.pyandhdf5_reconcile.py, which read per-tile files.final_catFITS directly.Parquet is out of scope. If the analysis side moves to dataframe tooling, the merged HDF5 can be exported to Parquet then.
Claude Opus 5.5 on behalf of Cail