Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,4 @@ _build/
.env
**/.DS_Store
build/
.hypothesis/
1 change: 1 addition & 0 deletions changelog.d/530.changed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
US files written before the WIC take-up mapping are no longer reused, because they may have lost the draw or were calculated without it. `Simulation.load()` raises for a saved US output that has no record of renamed stored inputs (`Simulation.ensure()` calculates it again), `load_datasets` raises for such a year file, and `ensure_datasets` creates such year files again.
1 change: 1 addition & 0 deletions changelog.d/530.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Map the stored US WIC take-up draw `would_claim_wic` onto `takes_up_wic_if_eligible` when loading US data, so policyengine-us 2.x no longer gives WIC to every WIC-eligible person in data that stores the draw under its old name, such as the certified default. `Simulation.run()`, `managed_microsimulation` and `create_datasets` apply the mapping and record the renames applied (output dataset metadata, `release_bundle`, saved output and year files, run records and `policyengine_bundle`). A stored table that is not in the simulation's person order, or that stores the draw for only part of a year, is refused. The mapping turns itself off once the data stores `takes_up_wic_if_eligible` or the engine defines `would_claim_wic` again.
47 changes: 47 additions & 0 deletions docs/microsim.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,6 +324,53 @@ mode. GCS dataset URIs are not supported.
For managed simulations, `sim.policyengine_bundle` records the actual source
package, repository type, revision, verified SHA-256, and local path.

## Renamed inputs in stored data

A country engine sets a stored column as an input only when it defines a
variable of that name. When policyengine-us renames an input, data written
before the rename keeps the old name, and the engine would skip it.
`policyengine.tax_benefit_models.us.legacy_inputs.LEGACY_INPUT_RENAMES` lists
each such rename; today it holds one entry, `would_claim_wic` →
`takes_up_wic_if_eligible` (the WIC take-up draw, PolicyEngine/microcosm#1026).

Every US load path applies it: `Simulation.run()`, `managed_microsimulation`
and `create_datasets` (and so `ensure_datasets` when it creates year files).
A rename applies when the data stores the old name, the engine does not define
the old name but does define the new one, and the data does not already store
the new name. The new input is then set from the stored values for every month
of every dataset year. The stored table must list the simulation's entity IDs
in the simulation's order, or loading fails rather than attach values to the
wrong people. Only values stored for a whole year are mapped, so a file that
stores the old name for part of a year, and not the new name, is refused. The
mapping turns itself off once the data stores the new name, or if the engine
defines the old name again.

The renames applied are recorded as `{old: new}` (`{}` when none applied):

- `Simulation.run()`: `simulation.output_dataset.metadata["legacy_input_renames"]`,
also shown as `simulation.release_bundle["legacy_input_renames"]`. It
includes renames applied when the input year file was cut, so a run over a
`create_datasets` year file records the rename even though the file already
stores the new name. A US `save()` writes the record into the output file,
`load()` restores it, and a run record's `results.json` binds it;
- `managed_microsimulation`: `sim.policyengine_bundle["legacy_input_renames"]`;
- `create_datasets`: each year file, and each returned dataset's
`metadata["legacy_input_renames"]`.

`{}` means that no stored column was mapped. It does not show that the data
carried the draw: data that stores neither name runs with the new input's
default, which for WIC is full take-up.

Files written before this mapping existed have no record, and are not reused:

- A saved US output may have been calculated without the mapped inputs.
`load()` refuses it, and `ensure()` runs the simulation again and saves the new
output.
- A year file that `ensure_datasets` or `create_datasets` wrote stores
neither name, so the draw is lost from it. `ensure_datasets` creates such
year files again, and `load_datasets` refuses them. A year file opened
directly, as `PolicyEngineUSDataset(filepath=...)`, is not checked.

## Pinned model versions

Every `policyengine` release pins specific country-model and country-data versions so results are reproducible. `pe.us.model` and `pe.uk.model` expose the pinned `TaxBenefitModelVersion`.
Expand Down
2 changes: 1 addition & 1 deletion docs/run-records.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ The directory contains:
| `bundle.trace.tro.jsonld` | The certified bundle TRO (model + data pins) |
| `reform.json` | The reform as parameter values with effective dates |
| `input.json` | Input dataset hash, dynamics, scoping, extra variables |
| `results.json` | Output dataset hash and per-entity table summaries |
| `results.json` | Output dataset hash and per-entity table summaries; for US runs, the SPM receipt and the renamed stored inputs mapped onto live ones (see [Renamed inputs in stored data](microsim.md#renamed-inputs-in-stored-data)) |

All payload files are written with the same canonical JSON used for
hashing, so the record verifies offline exactly as written.
Expand Down
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,7 @@ dev = [
"towncrier>=24.8.0",
"mypy>=1.11.0",
"pytest-cov>=5.0.0",
"hypothesis>=6.100.0",
"policyengine-core==3.32.5",
"policyengine-us==2.2.1",
"policyengine-uk==2.90.2",
Expand Down
9 changes: 9 additions & 0 deletions src/policyengine/core/run_record.py
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,15 @@ def build_simulation_run_record_payloads(
input_payload["spm"] = simulation.spm_config
results_payload["spm"] = simulation.spm_provenance()

# Stored inputs the country loader mapped onto renamed live inputs (for
# example the US WIC take-up draw) change the results, so the record
# binds them. See ``policyengine.tax_benefit_models.us.legacy_inputs``.
renames = (getattr(simulation.output_dataset, "metadata", None) or {}).get(
"legacy_input_renames"
)
if renames is not None:
results_payload["legacy_input_renames"] = dict(renames)

return {"reform": reform, "input": input_payload, "results": results_payload}


Expand Down
10 changes: 9 additions & 1 deletion src/policyengine/core/simulation.py
Original file line number Diff line number Diff line change
Expand Up @@ -240,7 +240,7 @@ def release_bundle(self) -> dict[str, Any]:
if self.tax_benefit_model_version is not None
else {}
)
result = {
result: dict[str, Any] = {
**bundle,
"dataset_filepath": self.dataset.filepath
if self.dataset is not None
Expand All @@ -250,4 +250,12 @@ def release_bundle(self) -> dict[str, Any]:
recorded = getattr(self.output_dataset, "metadata", {}).get("spm_config")
result["spm_config"] = dict(recorded or self.spm_config)
result["spm"] = self.spm_provenance()
# Stored inputs the country loader mapped onto renamed live inputs
# (for example the US WIC take-up draw), as the run recorded them.
# See ``policyengine.tax_benefit_models.us.legacy_inputs``.
renames = (getattr(self.output_dataset, "metadata", None) or {}).get(
"legacy_input_renames"
)
if renames is not None:
result["legacy_input_renames"] = dict(renames)
return result
32 changes: 31 additions & 1 deletion src/policyengine/tax_benefit_models/common/model_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -358,6 +358,9 @@ def save(self, simulation: Simulation) -> None:
serialized_spm = None
if self.country_code == "us":
from policyengine.core.spm import SPMProvenance
from policyengine.tax_benefit_models.us.legacy_inputs import (
RENAMES_RECORD_KEY,
)

receipt = SPMProvenance.model_validate(simulation.spm_provenance())
if (
Expand All @@ -374,6 +377,17 @@ def save(self, simulation: Simulation) -> None:
},
sort_keys=True,
)
# ``PolicyEngineUSDataset.save()`` writes this record into the
# file, and ``load()`` refuses an output without it.
renames = (getattr(simulation.output_dataset, "metadata", None) or {}).get(
RENAMES_RECORD_KEY
)
if renames is None:
raise ValueError(
"This US output does not record which renamed stored "
"inputs were mapped when it was calculated; run again "
"before saving"
)
simulation.output_dataset.save()
if serialized_spm is not None:
# Store UTF-8 JSON in a dataset rather than an attribute: the
Expand All @@ -399,13 +413,18 @@ def load(self, simulation: Simulation) -> None:
receipt = None
if self.country_code == "us":
from policyengine.core.spm import SPMProvenance
from policyengine.tax_benefit_models.us.legacy_inputs import (
RENAMES_RECORD_KEY,
read_renames_record,
)

with h5py.File(filepath, "r") as stream:
raw = (
stream["policyengine_spm"].asstr()[()]
if "policyengine_spm" in stream
else None
)
recorded_renames = read_renames_record(filepath)
if raw is None:
raise ValueError(
"Saved US simulation has no SPM configuration or receipt"
Expand All @@ -414,6 +433,15 @@ def load(self, simulation: Simulation) -> None:
if recorded["config"] != simulation.spm_config:
raise ValueError("Saved US simulation uses different SPM settings")
receipt = SPMProvenance.model_validate(recorded["provenance"])
# Outputs saved before stored inputs were mapped onto renamed
# live inputs were calculated without them (for example with
# every WIC-eligible person taking WIC up), so they are not
# reused. ``Simulation.ensure()`` runs such a simulation again.
if recorded_renames is None:
raise ValueError(
"Saved US simulation predates the mapping of renamed "
"stored inputs (it has no record of them); run it again"
)

simulation.output_dataset = self._dataset_class(
id=simulation.id,
Expand All @@ -428,7 +456,9 @@ def load(self, simulation: Simulation) -> None:

simulation.spm_receipt = receipt
simulation.spm = SPMSelection.model_validate(recorded["config"])
simulation.output_dataset.metadata["spm_config"] = recorded["config"]
simulation.output_dataset.metadata.update(
{"spm_config": recorded["config"], RENAMES_RECORD_KEY: recorded_renames}
)

if os.path.exists(filepath):
simulation.created_at = datetime.datetime.fromtimestamp(
Expand Down
92 changes: 86 additions & 6 deletions src/policyengine/tax_benefit_models/us/datasets.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,14 @@
from policyengine.tax_benefit_models.common.model_version import (
build_runtime_dataset_provenance,
)
from policyengine.tax_benefit_models.us.legacy_inputs import (
RENAMES_RECORD_KEY,
apply_legacy_input_renames_to_microsimulation,
check_yearly_periods,
pending_legacy_input_renames,
read_renames_record,
write_renames_record,
)
from policyengine.utils.hashing import sha256_file


Expand Down Expand Up @@ -97,12 +105,22 @@ def save(self) -> None:
store["spm_unit"] = pd.DataFrame(self.data.spm_unit)
store["tax_unit"] = pd.DataFrame(self.data.tax_unit)
store["household"] = pd.DataFrame(self.data.household)
# The renamed stored inputs mapped when this data was cut or
# calculated. Its presence marks a year file or output as written
# with the mapping (see ``legacy_inputs.RENAMES_H5_DATASET``).
renames = self.metadata.get(RENAMES_RECORD_KEY)
if renames is not None:
write_renames_record(filepath, renames)

def load(self) -> None:
"""Load dataset from HDF5 file into this instance."""
filepath = self.filepath
if _is_policyengine_core_h5(Path(filepath)):
self.data = _load_policyengine_core_h5(Path(filepath), self.year)
path = Path(filepath)
renames = read_renames_record(path)
if renames is not None:
self.metadata[RENAMES_RECORD_KEY] = renames
if _is_policyengine_core_h5(path):
self.data = _load_policyengine_core_h5(path, self.year)
return

with pd.HDFStore(filepath, mode="r") as store:
Expand Down Expand Up @@ -194,10 +212,20 @@ def _core_h5_entity_lengths(h5_file: h5py.File, year: int) -> dict[str, int]:
return lengths


def _core_h5_variable_entities() -> dict[str, str]:
def _core_h5_variable_entities() -> tuple[dict[str, str], dict[str, str]]:
"""Return each variable's entity, and the pending renames ``{legacy: live}``."""
from policyengine_us.system import system

return {name: variable.entity.key for name, variable in system.variables.items()}
entities = {
name: variable.entity.key for name, variable in system.variables.items()
}
# A stored input the engine has since renamed belongs to its live input's
# entity. Without this its entity is guessed from its length, and a column
# as long as two entities is dropped before the rename can map it.
pending = pending_legacy_input_renames(system.variables)
for legacy, live in pending.items():
entities[legacy] = entities[live]
return entities, pending


def _validate_entity_ids(data: dict[str, pd.DataFrame]) -> None:
Expand Down Expand Up @@ -288,11 +316,24 @@ def _load_policyengine_core_h5(path: Path, year: int) -> USYearData:
"""Load a PolicyEngine core variable/period H5 into .py entity DataFrames."""

data = {entity: pd.DataFrame() for entity in US_ENTITY_KEYS}
variable_entities = _core_h5_variable_entities()
variable_entities, pending_renames = _core_h5_variable_entities()

with h5py.File(path, "r") as h5_file:
entity_lengths = _core_h5_entity_lengths(h5_file, year)
for variable_name in h5_file.keys():
live_name = pending_renames.get(variable_name)
if live_name is not None and live_name not in h5_file:
# A stored legacy column is mapped onto every month of the
# year, so a part-year value is refused, as it is when
# ``managed_microsimulation`` reads this file. A file that
# also stores the live name loads that natively, and its
# legacy column is left alone on both paths.
node = h5_file[variable_name]
check_yearly_periods(
variable_name,
node.keys() if isinstance(node, h5py.Group) else [],
path,
)
values = _read_core_h5_period_values(h5_file, variable_name, year)
entity = variable_entities.get(variable_name)
if entity is None:
Expand Down Expand Up @@ -353,6 +394,9 @@ def create_datasets(
)
dataset_stem = source.name
sim = Microsimulation(dataset=source.path, spm=resolve_spm_selection())
# Map stored inputs the engine has since renamed before extracting,
# so each year's file stores the live input rather than dropping it.
legacy_input_renames = apply_legacy_input_renames_to_microsimulation(sim)

for year in years:
# Get all input variables from the simulation
Expand Down Expand Up @@ -496,6 +540,9 @@ def create_datasets(
description=f"US Dataset for year {year} based on {dataset_stem}",
filepath=f"{data_folder}/{dataset_stem}_year_{year}.h5",
year=int(year),
metadata={
RENAMES_RECORD_KEY: dict(sorted(legacy_input_renames.items()))
},
data=USYearData(
person=MicroDataFrame(person_df, weights="person_weight"),
household=MicroDataFrame(household_df, weights="household_weight"),
Expand All @@ -515,13 +562,28 @@ def create_datasets(
return result


def _year_file_records_renames(path: Path) -> bool:
"""Return whether a year file records the renamed stored inputs mapped.

``create_datasets`` writes the record into every year file (``{}`` when
nothing needed mapping). Files written before it did may have lost a
renamed input such as the WIC take-up draw.
"""
return read_renames_record(path) is not None


def load_datasets(
datasets: Optional[list[str]] = None,
years: list[int] = [2024, 2025, 2026, 2027, 2028],
data_folder: str = "./data",
) -> dict[str, PolicyEngineUSDataset]:
"""Load PolicyEngineUSDataset instances from saved HDF5 files.

A year file without the record of renamed stored inputs that
``create_datasets`` writes was cut before those inputs were mapped, and
may have lost one (see ``legacy_inputs.RENAMES_H5_DATASET``), so it is
refused.

Args:
datasets: List of HuggingFace dataset paths (used to derive file names)
years: List of years to load data for
Expand All @@ -537,6 +599,17 @@ def load_datasets(
dataset_stem = dataset_logical_name(resolved_dataset)
for year in years:
filepath = f"{data_folder}/{dataset_stem}_year_{year}.h5"
if Path(filepath).exists() and not _year_file_records_renames(
Path(filepath)
):
raise ValueError(
f"US year file {filepath} has no record of the renamed "
"stored inputs mapped when it was cut, so it was written "
"before policyengine.py mapped them and may have lost "
"the WIC take-up draw (every WIC-eligible person would "
"then take WIC up). Regenerate it with ensure_datasets() "
"or create_datasets()."
)
us_dataset = PolicyEngineUSDataset(
name=f"{dataset_stem}-year-{year}",
description=f"US Dataset for year {year} based on {dataset_stem}",
Expand Down Expand Up @@ -1182,6 +1255,11 @@ def ensure_datasets(
) -> dict[str, PolicyEngineUSDataset]:
"""Ensure datasets exist, loading if available or creating if not.

Year files without the record of renamed stored inputs that
``create_datasets`` writes were cut before those inputs were mapped, and
may have lost one such as the WIC take-up draw, so they are created
again rather than loaded.

Args:
datasets: List of HuggingFace dataset paths
years: List of years to load/create data for
Expand All @@ -1199,7 +1277,9 @@ def ensure_datasets(
dataset_stem = dataset_logical_name(resolved_dataset)
for year in years:
filepath = Path(f"{data_folder}/{dataset_stem}_year_{year}.h5")
if not filepath.exists():
# A year file written before renamed stored inputs were mapped
# may have lost them, so it is regenerated rather than reused.
if not filepath.exists() or not _year_file_records_renames(filepath):
all_exist = False
break
if not all_exist:
Expand Down
Loading
Loading