Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
73f788f
Give each per-variable QRF model its own seed
vahid-ahmadi Sep 16, 2026
146f0c4
Fix issues from review: bound QRF seeds and preserve tuning streams
juaristi22 Sep 17, 2026
a8c94bd
Fail loudly on an unknown variable and an invalid seed
vahid-ahmadi Sep 21, 2026
79ac62b
Prune failed Matching trials instead of scoring the training mean
vahid-ahmadi Sep 16, 2026
5a4f8d6
Fix issues from review: track Matching prediction failures
juaristi22 Sep 17, 2026
6a94c6d
Reset the failure counter, report the cause, expose it on the result
vahid-ahmadi Sep 21, 2026
09ba11b
Fix indentation that moved a raise out of its except block
vahid-ahmadi Sep 21, 2026
62d03ce
Fix issues from review: restore lint compatibility and preserve match…
juaristi22 Sep 21, 2026
d6ca1d5
Declare subpackages instead of shipping the source tree as package data
vahid-ahmadi Sep 16, 2026
60f822a
Fix issues from review: format documentation examples
juaristi22 Sep 17, 2026
6a0e9b7
Add per-version classifiers, an author email, and a find exclude
vahid-ahmadi Sep 21, 2026
ae2d831
Fix tests that pass for the wrong reason, and export ZeroInflatedImputer
vahid-ahmadi Sep 16, 2026
9a9929c
Allow Matching's own message in the missing-predictor test
vahid-ahmadi Sep 16, 2026
72542b3
Fix issues from review: preserve public config compatibility
juaristi22 Sep 17, 2026
94fb136
Act on review: loosen frozen tests, correct the weight docstring
vahid-ahmadi Sep 21, 2026
e4fc210
Give each per-variable QRF model its own seed
vahid-ahmadi Sep 16, 2026
e9fec87
Prune failed Matching trials instead of scoring the training mean
vahid-ahmadi Sep 16, 2026
1309a2a
Fix pre-submission imputation and evaluation correctness
juaristi22 Sep 16, 2026
398336b
Fix issues from review: preserve numeric and weight semantics
juaristi22 Sep 17, 2026
8fbd31e
Fix issues from review: repair prediction contracts and compatibility
juaristi22 Sep 21, 2026
9af2ba6
Integrate repaired QRF seed handling from PR #214
juaristi22 Sep 21, 2026
68d6acb
Integrate Matching failure reporting from PR #215
juaristi22 Sep 21, 2026
17d8699
Integrate packaging corrections from PR #216
juaristi22 Sep 21, 2026
5dfc865
Integrate public exports and compatibility cleanup from PR #217
juaristi22 Sep 21, 2026
296e47d
Fix issues from review: use native R donation-class argument
juaristi22 Sep 21, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/pr_code_changes.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -36,8 +36,8 @@ jobs:
echo "Changed files:"
echo "$CHANGED_FILES"

# Check if any MDN-related files were changed
if echo "$CHANGED_FILES" | grep -qE "(mdn|MDN)"; then
# Shared fitting, preprocessing and scoring code also affects MDN.
if echo "$CHANGED_FILES" | grep -qE '(^microimpute/|^tests/|^pyproject\.toml$|^uv\.lock$|^\.github/workflows/pr_code_changes\.yaml$|mdn|MDN)'; then
echo "mdn_changed=true" >> $GITHUB_OUTPUT
echo "MDN-related files were changed"
else
Expand Down
1 change: 1 addition & 0 deletions changelog.d/202.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Fit preprocessing statistics on donor training rows and reuse them for receivers, including single-row predictions. Cross-validation fits transformations within each fold and scores inverse-transformed predictions on the original target scale. Returned fitted models now replay donor preprocessing on future raw-data predictions and restore target units.
1 change: 1 addition & 0 deletions changelog.d/203.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Compute deterministic zero-inflated quantiles by inverting the ordered negative, zero, and positive mixture CDF. Preserve stochastic draws when quantiles are omitted, and reject unsupported sequential marginal quantiles or component predictions outside their sign support.
1 change: 1 addition & 0 deletions changelog.d/204.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Run core autoimpute tests without optional Matching or MDN dependencies instead of failing with an undefined Matching name.
1 change: 1 addition & 0 deletions changelog.d/205.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Apply the configured QRF defaults with at least 20 training observations per leaf and full leaf distributions. Compute survey-weighted conditional quantiles using each tree's normalized leaf weights and bootstrap donor multiplicities, and query sampled quantiles without discretizing them onto a truncated grid.
1 change: 1 addition & 0 deletions changelog.d/206-distributions.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Validate and propagate numeric zero-inflated sample weights to gate classifiers and sign-specific component fits, including aligned Series and array weights. Give numeric components reproducible independent random seeds and forward component fit parameters.
1 change: 1 addition & 0 deletions changelog.d/206-matching-scoring.breaking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Matching.predict without quantiles now returns a DataFrame containing donor draws. Explicit quantiles and return_probs=True raise NotImplementedError. autoimpute excludes Matching from distributional model selection, while impute_all includes its draws when requested alongside a distributional model. Distributional predictor analysis rejects Matching at entry. Rerun comparisons that previously scored replicated donor draws or fabricated probabilities.
1 change: 1 addition & 0 deletions changelog.d/206-models.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Allow QuantReg to fit requested prediction quantiles lazily on its original donor data; keep intercept columns on homogeneous receivers and reject nonfinite OLS/QuantReg inputs. Matching now explicitly rejects conditional quantile and class-probability requests instead of relabeling donor draws, retains and subsets donor weights during tuning, uses the documented weighted StatMatch API, and reports failed-record counts on returned data frames. Failed tuning trials remain pruned rather than scored using fallback values. Normalize OLS survey weights so arbitrary weight units cannot change predictive quantiles. Matching's R donor draws now use reproducible advancing child seeds while preserving the caller's R RNG state.
1 change: 1 addition & 0 deletions changelog.d/206-pipeline.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Honor autoimpute train_size by reproducibly sampling donor rows, forward random_state to model fitting, evaluate requested quantile grids, and retune final parameters on all training rows instead of selecting the luckiest outer test fold. Report the standard deviation of fold-average losses correctly. Distributional QRF comparisons use independent target fits so reported quantiles are conditional on the original predictors; stochastic sequential donor draws remain available through direct QRF use.
1 change: 1 addition & 0 deletions changelog.d/206-qrf-quantiles.breaking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Add QRF(sequential=False) to estimate each target's conditional marginal quantiles using the original predictors only, with target-order-independent seeds. Sequential multi-target QRF now rejects explicit quantiles because same-quantile chaining does not calculate marginal quantiles; stochastic sequential draws remain supported. Distributional comparison and imputation helpers use independent QRF fits. Numeric hyperparameter tuning now evaluates the actual predicted median instead of scoring a random draw as a median.
2 changes: 2 additions & 0 deletions changelog.d/207.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
Per-variable QRF models now derive distinct seeds, so variables imputed together no longer share one random quantile per row and come out comonotonic. `QRF` also accepts a `seed` argument.
Derived seeds stay within the supported uint32 range and are used consistently during numeric and classification tuning and target-specific subsampling.
1 change: 1 addition & 0 deletions changelog.d/208.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Draw one independent OLS residual per receiver from a persistent random generator, advancing across targets and prediction calls while preserving reproducibility from the same seed. QuantReg random grid sampling also advances independently across calls and targets.
1 change: 1 addition & 0 deletions changelog.d/209-target-types.breaking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Numeric columns, including integer counts and 0/1 integers, now remain numeric regardless of cardinality or the values present in a fold. Regression predictions may be fractional or outside observed support. Declare categorical targets with target_types={"column": "categorical"}, pandas categorical dtype, or boolean dtype for binary categories. Use consistent declarations in fitting and scoring. Log-loss comparisons require actual probabilities from predict(return_probs=True) and raise ValueError for missing or invalid probabilities. Comparison quantiles must be nonempty and finite within [0, 1].
1 change: 1 addition & 0 deletions changelog.d/209.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Compute categorical comparison and cross-validation log loss from actual model probabilities. Reject class-label inputs rather than fabricating 0.99/0.01 probabilities, and align classes when a cross-validation fold lacks a category.
1 change: 1 addition & 0 deletions changelog.d/210.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Matching hyperparameter tuning now prunes a trial when matching fails, instead of silently scoring it as if it had predicted the training mean, and reports when no trial succeeds. Predictions report how many records could not be matched and reset the failure count on every returned result, including small unchunked predictions.
1 change: 1 addition & 0 deletions changelog.d/211.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
The wheel now declares its subpackages properly instead of sweeping the source tree in as package data, so a build from a working tree containing compiled bytecode or stray files no longer ships them. Adds `py.typed`, the full author list, per-version classifiers and project URLs.
1 change: 1 addition & 0 deletions changelog.d/212.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Compute normalized mutual information from consistent paired discretizations and natural-log entropies, correcting the previous mixture of incompatible entropy and mutual-information estimates. Continuous-variable bins are configurable; constants have zero information, including on the diagonal. Predictor analysis now scores categorical forecasts from their actual probabilities.
1 change: 1 addition & 0 deletions changelog.d/213-validation.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Reject duplicate predictor/target names, predictor-target overlap, duplicate DataFrame columns, incompatible numeric/string predictor dtypes, and infinite sample weights with actionable errors. Treat numeric int/float predictor dtypes as compatible.
1 change: 1 addition & 0 deletions changelog.d/213.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
The test suite now runs cleanly without `rpy2` or `pytorch-tabular` instead of erroring, and five tests that could pass without asserting anything now assert. `ZeroInflatedImputer` is exported from `microimpute` and `microimpute.models`. Keeps the legacy `VALID_YEARS` and `DEFAULT_MODEL_PARAMS` imports, and corrects the `Imputer.fit` weight docstring.
1 change: 1 addition & 0 deletions changelog.d/215-review.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Include the unmatched-record count in the default Matching prediction frame's metadata, consistently with explicit-quantile results.
1 change: 1 addition & 0 deletions changelog.d/219-adversarial.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Prevent integer overflow in quantile loss and Matching tuning; consistently honor numeric boolean targets, preserve survey weights when transforming a weight predictor, retain default constant-category probabilities, and support nullable numeric OLS predictors.
1 change: 1 addition & 0 deletions changelog.d/219-cache-compatibility.breaking.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Load historical fitted QRF, OLS and Matching pickles with missing random-state attributes; historical QRF preserves sequential conditioning. Refit from original donor data and regenerate saved predictions to adopt corrected weighted distributions: loading cannot reconstruct new donor-weight information from an old forest. Regenerate MDN caches when upgrading, using force_retrain=True. See the distributional workflow migration guide for target declarations, scoring and preprocessing replay.
1 change: 1 addition & 0 deletions changelog.d/219-review.fixed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Bound derived QRF and zero-inflated component seeds, reset target declarations on zero-inflated refits, and return exact point-mass probabilities for constant categorical targets. Honor numeric declarations in QuantReg final fitting. Preserve recipient/donor pairs and R column-major ordering in Matching fallback indices, and seed both native R adapters. Run MDN validation for shared implementation changes and bound setuptools below 82 for its Lightning dependency.
1 change: 1 addition & 0 deletions docs/_toc.yml
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ parts:
- file: imputation-benchmarking/index
sections:
- file: imputation-benchmarking/preprocessing
- file: imputation-benchmarking/migration
- file: imputation-benchmarking/cross-validation
- file: imputation-benchmarking/metrics
- file: imputation-benchmarking/visualizations
Expand Down
106 changes: 106 additions & 0 deletions docs/imputation-benchmarking/migration.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
# Migrating distributional imputation workflows

Declare target semantics before fitting or scoring. Numeric columns, including
integer counts and 0/1 integers, remain numeric. Strings, pandas categorical
columns and booleans use classification. Pass `target_types` to `fit`,
`autoimpute`, `cross_validate_model` and `compare_metrics` when the dtype does
not express the intended meaning. Supported declarations are `"numeric"`,
`"categorical"` and `"bool"`.

```python
import numpy as np
import pandas as pd
from microimpute.models import QRF
from microimpute.comparisons.metrics import compare_metrics

donor = pd.DataFrame({
"age": np.arange(20., 60.),
"count": np.tile([0, 1, 2, 3], 10),
"status": np.tile([0, 1], 20),
})
targets = ["count", "status"]
types = {"count": "numeric", "status": "categorical"}
fitted = QRF(sequential=False, seed=42).fit(
donor, ["age"], targets, target_types=types,
)
predictions = fitted.predict(
donor[["age"]], quantiles=[0.1, 0.5, 0.9], return_probs=True,
)
scores = compare_metrics(
donor, {"QRF": predictions}, targets, target_types=types,
)
```

Numeric regression can produce fractional counts or values outside a target's
observed support. Declaring a target categorical restricts predictions to fitted
classes, but changes the model and scoring metric. Choose the intended model
explicitly; rounding predictions does not establish a valid count distribution.

## QRF draws and quantiles

`QRF()` defaults to `sequential=True`. With multiple targets, each target uses
previous targets as predictors. Call `fitted.predict(receiver)` for a stochastic
joint donor draw. Repeated calls advance the fitted random streams.

Use `QRF(sequential=False)` for target-specific conditional marginal quantiles.
Each target then uses only the original predictors, and its seed depends on its
name rather than target ordering. A sequential fit with multiple targets rejects
explicit quantiles: chaining the same quantile across targets does not calculate
their marginal quantiles. Single-target fits support explicit quantiles.

The helper `microimpute.models.imputer.create_distributional_model(QRF, seed=42)`
constructs the independent mode used by comparison, cross-validation and
predictor-analysis workflows. Direct construction retains the sequential default.

Weighted QRF retains in-bag donor multiplicities and survey weights to calculate
the conditional empirical distribution. Retaining weighted leaves adds storage,
and prediction visits trees and rows in Python. Runtime and memory depend on
training size, forest settings and query size; this release does not establish a
universal performance improvement or a fixed memory bound.

## Categorical probabilities and scoring

Request `return_probs=True` when evaluating categorical targets. The result's
`"probabilities"` entry maps each target to a probability matrix and its ordered
`"classes"`. OLS, QRF and MDN include exact one-class point masses for constant
categorical and boolean targets. Numeric constants remain numeric forecasts.

`compare_metrics` aligns supplied class labels and probabilities. It raises
`ValueError` for missing or invalid categorical probabilities; hard labels cannot
substitute for probability forecasts. Comparison quantiles must be a nonempty
list of finite values in `[0, 1]`. Input validation also rejects duplicate column
names, overlapping predictors/targets, and incompatible predictor dtypes.

## Matching donor draws

`Matching.fit(...).predict(receiver)` returns a DataFrame containing one matched
donor draw per recipient. Quantile and probability requests raise
`NotImplementedError`. `autoimpute(..., impute_all=True, models=[OLS, Matching])`
includes Matching draws in the output while excluding Matching from
distributional ranking. Supply at least one distributional model for selection.
Distributional predictor analysis also rejects Matching at its public entry
points. Direct Matching hyperparameter tuning evaluates donor-draw errors.

Matching's seed controls native R adapter calls without changing the caller's
global R random stream. Both the default lazy adapter and directly imported
`microimpute.utils.statmatch_hotdeck.nnd_hotdeck_using_rpy2` support this behavior.
Custom adapters retain their existing argument contract. Weighted R matching
uses `RANDwNND.hotdeck`; constrained weighted matching remains unsupported.

## Existing fitted models and preprocessing

Historical pickled OLS, QRF and Matching results initialize missing random-state
attributes on loading. Historical QRF results preserve sequential conditioning.
Deserialization cannot reconstruct corrected survey-weight distributions from
an old fitted forest: refit from the original donor data and regenerate saved
predictions and comparison results to adopt the corrected algorithms.

For MDN, remove or regenerate old model-cache directories when upgrading; set
`force_retrain=True` when fitting to rebuild them. The MDN extra currently bounds
setuptools below 82 because its Lightning dependency imports `pkg_resources`.

Models returned by `autoimpute` retain transformations fitted on the donor data.
Pass raw receiver predictors to their `predict` method; the wrapper transforms
inputs and returns predictions in original units. Do not normalize the receiver
independently or transform it a second time. Cross-validation fits transformations
inside each training fold and scores predictions in original units.
4 changes: 3 additions & 1 deletion docs/models/imputer/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,9 @@ The `Imputer` class is the abstract base class that defines the common interface

## Key features

All models share standardized `fit()` and `predict()` methods, so they can be used interchangeably regardless of underlying implementation. All imputers also share functionality like weighted data handling through a `weight_col` parameter.
All models share `fit()` and `predict()` methods, with model-specific capabilities. Matching returns donor draws and rejects quantile and probability requests. MDN rejects sample weights; OLS, QRF and Matching support weighted fitting through `weight_col`.

Use `fit(..., target_types={"status": "categorical"})` to declare numeric-coded categories explicitly. Numeric counts and 0/1 integers otherwise remain numeric. See the [migration guide](../../imputation-benchmarking/migration.md) for probability scoring, QRF modes and fitted-model compatibility.

The design enforces that `predict()` cannot be called before `fit()`. The base implementation also handles parameter and input data validation, so individual models don't need to duplicate those checks.

Expand Down
6 changes: 3 additions & 3 deletions docs/models/matching/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ The `Matching` model imputes missing values using nearest neighbor distance hot

## Variable type support

Matching handles any variable type: numerical, categorical, boolean, or mixed. Because it transfers actual observed values rather than generating predictions, it preserves the original data type and distribution of each variable.
Matching handles numerical, categorical, boolean and mixed targets. It transfers observed donor values, preserving their support and within-record target combinations. The distribution among recipients depends on donor selection and can differ from the donor marginal distribution.

## How it works

Expand All @@ -18,6 +18,6 @@ Because the imputed values are drawn from actually observed records, the natural

Matching is non-parametric: it makes no assumptions about the data distribution. This makes it useful when the data doesn't fit standard parametric models, or when the relationships between predictors and targets are hard to specify in closed form.

The method preserves the empirical distribution of the imputed variables. Since values come directly from observed data points, features like multimodality, skewness, and natural bounds are maintained. A model-based approach might smooth these away.
Donated values respect the observed target bounds and discrete categories. Nearest-neighbor selection determines how often donors contribute, so multimodality and skewness need assessment in the resulting recipient population.

One limitation is that Matching does not incorporate quantile information. It matches donor and receiver units identically regardless of the quantile being predicted, which means it cannot distinguish between different parts of the conditional distribution. It may also fail to capture non-linear predictor-target relationships despite producing a plausible marginal distribution.
`fitted.predict(receiver)` returns a DataFrame with one donor draw per recipient. Matching rejects explicit quantiles and `return_probs=True`, because a donor draw does not estimate conditional quantiles or class probabilities. AutoImpute excludes Matching from distributional ranking, but includes its draws when explicitly requested with `impute_all=True`. Distributional predictor analysis rejects Matching. See the [migration guide](../../imputation-benchmarking/migration.md).
Loading
Loading