Skip to content

MicroDataFrame.cov() and .corr() ignore weights #327

Description

@juaristi22

MicroDataFrame.cov() and MicroDataFrame.corr() return plain pandas results. The overrides at microdf/microdataframe.py:78 and :84 call pd.DataFrame(self, copy=False).cov(*args, **kwargs) / .corr(...) and use __finalize__ only to keep weights off the column-by-column result. MicroSeries.cov and MicroSeries.corr are frequency-weighted, so the same pair of columns gives a different answer depending on the entry point.

import numpy as np, pandas as pd, microdf as mdf

df = pd.DataFrame({"x": [1.0, 2.0, 3.0, 4.0], "y": [1.0, 4.0, 2.0, 8.0]})
w = np.array([1, 1, 1, 5])
m = mdf.MicroDataFrame(df, weights=w)
rep = df.loc[df.index.repeat(w)].reset_index(drop=True)  # frequency-replicated sample

m.cov().loc["x", "y"]     # 3.1667  == df.cov()  (unweighted)
m.x.cov(m.y)              # 3.1786  == rep.cov() (weighted)
m.corr().loc["x", "y"]    # 0.792   == df.corr() (unweighted)
rep.corr().loc["x", "y"]  # 0.896

Observed on 1.5.2 (main at de08443).

Expected. MicroDataFrame.cov() and .corr() return the pairwise frequency-weighted matrix, consistent with MicroSeries.cov / .corr and with the package convention that weights are frequency counts (sum(w) - ddof in the denominator), so integer weights agree with pandas on the replicated sample.

Context. Surfaced while reviewing #324, where the API reference described these two methods as frequency-weighted. The docs are being corrected there; this issue is the code side. Whichever way it goes, the reference rows for these two methods should state which behaviour they have.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions