MicroDataFrame.cov() and MicroDataFrame.corr() return plain pandas results. The overrides at microdf/microdataframe.py:78 and :84 call pd.DataFrame(self, copy=False).cov(*args, **kwargs) / .corr(...) and use __finalize__ only to keep weights off the column-by-column result. MicroSeries.cov and MicroSeries.corr are frequency-weighted, so the same pair of columns gives a different answer depending on the entry point.
import numpy as np, pandas as pd, microdf as mdf
df = pd.DataFrame({"x": [1.0, 2.0, 3.0, 4.0], "y": [1.0, 4.0, 2.0, 8.0]})
w = np.array([1, 1, 1, 5])
m = mdf.MicroDataFrame(df, weights=w)
rep = df.loc[df.index.repeat(w)].reset_index(drop=True) # frequency-replicated sample
m.cov().loc["x", "y"] # 3.1667 == df.cov() (unweighted)
m.x.cov(m.y) # 3.1786 == rep.cov() (weighted)
m.corr().loc["x", "y"] # 0.792 == df.corr() (unweighted)
rep.corr().loc["x", "y"] # 0.896
Observed on 1.5.2 (main at de08443).
Expected. MicroDataFrame.cov() and .corr() return the pairwise frequency-weighted matrix, consistent with MicroSeries.cov / .corr and with the package convention that weights are frequency counts (sum(w) - ddof in the denominator), so integer weights agree with pandas on the replicated sample.
Context. Surfaced while reviewing #324, where the API reference described these two methods as frequency-weighted. The docs are being corrected there; this issue is the code side. Whichever way it goes, the reference rows for these two methods should state which behaviour they have.
MicroDataFrame.cov()andMicroDataFrame.corr()return plain pandas results. The overrides atmicrodf/microdataframe.py:78and:84callpd.DataFrame(self, copy=False).cov(*args, **kwargs)/.corr(...)and use__finalize__only to keep weights off the column-by-column result.MicroSeries.covandMicroSeries.corrare frequency-weighted, so the same pair of columns gives a different answer depending on the entry point.Observed on 1.5.2 (main at de08443).
Expected.
MicroDataFrame.cov()and.corr()return the pairwise frequency-weighted matrix, consistent withMicroSeries.cov/.corrand with the package convention that weights are frequency counts (sum(w) - ddofin the denominator), so integer weights agree with pandas on the replicated sample.Context. Surfaced while reviewing #324, where the API reference described these two methods as frequency-weighted. The docs are being corrected there; this issue is the code side. Whichever way it goes, the reference rows for these two methods should state which behaviour they have.