Skip to content

Preserve MicroSeries weights in np.maximum and plain-Series-left divmod #322

Description

@juaristi22

Two binary-operation paths silently lose observation weights. np.maximum returns a MicroSeries with unit weights, and divmod returns plain pandas Series when its left operand is a plain Series and its right operand is a MicroSeries. Subsequent sums therefore produce unweighted totals.

These bugs already occur at base commit ce1ad0b and remain at PR #321 commit b779c4b. They are pre-existing limitations, not regressions introduced by #320 or #321. Each reproduction below gives identical results with pandas 2.3.1 and 3.0.5 at both commits.

Reproduction

Run this with microdf, pandas and NumPy installed:

import numpy as np
import pandas as pd
from microdf import MicroSeries

s = MicroSeries([10, 20], weights=[2, 9])

maximum = np.maximum(s, pd.Series([12, 10]))
quotient, remainder = divmod(pd.Series([23, 41]), s)

for name, result in [
    ("maximum", maximum),
    ("quotient", quotient),
    ("remainder", remainder),
]:
    print(
        name,
        type(result).__name__,
        result.tolist(),
        result.weights.tolist() if isinstance(result, MicroSeries) else None,
        float(result.sum()),
    )

Actual output:

maximum MicroSeries [12, 20] [1.0, 1.0] 32.0
quotient Series [2, 2] None 4.0
remainder Series [3, 1] None 4.0

Expected behavior

All three results represent the same two observations as s, so each should retain its weights [2.0, 9.0] in an independently mutable MicroSeries.

Result Values Expected weighted sum Actual sum
Maximum [12, 20] 12 * 2 + 20 * 9 = 204 32
Quotient [2, 2] 2 * 2 + 2 * 9 = 22 4
Remainder [3, 1] 3 * 2 + 1 * 9 = 15 4

The elementwise values are correct; the lost weights change the reductions. These examples use matching indexes and only one weighted operand, so they do not involve ambiguous row weights or conflicting input weight vectors.

Acceptance criteria

  • Preserve the weighted operand's observation weights for np.maximum(MicroSeries, Series) and for both results of divmod(Series, MicroSeries).
  • Preserve pandas' elementwise values, index alignment and error behavior, while retaining microdf's existing rejection of unknown or ambiguous row weights.
  • Keep result weights independently mutable without changing the input weights.
  • Add regression tests for the examples above on supported pandas 2 and 3 versions, including label-aligned operands and both members of the divmod tuple.

This issue tracks the two dispatch paths above. The mismatched-input-weight warning requested in #170 remains a separate concern.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions