Skip to content

feat: Add LumpFeatures transfomer - #1941

Open
dylanw-oss wants to merge 3 commits into
microsoft:masterfrom
dylanw-oss:lump
Open

feat: Add LumpFeatures transfomer#1941
dylanw-oss wants to merge 3 commits into
microsoft:masterfrom
dylanw-oss:lump

Conversation

@dylanw-oss

Copy link
Copy Markdown
Contributor

Related Issues/PRs

#1891

What changes are proposed in this pull request?

A transformer can be used to handle data with high cardinality skewed categorical before doing other featurization processing.

How is this patch tested?

unit test

Does this PR add a new feature? If so, have you added samples on website?

will add document in next commit (I'd like to ensure it makes sense before doing the next step)

  1. Find the corresponding markdown file for your new feature in website/docs/documentation folder.
    Make sure you choose the correct class estimators/transformers and namespace.
  2. Follow the pattern in markdown file and add another section for your new API, including pyspark, scala (and .NET potentially) samples.
  3. Make sure the DocTable points to correct API link.
  4. Navigate to website folder, and run yarn run start to make sure the website renders correctly.
  5. Don't forget to add <!--pytest-codeblocks:cont--> before each python code blocks to enable auto-tests for python samples.
  6. Make sure the WebsiteSamplesTests job pass in the pipeline.

@github-actions

Copy link
Copy Markdown

Hey dylanw-oss 👋!
Thank you so much for contributing to our repository 🙌.
Someone from SynapseML Team will be reviewing this pull request soon.

We use semantic commit messages to streamline the release process.
Before your pull request can be merged, you should make sure your first commit and PR title start with a semantic prefix.
This helps us to create release messages and credit you for your hard work!

Examples of commit messages with semantic prefixes:

  • fix: Fix LightGBM crashes with empty partitions
  • feat: Make HTTP on Spark back-offs configurable
  • docs: Update Spark Serving usage
  • build: Add codecov support
  • perf: improve LightGBM memory usage
  • refactor: make python code generation rely on classes
  • style: Remove nulls from CNTKModel
  • test: Add test coverage for CNTKModel

To test your commit locally, please follow our guild on building from source.
Check out the developer guide for additional guidance on testing your change.

@dylanw-oss dylanw-oss changed the title Feat, Add LumpFeatures transfomer feat, Add LumpFeatures transfomer Apr 26, 2023
@dciborow Daniel Ciborowski (dciborow) changed the title feat, Add LumpFeatures transfomer feat: Add LumpFeatures transfomer Apr 27, 2023
@dylanw-oss

Copy link
Copy Markdown
Contributor Author

if there is no objective for this feature, I'll add document. sarahshy Jason Wang (@memoryz)

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary by GPT-4

The LumpFeatures transformer is a custom transformer that takes a DataFrame and a list of lumping rules as input and returns a DataFrame comprised of the original columns, but the columns defined in lumping rules will be indexed and lumped to top k. This transformer can be used to handle high cardinality skewed categorical features before doing encoding.

In the given code, the LumpFeatures class extends Transformer and implements the following methods:

  1. transform: This method takes an input dataset and applies the lumping rules to it. It first creates a pipeline with StringIndexer transformers for each column specified in the lumping rules. Then, it fits and transforms the input dataset using this pipeline. Finally, it keeps only the top k levels for each categorical column according to the lumping rules.

  2. transformSchema: This method returns the schema of the output DataFrame after applying the transformation.

  3. copy: This method creates a copy of this instance with extra parameters.

The test suite LumpFeaturesSuite tests this transformer's basic functionality by creating an input DataFrame with categorical columns, applying lumping rules using an instance of LumpFeatures, and comparing the output DataFrame with an expected result.

In summary, this custom transformer helps in handling high cardinality skewed categorical features by indexing and lumping them according to specified rules before encoding them.

Suggestions

The changes in this PR look good and no suggestions are needed.

Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 1, 2026
## Summary
Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model.

## Prompting Intent
Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests.

## Linked Sources
- Pull request: microsoft#2596
- Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360
- Original proposal: microsoft#1941
- Feature request: microsoft#1891

## Rationale
Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 1, 2026
## Summary
Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models.

## Prompting Intent
Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths.

## Linked Sources
- Pull request: microsoft#2596
- Original proposal: microsoft#1941
- Feature request: microsoft#1891
- Independent review finding supplied in the PR follow-up request
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 4, 2026
## Summary
Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests.

## Prompting Intent
Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation.

## Linked Sources
- Original proposal PR: microsoft#1941
- Feature request: microsoft#1891
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 4, 2026
## Summary
Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model.

## Prompting Intent
Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests.

## Linked Sources
- Pull request: microsoft#2596
- Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360
- Original proposal: microsoft#1941
- Feature request: microsoft#1891

## Rationale
Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 4, 2026
## Summary
Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models.

## Prompting Intent
Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths.

## Linked Sources
- Pull request: microsoft#2596
- Original proposal: microsoft#1941
- Feature request: microsoft#1891
- Independent review finding supplied in the PR follow-up request
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests.

## Prompting Intent
Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation.

## Linked Sources
- Original proposal PR: microsoft#1941
- Feature request: microsoft#1891
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model.

## Prompting Intent
Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests.

## Linked Sources
- Pull request: microsoft#2596
- Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360
- Original proposal: microsoft#1941
- Feature request: microsoft#1891

## Rationale
Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models.

## Prompting Intent
Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths.

## Linked Sources
- Pull request: microsoft#2596
- Original proposal: microsoft#1941
- Feature request: microsoft#1891
- Independent review finding supplied in the PR follow-up request
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) pushed a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Persist the fit-time fallback inside learned categorical state, prevent all direct and generic mutation bypasses, keep copy and load behavior atomic, and make nullability match runtime output.

## Prompting Intent
Rebase PR microsoft#2596 onto current master and resolve review findings while keeping top-K scoring lean, deterministic, persisted, and compatible with Spark ML copy/load and generated language bindings.

## Linked Sources
- Pull request: microsoft#2596
- Original categorical lumping proposal: microsoft#1941
- Related issue: microsoft#1891
- Nullability review: microsoft#2596 (comment)

## Rationale
The reserved fallback is stored with the existing learned-state JSON rather than as a separately clearable Spark Param, so generated bindings expose no internal escape hatch. A compact in-memory snapshot restores and rejects generic mutations, copy extras are checked before transfer, and legacy artifacts are upgraded from their persisted fallback. Validation occurs once per schema/transform rather than per row.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 81d39bfc-927c-418a-90a8-e0f2cd8fc128
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests.

## Prompting Intent
Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation.

## Linked Sources
- Original proposal PR: microsoft#1941
- Feature request: microsoft#1891
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model.

## Prompting Intent
Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests.

## Linked Sources
- Pull request: microsoft#2596
- Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360
- Original proposal: microsoft#1941
- Feature request: microsoft#1891

## Rationale
Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) added a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models.

## Prompting Intent
Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths.

## Linked Sources
- Pull request: microsoft#2596
- Original proposal: microsoft#1941
- Feature request: microsoft#1891
- Independent review finding supplied in the PR follow-up request
- No Azure DevOps work item was supplied; tracking is through the linked GitHub issue.

## Rationale
Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rana Singh (ranadeepsingh) pushed a commit to ranadeepsingh/SynapseML that referenced this pull request Aug 7, 2026
## Summary
Persist the fit-time fallback inside learned categorical state, prevent all direct and generic mutation bypasses, keep copy and load behavior atomic, and make nullability match runtime output.

## Prompting Intent
Rebase PR microsoft#2596 onto current master and resolve review findings while keeping top-K scoring lean, deterministic, persisted, and compatible with Spark ML copy/load and generated language bindings.

## Linked Sources
- Pull request: microsoft#2596
- Original categorical lumping proposal: microsoft#1941
- Related issue: microsoft#1891
- Nullability review: microsoft#2596 (comment)

## Rationale
The reserved fallback is stored with the existing learned-state JSON rather than as a separately clearable Spark Param, so generated bindings expose no internal escape hatch. A compact in-memory snapshot restores and rejects generic mutations, copy extras are checked before transfer, and legacy artifacts are upgraded from their persisted fallback. Validation occurs once per schema/transform rather than per row.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 81d39bfc-927c-418a-90a8-e0f2cd8fc128
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants