feat: Add LumpFeatures transfomer - #1941
Conversation
|
Hey dylanw-oss 👋! We use semantic commit messages to streamline the release process. Examples of commit messages with semantic prefixes:
To test your commit locally, please follow our guild on building from source. |
|
if there is no objective for this feature, I'll add document. sarahshy Jason Wang (@memoryz) |
There was a problem hiding this comment.
Summary by GPT-4
The LumpFeatures transformer is a custom transformer that takes a DataFrame and a list of lumping rules as input and returns a DataFrame comprised of the original columns, but the columns defined in lumping rules will be indexed and lumped to top k. This transformer can be used to handle high cardinality skewed categorical features before doing encoding.
In the given code, the LumpFeatures class extends Transformer and implements the following methods:
-
transform: This method takes an input dataset and applies the lumping rules to it. It first creates a pipeline with StringIndexer transformers for each column specified in the lumping rules. Then, it fits and transforms the input dataset using this pipeline. Finally, it keeps only the top k levels for each categorical column according to the lumping rules. -
transformSchema: This method returns the schema of the output DataFrame after applying the transformation. -
copy: This method creates a copy of this instance with extra parameters.
The test suite LumpFeaturesSuite tests this transformer's basic functionality by creating an input DataFrame with categorical columns, applying lumping rules using an instance of LumpFeatures, and comparing the output DataFrame with an expected result.
In summary, this custom transformer helps in handling high cardinality skewed categorical features by indexing and lumping them according to specified rules before encoding them.
Suggestions
The changes in this PR look good and no suggestions are needed.
## Summary Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model. ## Prompting Intent Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests. ## Linked Sources - Pull request: microsoft#2596 - Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 ## Rationale Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models. ## Prompting Intent Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths. ## Linked Sources - Pull request: microsoft#2596 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 - Independent review finding supplied in the PR follow-up request - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests. ## Prompting Intent Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation. ## Linked Sources - Original proposal PR: microsoft#1941 - Feature request: microsoft#1891 - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model. ## Prompting Intent Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests. ## Linked Sources - Pull request: microsoft#2596 - Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 ## Rationale Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models. ## Prompting Intent Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths. ## Linked Sources - Pull request: microsoft#2596 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 - Independent review finding supplied in the PR follow-up request - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests. ## Prompting Intent Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation. ## Linked Sources - Original proposal PR: microsoft#1941 - Feature request: microsoft#1891 - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model. ## Prompting Intent Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests. ## Linked Sources - Pull request: microsoft#2596 - Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 ## Rationale Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models. ## Prompting Intent Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths. ## Linked Sources - Pull request: microsoft#2596 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 - Independent review finding supplied in the PR follow-up request - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist the fit-time fallback inside learned categorical state, prevent all direct and generic mutation bypasses, keep copy and load behavior atomic, and make nullability match runtime output. ## Prompting Intent Rebase PR microsoft#2596 onto current master and resolve review findings while keeping top-K scoring lean, deterministic, persisted, and compatible with Spark ML copy/load and generated language bindings. ## Linked Sources - Pull request: microsoft#2596 - Original categorical lumping proposal: microsoft#1941 - Related issue: microsoft#1891 - Nullability review: microsoft#2596 (comment) ## Rationale The reserved fallback is stored with the existing learned-state JSON rather than as a separately clearable Spark Param, so generated bindings expose no internal escape hatch. A compact in-memory snapshot restores and rejects generic mutations, copy extras are checked before transfer, and legacy artifacts are upgraded from their persisted fallback. Validation occurs once per schema/transform rather than per row. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 81d39bfc-927c-418a-90a8-e0f2cd8fc128
## Summary Add LumpFeatures as a Spark ML estimator with a persisted model, deterministic top-K learning, explicit other-bucket and null semantics, schema-safe transforms, generated bindings, and comprehensive tests. ## Prompting Intent Recreate the valuable proposal from GitHub PR microsoft#1941 for current SynapseML without fitting during transform. Preserve lumpRules compatibility while covering persistence, copy behavior, special column names, unseen values, collisions, and Scala-first Python code generation. ## Linked Sources - Original proposal PR: microsoft#1941 - Feature request: microsoft#1891 - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Learn category frequencies once in fit and persist only retained values so scoring is stable and side-effect free. Restrict v1 to string columns, rank ties by value, preserve nulls by default, and reject other-bucket collisions rather than silently merging real categories. Use Spark SQL expressions instead of UDFs and retain the multi-column lumpRules API. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Add dedicated transformer fuzzing coverage for LumpFeaturesModel so global experiment, serialization, Python, and R coverage gates recognize the persisted model. ## Prompting Intent Repair the concrete UnitTests core failure from PR microsoft#2596 after Azure build 229219360 reported that LumpFeaturesModel had no directly registered fuzzers, while preserving all estimator tests. ## Linked Sources - Pull request: microsoft#2596 - Failed Azure build: https://msdata.visualstudio.com/b9b2accc-2d1c-45b3-9d24-0eb5d78cc47f/_build/results?buildId=229219360 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 ## Rationale Register a real TransformerFuzzing test object instead of exempting the model. This exercises deterministic transforms and model persistence while generating Python and R correspondence coverage expected by the repository-wide FuzzingTest. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist each fitted top-K alongside retained values and reject incompatible LumpFeaturesModel lumpRules mutations across direct setters, generic generated-binding transfer, copy overrides, and loaded models. ## Prompting Intent Address the independent medium-severity API review finding on PR microsoft#2596 without removing API-compatible params. Ensure a fitted model can never silently score with learned values that disagree with a post-fit K, and cover persistence, copy, Scala, Java, JSON, and generated-binding paths. ## Linked Sources - Pull request: microsoft#2596 - Original proposal: microsoft#1941 - Feature request: microsoft#1891 - Independent review finding supplied in the PR follow-up request - No Azure DevOps work item was supplied; tracking is through the linked GitHub issue. ## Rationale Encode the fitted top-K inside the existing model-only keptValuesJson state instead of adding another generated mutable parameter. Direct model setters fail immediately, while transform-time state validation protects generic Param paths used by generated bindings. Exact no-op rule assignment remains allowed, and incompatible copy overrides fail before returning. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
## Summary Persist the fit-time fallback inside learned categorical state, prevent all direct and generic mutation bypasses, keep copy and load behavior atomic, and make nullability match runtime output. ## Prompting Intent Rebase PR microsoft#2596 onto current master and resolve review findings while keeping top-K scoring lean, deterministic, persisted, and compatible with Spark ML copy/load and generated language bindings. ## Linked Sources - Pull request: microsoft#2596 - Original categorical lumping proposal: microsoft#1941 - Related issue: microsoft#1891 - Nullability review: microsoft#2596 (comment) ## Rationale The reserved fallback is stored with the existing learned-state JSON rather than as a separately clearable Spark Param, so generated bindings expose no internal escape hatch. A compact in-memory snapshot restores and rejects generic mutations, copy extras are checked before transfer, and legacy artifacts are upgraded from their persisted fallback. Validation occurs once per schema/transform rather than per row. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 81d39bfc-927c-418a-90a8-e0f2cd8fc128
Related Issues/PRs
#1891
What changes are proposed in this pull request?
A transformer can be used to handle data with high cardinality skewed categorical before doing other featurization processing.
How is this patch tested?
unit test
Does this PR add a new feature? If so, have you added samples on website?
will add document in next commit (I'd like to ensure it makes sense before doing the next step)
website/docs/documentationfolder.Make sure you choose the correct class
estimators/transformersand namespace.DocTablepoints to correct API link.yarn run startto make sure the website renders correctly.<!--pytest-codeblocks:cont-->before each python code blocks to enable auto-tests for python samples.WebsiteSamplesTestsjob pass in the pipeline.