Add a small Random Forest baseline demonstration - #2
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add a runnable CPU-only Random Forest example for commitment-outcome classification, using scikit-learn in an optional pinned ML environment. The default run generates 300 explicitly fictional tabular snapshots, fits a fixed forest, and compares its probability estimates with the historical training fulfilment rate on a later test period.
An explicit feature allowlist excludes identities and outcome fields. Training labels must be available by the fit cutoff; unknown, disputed, censored, and unavailable labels are accounted for separately. The optional JSON report records generation provenance, split dates, model settings, exclusions, and scored test predictions. The generator's assumptions and limits are documented; it does not label the challenge cards or extract features from the event ledger.
The guide also positions TabFM and TabPFN as future tabular comparisons, with primary-source links. No PyTorch, GPU, provider API, model-weight download, real dataset, or empirical organisational claim is introduced.
Validation: 22 local tests passed, existing 11-event fixture validated, full demo completed, and documentation links/whitespace checked. A dedicated Python 3.14 ML CI job runs the optional dependencies and full suite; base event validation remains on Python 3.11 and 3.14.