Skip to content

Add resumable DART100 dataset lifecycle evidence tooling - #3

Open
kimminhyun-ai wants to merge 6 commits into
mainfrom
year3/m6-eval-incentive
Open

Add resumable DART100 dataset lifecycle evidence tooling#3
kimminhyun-ai wants to merge 6 commits into
mainfrom
year3/m6-eval-incentive

Conversation

@kimminhyun-ai

Copy link
Copy Markdown
Contributor

Scope

Deployment-specific dataset → teach → compare-inference evidence harness. Verifies the immutable public HF dataset manifest against 100 registered IDs and canonical hashes; adopts existing jobs, persists submission intent, and stops on ambiguous outcomes instead of resubmitting.

Separates execution coverage from answer accuracy, includes every canonical row, hashes raw responses for resume, and refuses to unload unowned runtime patches. No automatic publication, consent/PII bypass, incentive payout, or claim that 100 patches are 100 distinct models.

Validation

Nine regression tests pass in Node 24 inside a read-only Docker container (2 CPU, 2 GiB, network disabled). A nested CommonJS package boundary fixes the initial ES-module import failure, whose log is retained.

Live deployment: existing first job reused without retraining; 16 compare calls complete with 5/8 primary and 1/8 alternate answers correct. Second dataset is training. All 100 lessons, HF model/service deployment and public Ainize listings remain incomplete. The live snapshot predates two extra resume/error-log guards documented in README.

Limitations

Fixed existing deployment paths and prerequisites; not a fresh-machine bootstrap. Source publication only, not an npm release or a merged/deployed production change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants