Skip to content

feat(scripts): Wilson CI + paired McNemar reporting - #231

Open
KeilerHirsch wants to merge 2 commits into
THUDM:mainfrom
KeilerHirsch:ci/wilson-ci-reporting
Open

feat(scripts): Wilson CI + paired McNemar reporting#231
KeilerHirsch wants to merge 2 commits into
THUDM:mainfrom
KeilerHirsch:ci/wilson-ci-reporting

Conversation

@KeilerHirsch

Copy link
Copy Markdown

Summary

Adds scripts/ci_report.py: a stdlib-only tool that reports aggregate pass rates with 95% Wilson score intervals and, with --compare, an exact paired McNemar test between two agents over their shared task set.

Input is a per-instance results JSON (mapping {task_id: bool} or a list of {"id", "success"} records), which can be produced from AgentBench run outputs.

Motivation

On fixed item sets, strict leaderboard ranks overstate differences: every aggregate should carry an uncertainty measure, and the paired McNemar test answers whether a #1-vs-#k gap is real. This is the "measurement before ranking" tooling for AgentBench-style evaluations.

Testing

Smoke-tested locally against two small result files.

…results

Add a stdlib-only script that reports aggregate pass rates with 95% Wilson
score intervals and, with --compare, an exact paired McNemar test between
two agents over their shared task set.
… format

The tool prints aggregate pass rates plus an optional paired comparison,
not per-task rates. Also clarify that the input is a per-instance JSON
derived from runs, not native overall.json/db_out_new.jsonl artifacts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant