Update leaderboard - #1
Merged
Merged
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The documented scoring workflow references missing files, and several validation, accessibility, date, and documentation inconsistencies remain.
Review effort: Balanced
Findings: 3
Open (7)
Handle missing CSV fields without crashing validation · New Use row-group semantics for model category headings · New Add missing submission scorer and example or fix workflow instructions · New Correct exact 50% claim for the finite baseline result · New Update README schema documentation for new metrics and readout · New Fix unavailable artifact links or include the referenced files · New Replace exact 50% claim with an approximate qualification · New
What changed in this PR
Expands the leaderboard with real benchmark results and FLIP/STAY metrics while updating the site and validation workflow.
Changes:
- Replaces placeholder results with model, human, and baseline scores.
- Adds FLIP/STAY metrics, grouping, sorting, and documentation.
- Updates schema validation and deployment configuration.
| File | Description |
|---|---|
leaderboard.csv |
Adds expanded schema and results. |
scripts/validate.py |
Validates new metrics and readout modes. |
scripts/build.py |
Updates anonymity leak terms. |
site/index.html |
Expands leaderboard columns and descriptions. |
site/about.html |
Documents dataset, metrics, and findings. |
site/submit.html |
Updates submission instructions and schema. |
site/assets/app.js |
Adds metric rendering, sorting, and grouping. |
site/assets/style.css |
Styles the expanded leaderboard. |
site/assets/config.js |
Updates repository URL and item count. |
site/assets/config.anon.js |
Configures the anonymous repository mirror. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if not 0 <= v <= 100: | ||
| raise ValueError | ||
| return v | ||
| except ValueError: |
Comment on lines
+126
to
+127
| tbody.innerHTML = groups.map(([title, g]) => | ||
| `<tr class="group"><th scope="colgroup" colspan="${COLS}">${esc(title)}</th></tr>` + g.map(rowHtml).join("")).join(""); |
Comment on lines
+32
to
+33
| <p>Answer the 3,000 test items of the v1.0 test split (1,000 clean, 1,000 FLIP and 1,000 STAY items) and write one JSON Lines file: a first line with the model name and version, how the answer was read (letter probabilities or generated text), whether the model used the training split, the date and an optional link to code, then one record per item with the item ID and the letter ‘A’ or ‘B’. Missing or invalid answers count as wrong. Training on test items is not allowed. The format is described at the top of <code>musiclistenbench/scoring/score_submission.py</code> in the <a data-link="githubRepo" href="#">GitHub repository</a>, and <code>data/example_submission.jsonl</code> is an example. Score the file with the official script, run from the repository root:</p> | ||
| <pre><code>python -m musiclistenbench.scoring.score_submission submission.jsonl</code></pre></li> |
| model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,SFT,60.8,54.6,71.4,51.8,59.65,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results | ||
| model,PLACEHOLDER closed model B,PLACEHOLDER org,https://example.com/,closed,,zero-shot,66.4,58.0,73.6,55.2,63.3,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results | ||
| type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes | ||
| baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,50.0,50.0,50.0,,—,yes,maintainers,2026-09-30,,1.0,"Guessing at random. A model that ignores the audio scores 50% on FLIP/STAY, whatever letter it prefers" |
Comment on lines
6
to
+8
| COLUMNS = ["type","model","organization","model_url","access","params_b","method", | ||
| "melody","rhythm","timbre","harmony","overall","verified","submitted_by", | ||
| "date","source_url","benchmark_version","notes"] | ||
| "melody","rhythm","timbre","harmony","overall","acc_fs","miss_rate","false_flip_rate","a_rate", | ||
| "readout","verified","submitted_by","date","source_url","benchmark_version","notes"] |
| <div class="prose"> | ||
| <p>The full dataset is on <a data-link="datasetUrl" href="#">Hugging Face</a>. A test split of 2,000 items (500 per attribute) is set aside for the leaderboard.</p> | ||
| <p><span class="todo">TODO: source of the musical material, clip length, sample rate, licence, and whether test labels are public or held out.</span></p> | ||
| <p>The benchmark has 10,000 training items (2,500 per task) and 3,000 test items: 1,000 clean items (250 per task) and their 2,000 FLIP/STAY variants. All audio is generated from symbolic music and rendered with FluidSynth and freely available soundfonts, so every answer is known exactly and no human labelling is needed. Test items use pitch ranges, a soundfont and sound effects that never appear in training. The item files, the generator and the scoring script are in the <a data-link="githubRepo" href="#">code repository</a>.</p> |
| <div> | ||
| <h1>Can audio LLMs hear the difference yet?</h1> | ||
| <p class="lede">Each test item is a pair of clips that differ in exactly one attribute: melody, rhythm, timbre or harmony. A model scores only if it hears that one change and ignores everything else.</p> | ||
| <p class="lede">Each test item is two short clips and one question about melody, rhythm, timbre or harmony. Every item comes in two variants: in FLIP the music changes and the answer switches, in STAY only the sound changes (room echo, EQ, hiss, transposition) and the answer stays. A model that ignores the audio scores 50%.</p> |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


New leaderboard result
leaderboard.csvand changed no other files<commit hash>source_urllinks to a paper, model card, or eval logsverifiedtono(maintainers change it after reproducing)Model:
Prompt / decoding settings:
Anything unusual about the setup: