Skip to content

Update leaderboard - #1

Merged
aroy1990-dev merged 1 commit into
mainfrom
leaderboard-only
Sep 29, 2026
Merged

aroy1990-dev merged 1 commit into
mainfrom
leaderboard-only

Conversation

@aroy1990-dev

Copy link
Copy Markdown
Contributor

New leaderboard result

  • I added one row per model + method to leaderboard.csv and changed no other files
  • Scores are percent accuracy on the MusicListenBench v1.0 test split (500 items per attribute)
  • I used the official evaluation script at commit: <commit hash>
  • source_url links to a paper, model card, or eval logs
  • I set verified to no (maintainers change it after reproducing)

Model:
Prompt / decoding settings:
Anything unusual about the setup:

Copilot AI balanced review requested due to automatic review settings September 29, 2026 17:05
@aroy1990-dev
aroy1990-dev merged commit b915b41 into main Sep 29, 2026
2 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The documented scoring workflow references missing files, and several validation, accessibility, date, and documentation inconsistencies remain.

Review effort: Balanced
Findings: 3 Medium severity · 4 Low severity

Open (7)
What changed in this PR

Expands the leaderboard with real benchmark results and FLIP/STAY metrics while updating the site and validation workflow.

Changes:

  • Replaces placeholder results with model, human, and baseline scores.
  • Adds FLIP/STAY metrics, grouping, sorting, and documentation.
  • Updates schema validation and deployment configuration.
File Description
leaderboard.csv Adds expanded schema and results.
scripts/​validate.py Validates new metrics and readout modes.
scripts/​build.py Updates anonymity leak terms.
site/​index.html Expands leaderboard columns and descriptions.
site/​about.html Documents dataset, metrics, and findings.
site/​submit.html Updates submission instructions and schema.
site/​assets/​app.js Adds metric rendering, sorting, and grouping.
site/​assets/​style.css Styles the expanded leaderboard.
site/​assets/​config.js Updates repository URL and item count.
site/​assets/​config.anon.js Configures the anonymous repository mirror.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/validate.py
if not 0 <= v <= 100:
raise ValueError
return v
except ValueError:
Comment thread site/assets/app.js
Comment on lines +126 to +127
tbody.innerHTML = groups.map(([title, g]) =>
`<tr class="group"><th scope="colgroup" colspan="${COLS}">${esc(title)}</th></tr>` + g.map(rowHtml).join("")).join("");
Comment thread site/submit.html
Comment on lines +32 to +33
<p>Answer the 3,000 test items of the v1.0 test split (1,000 clean, 1,000 FLIP and 1,000 STAY items) and write one JSON Lines file: a first line with the model name and version, how the answer was read (letter probabilities or generated text), whether the model used the training split, the date and an optional link to code, then one record per item with the item ID and the letter ‘A’ or ‘B’. Missing or invalid answers count as wrong. Training on test items is not allowed. The format is described at the top of <code>musiclistenbench/scoring/score_submission.py</code> in the <a data-link="githubRepo" href="#">GitHub repository</a>, and <code>data/example_submission.jsonl</code> is an example. Score the file with the official script, run from the repository root:</p>
<pre><code>python -m musiclistenbench.scoring.score_submission submission.jsonl</code></pre></li>
Comment thread leaderboard.csv
model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,SFT,60.8,54.6,71.4,51.8,59.65,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
model,PLACEHOLDER closed model B,PLACEHOLDER org,https://example.com/,closed,,zero-shot,66.4,58.0,73.6,55.2,63.3,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes
baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,50.0,50.0,50.0,,—,yes,maintainers,2026-09-30,,1.0,"Guessing at random. A model that ignores the audio scores 50% on FLIP/STAY, whatever letter it prefers"
Comment thread scripts/validate.py
Comment on lines 6 to +8
COLUMNS = ["type","model","organization","model_url","access","params_b","method",
"melody","rhythm","timbre","harmony","overall","verified","submitted_by",
"date","source_url","benchmark_version","notes"]
"melody","rhythm","timbre","harmony","overall","acc_fs","miss_rate","false_flip_rate","a_rate",
"readout","verified","submitted_by","date","source_url","benchmark_version","notes"]
Comment thread site/about.html
<div class="prose">
<p>The full dataset is on <a data-link="datasetUrl" href="#">Hugging Face</a>. A test split of 2,000 items (500 per attribute) is set aside for the leaderboard.</p>
<p><span class="todo">TODO: source of the musical material, clip length, sample rate, licence, and whether test labels are public or held out.</span></p>
<p>The benchmark has 10,000 training items (2,500 per task) and 3,000 test items: 1,000 clean items (250 per task) and their 2,000 FLIP/STAY variants. All audio is generated from symbolic music and rendered with FluidSynth and freely available soundfonts, so every answer is known exactly and no human labelling is needed. Test items use pitch ranges, a soundfont and sound effects that never appear in training. The item files, the generator and the scoring script are in the <a data-link="githubRepo" href="#">code repository</a>.</p>
Comment thread site/index.html
<div>
<h1>Can audio LLMs hear the difference yet?</h1>
<p class="lede">Each test item is a pair of clips that differ in exactly one attribute: melody, rhythm, timbre or harmony. A model scores only if it hears that one change and ignores everything else.</p>
<p class="lede">Each test item is two short clips and one question about melody, rhythm, timbre or harmony. Every item comes in two variants: in FLIP the music changes and the answer switches, in STAY only the sound changes (room echo, EQ, hiss, transposition) and the answer stays. A model that ignores the audio scores 50%.</p>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants