From 6571e57c098bc701b6641bb509f51c762b304111 Mon Sep 17 00:00:00 2001 From: Abhinaba Roy Date: Wed, 30 Sep 2026 00:59:59 +0800 Subject: [PATCH] Update leaderboard --- leaderboard.csv | 30 +++++++++++++---- scripts/build.py | 2 +- scripts/validate.py | 48 +++++++++++++++++---------- site/about.html | 26 +++++++++------ site/assets/app.js | 66 ++++++++++++++++++++++++++------------ site/assets/config.anon.js | 8 ++--- site/assets/config.js | 4 +-- site/assets/style.css | 7 ++++ site/index.html | 14 +++++--- site/submit.html | 12 +++---- 10 files changed, 146 insertions(+), 71 deletions(-) diff --git a/leaderboard.csv b/leaderboard.csv index 1035853..a8f31c2 100644 --- a/leaderboard.csv +++ b/leaderboard.csv @@ -1,6 +1,24 @@ -type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,verified,submitted_by,date,source_url,benchmark_version,notes -baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,yes,maintainers,2026-09-24,,1.0,TODO confirm chance level (assumes two answer options) -model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,zero-shot,61.2,55.4,70.8,52.6,60.0,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results -model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,GRPO,74.0,63.8,78.2,60.4,69.1,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results -model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,SFT,60.8,54.6,71.4,51.8,59.65,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results -model,PLACEHOLDER closed model B,PLACEHOLDER org,https://example.com/,closed,,zero-shot,66.4,58.0,73.6,55.2,63.3,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results +type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes +baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,50.0,50.0,50.0,,—,yes,maintainers,2026-09-30,,1.0,"Guessing at random. A model that ignores the audio scores 50% on FLIP/STAY, whatever letter it prefers" +baseline,Always answer A,—,,—,,—,54.0,50.8,56.8,49.6,52.8,49.5,53.0,48.1,100.0,—,yes,maintainers,2026-09-30,,1.0,Ignores the audio and always answers A +baseline,Human listeners (12),—,,—,,—,79.2,87.5,95.8,87.5,87.5,87.0,5.2,20.8,,—,yes,maintainers,2026-09-30,,1.0,"Paper Table 2. Twelve listeners, 24 items each: 24 clean answers per task, 96 FLIP and 96 STAY answers" +model,Qwen2-Audio,Alibaba,https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct,open,8.4,zero-shot,46.0,49.2,43.2,50.4,47.2,50.7,47.7,50.9,1.4,logprob,yes,maintainers,2026-09-30,,1.0, +model,Qwen2.5-Omni,Alibaba,https://huggingface.co/Qwen/Qwen2.5-Omni-7B,open,10.7,zero-shot,68.8,50.0,66.8,59.2,61.2,61.2,36.7,41.0,52.4,logprob,yes,maintainers,2026-09-30,,1.0, +model,Audio Flamingo 3,NVIDIA,https://huggingface.co/nvidia/audio-flamingo-3-hf,open,16.5,zero-shot,51.2,52.8,58.0,48.8,52.7,51.4,49.6,47.6,87.7,logprob,yes,maintainers,2026-09-30,,1.0, +model,Phi-4-multimodal,Microsoft,https://huggingface.co/microsoft/Phi-4-multimodal-instruct,open,5.6,zero-shot,54.4,50.8,56.8,47.2,52.3,49.1,53.5,48.3,95.9,logprob,yes,maintainers,2026-09-30,,1.0, +model,Kimi-Audio,Moonshot AI,https://huggingface.co/moonshotai/Kimi-Audio-7B-Instruct,open,7.9,zero-shot,57.6,52.8,53.2,47.2,52.7,51.0,52.3,45.7,90.1,logprob,yes,maintainers,2026-09-30,,1.0, +model,MiMo-Audio,Xiaomi,https://huggingface.co/XiaomiMiMo/MiMo-Audio-7B-Instruct,open,8.0,zero-shot,54.0,59.6,58.0,54.4,56.5,52.3,49.4,46.1,91.3,logprob,yes,maintainers,2026-09-30,,1.0, +model,Fun-Audio-Chat,FunAudioLLM,https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B,open,9.5,zero-shot,50.0,49.6,60.0,51.6,52.8,53.7,46.6,46.1,62.2,logprob,yes,maintainers,2026-09-30,,1.0, +model,Step-Audio 2,StepFun,https://huggingface.co/stepfun-ai/Step-Audio-2-mini,open,8.3,zero-shot,52.4,52.4,62.8,50.0,54.4,52.7,49.1,45.5,69.4,logprob,yes,maintainers,2026-09-30,,1.0, +model,GPT-Audio-mini,OpenAI,,closed,,zero-shot,47.6,50.8,54.4,51.6,51.1,49.6,50.6,50.3,81.1,generated,yes,maintainers,2026-09-30,,1.0, +model,GPT-Audio-1.5,OpenAI,,closed,,zero-shot,45.6,44.0,52.0,47.6,47.3,52.1,45.8,50.1,18.3,generated,yes,maintainers,2026-09-30,,1.0, +model,Gemini 2.5 Pro,Google,,closed,,zero-shot,80.0,79.6,88.4,68.8,79.2,74.2,24.7,26.9,60.8,generated,yes,maintainers,2026-09-30,,1.0, +model,Gemini 3.8 Flash,Google,,closed,,zero-shot,53.6,72.4,85.6,54.4,66.5,64.0,33.0,39.1,43.3,generated,yes,maintainers,2026-09-30,,1.0, +model,Qwen2-Audio,This paper,,open,8.4,GRPO (LoRA),70.8,51.6,92.4,60.8,68.9,63.8,29.7,42.7,31.3,logprob,yes,maintainers,2026-09-30,,1.0, +model,Qwen2-Audio,This paper,,open,8.4,GRPO (full),70.4,52.0,80.8,61.6,66.2,61.5,31.7,45.4,23.0,logprob,yes,maintainers,2026-09-30,,1.0, +model,Qwen2.5-Omni,This paper,,open,10.7,GRPO (LoRA),99.6,100.0,99.6,96.8,99.0,86.0,1.9,26.1,52.2,logprob,yes,maintainers,2026-09-30,,1.0, +model,Qwen2.5-Omni,This paper,,open,10.7,GRPO (full),100.0,99.6,100.0,85.2,96.2,83.5,4.2,28.9,49.0,logprob,yes,maintainers,2026-09-30,,1.0, +model,Audio Flamingo 3,This paper,,open,16.5,GRPO (LoRA),57.2,50.8,97.6,74.8,70.1,64.5,29.7,41.4,61.1,logprob,yes,maintainers,2026-09-30,,1.0, +model,Audio Flamingo 3,This paper,,open,16.5,GRPO (full),68.0,69.2,99.6,80.8,79.4,70.4,22.5,36.7,43.2,logprob,yes,maintainers,2026-09-30,,1.0,"Stopped at step 500, where validation accuracy peaks; the other GRPO runs use step 1000" +model,Phi-4-multimodal,This paper,,open,5.6,GRPO (LoRA),52.0,50.8,56.8,49.6,52.3,49.1,53.4,48.4,99.5,logprob,yes,maintainers,2026-09-30,,1.0, +model,Phi-4-multimodal,This paper,,open,5.6,GRPO (full),55.6,50.8,88.8,54.0,62.3,58.9,42.9,39.3,79.3,logprob,yes,maintainers,2026-09-30,,1.0, diff --git a/scripts/build.py b/scripts/build.py index b8557f5..31bba5e 100644 --- a/scripts/build.py +++ b/scripts/build.py @@ -19,7 +19,7 @@ # TODO: add every author's surname, and any other giveaway (grant names, emails). LEAK_TERMS = [ "AMAAI", "amaai-lab", "SUTD", "Singapore University", - "sleeping-ai", # style reference, not needed in either build + "roy", "heremmans", "mehrish", # style reference, not needed in either build ] ANON_BRAND = "Anonymous submission" diff --git a/scripts/validate.py b/scripts/validate.py index 6cd0e08..c5d1dbc 100644 --- a/scripts/validate.py +++ b/scripts/validate.py @@ -4,21 +4,36 @@ from datetime import date COLUMNS = ["type","model","organization","model_url","access","params_b","method", - "melody","rhythm","timbre","harmony","overall","verified","submitted_by", - "date","source_url","benchmark_version","notes"] + "melody","rhythm","timbre","harmony","overall","acc_fs","miss_rate","false_flip_rate","a_rate", + "readout","verified","submitted_by","date","source_url","benchmark_version","notes"] ASPECTS = ["melody","rhythm","timbre","harmony"] ENUMS = {"type": {"model","baseline"}, "verified": {"yes","no"}} MODEL_ACCESS = {"open","closed"} +MODEL_READOUT = {"logprob","generated"} KNOWN_VERSIONS = {"1.0"} -# TODO(maintainers): confirm "overall" is the macro-average of the four aspects. -# Set to False if overall is scored as its own task. -OVERALL_IS_MEAN = True +# The test split has 250 clean items per task, so "overall" (clean accuracy on all 1,000 items) +# equals the mean of the four task scores. +# acc_fs = 100 - (miss_rate + false_flip_rate) / 2 (FLIP/STAY accuracy, paper Section 3.3). +# Scores are rounded to one decimal, so allow a small tolerance. TOLERANCE = 0.1 errors = [] def err(line, msg): errors.append(f"::error file=leaderboard.csv,line={line}::{msg}") +def number(row, col, line, required=True): + """Return the value of a percent column, or None (and report) if it is missing or out of range.""" + if row[col] == "" and not required: + return None + try: + v = float(row[col]) + if not 0 <= v <= 100: + raise ValueError + return v + except ValueError: + err(line, f"{col} must be a number between 0 and 100 (percent)") + return None + def main(path): with open(path, newline="", encoding="utf-8") as f: reader = csv.DictReader(f) @@ -32,24 +47,23 @@ def main(path): err(i, f"{col} must be one of {sorted(allowed)}, got '{row[col]}'") if not row["model"].strip(): err(i, "model is required") - scores = {} - for a in ASPECTS + ["overall"]: - try: - v = float(row[a]) - if not 0 <= v <= 100: - raise ValueError - scores[a] = v - except ValueError: - err(i, f"{a} must be a number between 0 and 100 (percent accuracy)") - if OVERALL_IS_MEAN and len(scores) == 5: + scores = {a: number(row, a, i) for a in ASPECTS + ["overall", "acc_fs", "miss_rate", "false_flip_rate"]} + number(row, "a_rate", i, required=False) # share of 'A' answers on the clean items; empty if unknown + if None not in (scores[a] for a in ASPECTS + ["overall"]): mean = sum(scores[a] for a in ASPECTS) / 4 if abs(mean - scores["overall"]) > TOLERANCE: - err(i, f"overall ({scores['overall']}) should equal the mean of the four aspects ({mean:.2f})") + err(i, f"overall ({scores['overall']}) should equal the mean of the four tasks ({mean:.2f})") + if None not in (scores["acc_fs"], scores["miss_rate"], scores["false_flip_rate"]): + fs = 100 - (scores["miss_rate"] + scores["false_flip_rate"]) / 2 + if abs(fs - scores["acc_fs"]) > TOLERANCE: + err(i, f"acc_fs ({scores['acc_fs']}) should equal 100 - (miss_rate + false_flip_rate) / 2 = {fs:.2f}") if row["type"] == "model": if row["access"] not in MODEL_ACCESS: err(i, "access must be 'open' or 'closed'") + if row["readout"] not in MODEL_READOUT: + err(i, "readout must be 'logprob' (letter probabilities) or 'generated' (generated text)") if not row["method"].strip(): - err(i, "method is required (e.g. zero-shot, SFT, GRPO)") + err(i, "method is required: 'zero-shot', or how the released training split was used (e.g. GRPO (LoRA))") if row["params_b"] and not re.fullmatch(r"\d+(\.\d+)?", row["params_b"]): err(i, "params_b must be a number in billions, or empty if unknown") if not row["source_url"] and row["verified"] == "no": diff --git a/site/about.html b/site/about.html index 9b45e0a..6227db8 100644 --- a/site/about.html +++ b/site/about.html @@ -28,31 +28,37 @@

What the benchmark tests

Audio large language models can caption a track and answer broad questions about it, which suggests they understand music. MusicListenBench tests whether that understanding holds at the level of individual musical attributes.

-

Every item asks one question: given two clips, are they the same or different in a named attribute? Each pair changes exactly one attribute and holds the rest fixed, so a correct answer requires hearing that attribute and ignoring the others.

+

Each item is one audio file: clip A, one second of silence, and clip B. One question follows, and the model answers with a single letter. Melody, harmony and timbre ask whether the two clips are the same; rhythm asks which clip is faster. When the clips differ, they differ in exactly the queried attribute.

+
+

FLIP and STAY

+
+

For each clean test item there are two more variants, made by replacing clip B. In FLIP, clip B is an edited clip that changes the queried attribute, so the correct answer switches. In STAY, clip B sounds different but keeps the attribute: room echo, a mild EQ change, background hiss, or a transposition of up to 200 cents. The correct answer stays.

+

The two variants of an item have opposite answers. A model that ignores the audio therefore gets exactly one of them right and scores 50%, whatever letter it prefers (on rhythm this holds on average). The miss rate (MR) is the share of FLIP items answered wrongly; the false-flip rate (FFR) is the share of STAY items answered wrongly. The FLIP/STAY accuracy is AccFS = 100 − (MR + FFR) / 2.

+
+
+

Data

-

The full dataset is on Hugging Face. A test split of 2,000 items (500 per attribute) is set aside for the leaderboard.

-

TODO: source of the musical material, clip length, sample rate, licence, and whether test labels are public or held out.

+

The benchmark has 10,000 training items (2,500 per task) and 3,000 test items: 1,000 clean items (250 per task) and their 2,000 FLIP/STAY variants. All audio is generated from symbolic music and rendered with FluidSynth and freely available soundfonts, so every answer is known exactly and no human labelling is needed. Test items use pitch ranges, a soundfont and sound effects that never appear in training. The item files, the generator and the scoring script are in the code repository.

Evaluation

-

Models answer a multiple-choice question for each pair. We report accuracy per attribute and overall.

-

TODO: exact prompt template, answer options, answer-extraction rule, how invalid answers are scored, and the chance level.

+

Each question ends with a fixed instruction that says which letter means which answer. Open models are scored by comparing the probabilities of the tokens ‘A’ and ‘B’; API models are scored by the letter they generate. Missing or invalid answers count as wrong. Zero-shot models and models that used the training split are listed separately.

Findings from the paper

-

We evaluate open and closed-source audio LLMs, and show that GRPO post-training improves scores while supervised fine-tuning does not. TODO: one or two headline numbers.

+

Without training, seven of eight open audio LLMs score within 4 points of the 50% floor on FLIP/STAY. The best commercial model, Gemini 2.5 Pro, reaches 74.2%, and human listeners reach 87.0%. Post-training Qwen2.5-Omni with GRPO on the training split raises its clean accuracy from 61.2% to 99.0%, but its FLIP/STAY accuracy only reaches 86.0%: it misses almost no change (1.9%) and still false-flips on 26.1% of STAY items.

diff --git a/site/assets/app.js b/site/assets/app.js index ae54f60..0a89315 100644 --- a/site/assets/app.js +++ b/site/assets/app.js @@ -1,6 +1,8 @@ (() => { const C = window.MLB_CONFIG; const ASPECTS = ["melody", "rhythm", "timbre", "harmony"]; + const NUMERIC = ASPECTS.concat(["overall", "acc_fs", "miss_rate", "false_flip_rate", "a_rate"]); + const LOWER_IS_BETTER = ["miss_rate", "false_flip_rate"]; /* ---------- Config-driven links and footer ---------- */ document.querySelectorAll("[data-link]").forEach(a => { @@ -44,14 +46,14 @@ /* ---------- Leaderboard ---------- */ const board = document.getElementById("board"); if (board) { - const state = { rows: [], sort: "overall", dir: -1, q: "", access: "all", verifiedOnly: false, chance: 50 }; + const state = { rows: [], sort: "acc_fs", dir: -1, q: "", access: "all", verifiedOnly: false, chance: 50 }; const tbody = board.querySelector("tbody"); fetch(C.csvPath, { cache: "no-cache" }) .then(r => { if (!r.ok) throw new Error(r.status); return r.text(); }) .then(t => { state.rows = parseCSV(t).map(r => { - ASPECTS.concat("overall").forEach(k => r[k] = parseFloat(r[k])); + NUMERIC.forEach(k => r[k] = parseFloat(r[k])); return r; }); const base = state.rows.find(r => r.type === "baseline" && /chance/i.test(r.model)); @@ -63,48 +65,71 @@ render(); }) .catch(() => { - tbody.innerHTML = `Couldn't load ${esc(C.csvPath)}. If you opened this file directly from disk, run python -m http.server in the site folder instead.`; + tbody.innerHTML = `Couldn't load ${esc(C.csvPath)}. If you opened this file directly from disk, run python -m http.server in the site folder instead.`; }); + const COLS = 12; + const fmt = v => Number.isNaN(v) ? "—" : v.toFixed(1); + const isTrained = r => !/^zero-shot$/i.test(r.method); + function render() { const q = state.q.toLowerCase(); - let rows = state.rows.filter(r => + const rows = state.rows.filter(r => (r.type === "baseline" || state.access === "all" || r.access === state.access) && (r.type === "baseline" || !state.verifiedOnly || r.verified === "yes") && (!q || `${r.model} ${r.organization} ${r.method}`.toLowerCase().includes(q))); const key = state.sort; rows.sort((a, b) => { const av = a[key], bv = b[key]; - return typeof av === "number" ? (av - bv) * state.dir : String(av).localeCompare(String(bv)) * state.dir; + if (typeof av === "number") { + if (Number.isNaN(av) || Number.isNaN(bv)) return Number.isNaN(av) - Number.isNaN(bv); // empty values last + return (av - bv) * state.dir; + } + return String(av).localeCompare(String(bv)) * state.dir; }); - let rank = 0; - const byOverall = [...rows].filter(r => r.type === "model").sort((a, b) => b.overall - a.overall); - const rankOf = new Map(byOverall.map(r => [r, ++rank])); + // Zero-shot models and models trained on the released training split are listed separately. + const groups = [ + ["Zero-shot models", rows.filter(r => r.type === "model" && !isTrained(r))], + ["Models trained on the released training split", rows.filter(r => r.type === "model" && isTrained(r))], + ["References", rows.filter(r => r.type === "baseline")], + ].filter(g => g[1].length); + + if (!groups.length) { tbody.innerHTML = `No results match these filters. Clear the search or show all models.`; return; } - if (!rows.length) { tbody.innerHTML = `No results match these filters. Clear the search or show all models.`; return; } + // Rank inside each group by FLIP/STAY accuracy. + const rankOf = new Map(); + groups.filter(g => g[1][0].type === "model").forEach(([, g]) => + [...g].sort((a, b) => b.acc_fs - a.acc_fs).forEach((r, i) => rankOf.set(r, i + 1))); - tbody.innerHTML = rows.map(r => { + const rowHtml = r => { const isBase = r.type === "baseline"; - const cell = (k, cls = "num") => ` -
${r[k].toFixed(1)}
`; + const bar = (k, cls = "num") => ` +
${fmt(r[k])}
`; + const plain = k => `${fmt(r[k])}`; const name = r.model_url ? `${esc(r.model)}` : esc(r.model); const tags = isBase ? "" : `${esc(r.access)}` + (r.verified === "yes" ? "" : `self-reported`); const meta = isBase ? "Reference line" : [r.organization, r.params_b && `${r.params_b}B`].filter(Boolean).map(esc).join(", "); const src = r.source_url ? ` source` : ""; + const note = isBase ? "" : (r.notes ? `${esc(r.notes)}` : ""); return ` ${isBase ? "" : rankOf.get(r)} - ${name}${meta}${src}${tags ? `
${tags}
` : ""} + ${name}${meta}${src}${note}${tags ? `
${tags}
` : ""} ${esc(r.method)} - ${ASPECTS.map(k => cell(k)).join("")} - ${cell("overall", "num overall")} + ${ASPECTS.map(k => bar(k)).join("")} + ${bar("overall", "num overall")} + ${bar("acc_fs", "num overall")} + ${plain("miss_rate")}${plain("false_flip_rate")}${plain("a_rate")} `; - }).join(""); + }; + + tbody.innerHTML = groups.map(([title, g]) => + `${esc(title)}` + g.map(rowHtml).join("")).join(""); } board.querySelectorAll("th[data-sort] button").forEach(btn => btn.addEventListener("click", () => { const th = btn.closest("th"), k = th.dataset.sort; - state.dir = state.sort === k ? -state.dir : (["model", "method"].includes(k) ? 1 : -1); + state.dir = state.sort === k ? -state.dir : (["model", "method"].concat(LOWER_IS_BETTER).includes(k) ? 1 : -1); state.sort = k; board.querySelectorAll("th[data-sort]").forEach(h => h.removeAttribute("aria-sort")); th.setAttribute("aria-sort", state.dir === -1 ? "descending" : "ascending"); @@ -133,8 +158,9 @@ const variants = { melody: { mel: m => m.map((n, i) => i === 4 ? [7, 12, 1] : n), changed: { mel: [4] }, caption: "In clip B one note moves up. Rhythm, timbre and harmony are identical, so the answer is: different melody." }, - rhythm: { mel: m => m.map((n, i) => i === 2 ? [4, 11, 1] : i === 3 ? [5, 9, 2] : n), changed: { mel: [2, 3] }, - caption: "Clip B plays the same pitches with different timing. The answer is: different rhythm." }, + rhythm: { mel: m => m.map(([t, p, d]) => [t * 0.8, p, d * 0.8]), ch: c => c.map(([t, ps, d]) => [t * 0.8, ps, d * 0.8]), + changed: { mel: [0, 1, 2, 3, 4, 5, 6, 7], ch: [0, 1] }, + caption: "Rhythm items ask which clip is faster. Clip B plays the same notes at a faster tempo, so the answer is: clip B." }, timbre: { timbre: true, changed: { mel: [0, 1, 2, 3, 4, 5, 6, 7], ch: [0, 1] }, caption: "Every note is the same, played on a different instrument. The answer is: different timbre." }, harmony: { ch: c => c.map((n, i) => i === 1 ? [8, [0, 4], 8] : n), changed: { ch: [1] }, @@ -172,7 +198,7 @@ }); svg.innerHTML = out; caption.textContent = variants[aspect].caption; - q.textContent = `Same or different ${aspect}?`; + q.textContent = aspect === "rhythm" ? "Which clip is faster?" : `Same or different ${aspect}?`; pair.querySelectorAll(".aspects button").forEach(b => b.setAttribute("aria-pressed", b.dataset.aspect === aspect)); } svg.setAttribute("viewBox", `0 0 ${W} ${top * 2 + laneH * 2 + gap}`); diff --git a/site/assets/config.anon.js b/site/assets/config.anon.js index 6af614e..895274a 100644 --- a/site/assets/config.anon.js +++ b/site/assets/config.anon.js @@ -5,13 +5,13 @@ window.MLB_CONFIG = { title: "MusicListenBench", lab: "Anonymous authors (under review)", labUrl: null, - githubRepo: null, // TODO: Anonymous GitHub mirror, e.g. "https://anonymous.4open.science/r/XXXX" - csvSourceUrl: null, // optional: the CSV inside the mirror, e.g. ".../r/XXXX/leaderboard.csv" - datasetUrl: null, // TODO: anonymous dataset location, or leave null if it is in the supplementary material + githubRepo: "https://anonymous.4open.science/r/MusicListenBench-F38B", + csvSourceUrl: "https://anonymous.4open.science/r/MusicListenBench-F38B/leaderboard.csv", + datasetUrl: null, // the item files are in the code repository (data/); the About page links to it spaceUrl: null, paperUrl: null, contact: null, benchmarkVersion: "1.0", - itemsPerAspect: 500, + itemsPerAspect: 250, csvPath: "leaderboard.csv", }; diff --git a/site/assets/config.js b/site/assets/config.js index afb5491..2b481d2 100644 --- a/site/assets/config.js +++ b/site/assets/config.js @@ -3,12 +3,12 @@ window.MLB_CONFIG = { title: "MusicListenBench", lab: "AMAAI Lab, SUTD", labUrl: null, // TODO: lab homepage URL - githubRepo: "https://github.com/TODO-org/MusicListenBench", // TODO + githubRepo: "https://github.com/AMAAI-Lab/MusicListenBench", datasetUrl: "https://huggingface.co/datasets/amaai-lab/MusicListenBench", spaceUrl: "https://huggingface.co/spaces/amaai-lab/MusicListenBench-leaderboard", // TODO: confirm paperUrl: null, // TODO: arXiv / OpenReview link once public contact: null, // TODO: e.g. "someone@sutd.edu.sg" benchmarkVersion: "1.0", - itemsPerAspect: 500, + itemsPerAspect: 250, csvPath: "leaderboard.csv", }; diff --git a/site/assets/style.css b/site/assets/style.css index e942e46..ddb3830 100644 --- a/site/assets/style.css +++ b/site/assets/style.css @@ -168,3 +168,10 @@ code { font-size: .92em; } @media (max-width: 860px) { .hero { grid-template-columns: 1fr; } } + +/* leaderboard with FLIP/STAY columns */ +table { min-width: 1080px; } +td.num { width: 7%; } +td.plain span { font-weight: 600; font-variant-numeric: tabular-nums; } +tr.group th { background: color-mix(in srgb, var(--ink) 6%, transparent); color: var(--ink); font-size: .85rem; padding: .55rem .95rem; } +.model small.note { display: block; margin-top: .15rem; } diff --git a/site/index.html b/site/index.html index cbe5cde..952cce7 100644 --- a/site/index.html +++ b/site/index.html @@ -27,7 +27,7 @@

Can audio LLMs hear the difference yet?

-

Each test item is a pair of clips that differ in exactly one attribute: melody, rhythm, timbre or harmony. A model scores only if it hears that one change and ignores everything else.

+

Each test item is two short clips and one question about melody, rhythm, timbre or harmony. Every item comes in two variants: in FLIP the music changes and the answer switches, in STAY only the sound changes (room echo, EQ, hiss, transposition) and the answer stays. A model that ignores the audio scores 50%.

See the rankings Submit a result @@ -48,7 +48,7 @@

Can audio LLMs hear the difference yet?

Rankings

-

Accuracy (%) on the 2,000-item test split, 500 items per attribute. The thin tick on each bar marks chance. — models, last updated —.

+

Test split: 1,000 clean items (250 per task) and 2,000 FLIP/STAY items. All numbers are percentages; the thin tick on each bar marks 50%. — models, last updated —.

@@ -68,12 +68,16 @@

Rankings

- + + + + + - Loading results… + Loading results…
-

Overall is the mean of the four attribute scores. TODO: confirm this definition. Self-reported results have not yet been reproduced by the maintainers.

+

Melody … Clean (all): accuracy on the clean items; “all” is the mean of the four tasks. AccFS: accuracy on the FLIP and STAY items, 100 − (MR + FFR) / 2; a model that ignores the audio scores 50% whatever letter it prefers. MR (miss rate): share of FLIP items answered wrongly. FFR (false-flip rate): share of STAY items answered wrongly. Lower is better for MR and FFR. A rate: share of ‘A’ answers on the clean items, so letter bias is visible. Models are ranked by AccFS within each group. Self-reported results have not yet been reproduced by the maintainers.

diff --git a/site/submit.html b/site/submit.html index d652f5f..eaa31ad 100644 --- a/site/submit.html +++ b/site/submit.html @@ -26,16 +26,16 @@

Submit a result

-

The leaderboard is a single CSV file on GitHub. You add your result with a pull request; this page and the Hugging Face Space update when it merges.

+

The leaderboard is a single CSV file on GitHub. You score your model with the official script, add one row with a pull request, and this page and the Hugging Face Space update when it merges.

  1. Run the official evaluation

    -

    Use the evaluation script from the GitHub repository on the v1.0 test split. TODO: install and run command.

    -
    pip install TODO
    -python eval.py --model YOUR_MODEL --split test --out results.json
  2. +

    Answer the 3,000 test items of the v1.0 test split (1,000 clean, 1,000 FLIP and 1,000 STAY items) and write one JSON Lines file: a first line with the model name and version, how the answer was read (letter probabilities or generated text), whether the model used the training split, the date and an optional link to code, then one record per item with the item ID and the letter ‘A’ or ‘B’. Missing or invalid answers count as wrong. Training on test items is not allowed. The format is described at the top of musiclistenbench/scoring/score_submission.py in the GitHub repository, and data/example_submission.jsonl is an example. Score the file with the official script, run from the repository root:

    +
    python -m musiclistenbench.scoring.score_submission submission.jsonl
  3. Add one row per model and method

    Append to leaderboard.csv. Scores are percent accuracy with one decimal. Set verified to no; maintainers change it after reproducing your numbers.

    -
    type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,verified,submitted_by,date,source_url,benchmark_version,notes
    -model,My-Audio-LLM,My Lab,https://huggingface.co/...,open,7,zero-shot,61.2,55.4,70.8,52.6,60.0,no,@your-github,2026-10-01,https://arxiv.org/...,1.0,
  4. +
    type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes
    +model,My-Audio-LLM,My Lab,https://huggingface.co/...,open,7,zero-shot,60.0,55.2,66.8,52.0,58.5,57.6,40.1,44.7,52.0,logprob,no,@your-github,2026-10-01,https://arxiv.org/...,1.0,
    +

    The first four score columns are clean accuracy per task, overall is clean accuracy on all 1,000 clean items, and acc_fs, miss_rate, false_flip_rate and a_rate are FLIP/STAY accuracy, miss rate, false-flip rate and the share of ‘A’ answers on the clean items. method is zero-shot, or how you used the training split (for example GRPO (LoRA)); readout is logprob or generated.

  5. Open a pull request

    A check validates the file automatically and explains any problem. Fix it and push again; the check reruns.

  6. Maintainers review and merge