Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 24 additions & 6 deletions leaderboard.csv
Original file line number Diff line number Diff line change
@@ -1,6 +1,24 @@
type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,verified,submitted_by,date,source_url,benchmark_version,notes
baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,yes,maintainers,2026-09-24,,1.0,TODO confirm chance level (assumes two answer options)
model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,zero-shot,61.2,55.4,70.8,52.6,60.0,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,GRPO,74.0,63.8,78.2,60.4,69.1,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
model,PLACEHOLDER open model A,PLACEHOLDER org,https://huggingface.co/,open,7,SFT,60.8,54.6,71.4,51.8,59.65,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
model,PLACEHOLDER closed model B,PLACEHOLDER org,https://example.com/,closed,,zero-shot,66.4,58.0,73.6,55.2,63.3,yes,maintainers,2026-09-24,,1.0,PLACEHOLDER row - replace with paper results
type,model,organization,model_url,access,params_b,method,melody,rhythm,timbre,harmony,overall,acc_fs,miss_rate,false_flip_rate,a_rate,readout,verified,submitted_by,date,source_url,benchmark_version,notes
baseline,Random chance,—,,—,,—,50.0,50.0,50.0,50.0,50.0,50.0,50.0,50.0,,—,yes,maintainers,2026-09-30,,1.0,"Guessing at random. A model that ignores the audio scores 50% on FLIP/STAY, whatever letter it prefers"
baseline,Always answer A,—,,—,,—,54.0,50.8,56.8,49.6,52.8,49.5,53.0,48.1,100.0,—,yes,maintainers,2026-09-30,,1.0,Ignores the audio and always answers A
baseline,Human listeners (12),—,,—,,—,79.2,87.5,95.8,87.5,87.5,87.0,5.2,20.8,,—,yes,maintainers,2026-09-30,,1.0,"Paper Table 2. Twelve listeners, 24 items each: 24 clean answers per task, 96 FLIP and 96 STAY answers"
model,Qwen2-Audio,Alibaba,https://huggingface.co/Qwen/Qwen2-Audio-7B-Instruct,open,8.4,zero-shot,46.0,49.2,43.2,50.4,47.2,50.7,47.7,50.9,1.4,logprob,yes,maintainers,2026-09-30,,1.0,
model,Qwen2.5-Omni,Alibaba,https://huggingface.co/Qwen/Qwen2.5-Omni-7B,open,10.7,zero-shot,68.8,50.0,66.8,59.2,61.2,61.2,36.7,41.0,52.4,logprob,yes,maintainers,2026-09-30,,1.0,
model,Audio Flamingo 3,NVIDIA,https://huggingface.co/nvidia/audio-flamingo-3-hf,open,16.5,zero-shot,51.2,52.8,58.0,48.8,52.7,51.4,49.6,47.6,87.7,logprob,yes,maintainers,2026-09-30,,1.0,
model,Phi-4-multimodal,Microsoft,https://huggingface.co/microsoft/Phi-4-multimodal-instruct,open,5.6,zero-shot,54.4,50.8,56.8,47.2,52.3,49.1,53.5,48.3,95.9,logprob,yes,maintainers,2026-09-30,,1.0,
model,Kimi-Audio,Moonshot AI,https://huggingface.co/moonshotai/Kimi-Audio-7B-Instruct,open,7.9,zero-shot,57.6,52.8,53.2,47.2,52.7,51.0,52.3,45.7,90.1,logprob,yes,maintainers,2026-09-30,,1.0,
model,MiMo-Audio,Xiaomi,https://huggingface.co/XiaomiMiMo/MiMo-Audio-7B-Instruct,open,8.0,zero-shot,54.0,59.6,58.0,54.4,56.5,52.3,49.4,46.1,91.3,logprob,yes,maintainers,2026-09-30,,1.0,
model,Fun-Audio-Chat,FunAudioLLM,https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B,open,9.5,zero-shot,50.0,49.6,60.0,51.6,52.8,53.7,46.6,46.1,62.2,logprob,yes,maintainers,2026-09-30,,1.0,
model,Step-Audio 2,StepFun,https://huggingface.co/stepfun-ai/Step-Audio-2-mini,open,8.3,zero-shot,52.4,52.4,62.8,50.0,54.4,52.7,49.1,45.5,69.4,logprob,yes,maintainers,2026-09-30,,1.0,
model,GPT-Audio-mini,OpenAI,,closed,,zero-shot,47.6,50.8,54.4,51.6,51.1,49.6,50.6,50.3,81.1,generated,yes,maintainers,2026-09-30,,1.0,
model,GPT-Audio-1.5,OpenAI,,closed,,zero-shot,45.6,44.0,52.0,47.6,47.3,52.1,45.8,50.1,18.3,generated,yes,maintainers,2026-09-30,,1.0,
model,Gemini 2.5 Pro,Google,,closed,,zero-shot,80.0,79.6,88.4,68.8,79.2,74.2,24.7,26.9,60.8,generated,yes,maintainers,2026-09-30,,1.0,
model,Gemini 3.8 Flash,Google,,closed,,zero-shot,53.6,72.4,85.6,54.4,66.5,64.0,33.0,39.1,43.3,generated,yes,maintainers,2026-09-30,,1.0,
model,Qwen2-Audio,This paper,,open,8.4,GRPO (LoRA),70.8,51.6,92.4,60.8,68.9,63.8,29.7,42.7,31.3,logprob,yes,maintainers,2026-09-30,,1.0,
model,Qwen2-Audio,This paper,,open,8.4,GRPO (full),70.4,52.0,80.8,61.6,66.2,61.5,31.7,45.4,23.0,logprob,yes,maintainers,2026-09-30,,1.0,
model,Qwen2.5-Omni,This paper,,open,10.7,GRPO (LoRA),99.6,100.0,99.6,96.8,99.0,86.0,1.9,26.1,52.2,logprob,yes,maintainers,2026-09-30,,1.0,
model,Qwen2.5-Omni,This paper,,open,10.7,GRPO (full),100.0,99.6,100.0,85.2,96.2,83.5,4.2,28.9,49.0,logprob,yes,maintainers,2026-09-30,,1.0,
model,Audio Flamingo 3,This paper,,open,16.5,GRPO (LoRA),57.2,50.8,97.6,74.8,70.1,64.5,29.7,41.4,61.1,logprob,yes,maintainers,2026-09-30,,1.0,
model,Audio Flamingo 3,This paper,,open,16.5,GRPO (full),68.0,69.2,99.6,80.8,79.4,70.4,22.5,36.7,43.2,logprob,yes,maintainers,2026-09-30,,1.0,"Stopped at step 500, where validation accuracy peaks; the other GRPO runs use step 1000"
model,Phi-4-multimodal,This paper,,open,5.6,GRPO (LoRA),52.0,50.8,56.8,49.6,52.3,49.1,53.4,48.4,99.5,logprob,yes,maintainers,2026-09-30,,1.0,
model,Phi-4-multimodal,This paper,,open,5.6,GRPO (full),55.6,50.8,88.8,54.0,62.3,58.9,42.9,39.3,79.3,logprob,yes,maintainers,2026-09-30,,1.0,
2 changes: 1 addition & 1 deletion scripts/build.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
# TODO: add every author's surname, and any other giveaway (grant names, emails).
LEAK_TERMS = [
"AMAAI", "amaai-lab", "SUTD", "Singapore University",
"sleeping-ai", # style reference, not needed in either build
"roy", "heremmans", "mehrish", # style reference, not needed in either build
]

ANON_BRAND = "Anonymous submission"
Expand Down
48 changes: 31 additions & 17 deletions scripts/validate.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,21 +4,36 @@
from datetime import date

COLUMNS = ["type","model","organization","model_url","access","params_b","method",
"melody","rhythm","timbre","harmony","overall","verified","submitted_by",
"date","source_url","benchmark_version","notes"]
"melody","rhythm","timbre","harmony","overall","acc_fs","miss_rate","false_flip_rate","a_rate",
"readout","verified","submitted_by","date","source_url","benchmark_version","notes"]
Comment on lines 6 to +8
ASPECTS = ["melody","rhythm","timbre","harmony"]
ENUMS = {"type": {"model","baseline"}, "verified": {"yes","no"}}
MODEL_ACCESS = {"open","closed"}
MODEL_READOUT = {"logprob","generated"}
KNOWN_VERSIONS = {"1.0"}
# TODO(maintainers): confirm "overall" is the macro-average of the four aspects.
# Set to False if overall is scored as its own task.
OVERALL_IS_MEAN = True
# The test split has 250 clean items per task, so "overall" (clean accuracy on all 1,000 items)
# equals the mean of the four task scores.
# acc_fs = 100 - (miss_rate + false_flip_rate) / 2 (FLIP/STAY accuracy, paper Section 3.3).
# Scores are rounded to one decimal, so allow a small tolerance.
TOLERANCE = 0.1

errors = []
def err(line, msg):
errors.append(f"::error file=leaderboard.csv,line={line}::{msg}")

def number(row, col, line, required=True):
"""Return the value of a percent column, or None (and report) if it is missing or out of range."""
if row[col] == "" and not required:
return None
try:
v = float(row[col])
if not 0 <= v <= 100:
raise ValueError
return v
except ValueError:
err(line, f"{col} must be a number between 0 and 100 (percent)")
return None

def main(path):
with open(path, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
Expand All @@ -32,24 +47,23 @@ def main(path):
err(i, f"{col} must be one of {sorted(allowed)}, got '{row[col]}'")
if not row["model"].strip():
err(i, "model is required")
scores = {}
for a in ASPECTS + ["overall"]:
try:
v = float(row[a])
if not 0 <= v <= 100:
raise ValueError
scores[a] = v
except ValueError:
err(i, f"{a} must be a number between 0 and 100 (percent accuracy)")
if OVERALL_IS_MEAN and len(scores) == 5:
scores = {a: number(row, a, i) for a in ASPECTS + ["overall", "acc_fs", "miss_rate", "false_flip_rate"]}
number(row, "a_rate", i, required=False) # share of 'A' answers on the clean items; empty if unknown
if None not in (scores[a] for a in ASPECTS + ["overall"]):
mean = sum(scores[a] for a in ASPECTS) / 4
if abs(mean - scores["overall"]) > TOLERANCE:
err(i, f"overall ({scores['overall']}) should equal the mean of the four aspects ({mean:.2f})")
err(i, f"overall ({scores['overall']}) should equal the mean of the four tasks ({mean:.2f})")
if None not in (scores["acc_fs"], scores["miss_rate"], scores["false_flip_rate"]):
fs = 100 - (scores["miss_rate"] + scores["false_flip_rate"]) / 2
if abs(fs - scores["acc_fs"]) > TOLERANCE:
err(i, f"acc_fs ({scores['acc_fs']}) should equal 100 - (miss_rate + false_flip_rate) / 2 = {fs:.2f}")
if row["type"] == "model":
if row["access"] not in MODEL_ACCESS:
err(i, "access must be 'open' or 'closed'")
if row["readout"] not in MODEL_READOUT:
err(i, "readout must be 'logprob' (letter probabilities) or 'generated' (generated text)")
if not row["method"].strip():
err(i, "method is required (e.g. zero-shot, SFT, GRPO)")
err(i, "method is required: 'zero-shot', or how the released training split was used (e.g. GRPO (LoRA))")
if row["params_b"] and not re.fullmatch(r"\d+(\.\d+)?", row["params_b"]):
err(i, "params_b must be a number in billions, or empty if unknown")
if not row["source_url"] and row["verified"] == "no":
Expand Down
26 changes: 16 additions & 10 deletions site/about.html
Original file line number Diff line number Diff line change
Expand Up @@ -28,31 +28,37 @@
<h2>What the benchmark tests</h2>
<div class="prose">
<p>Audio large language models can caption a track and answer broad questions about it, which suggests they understand music. MusicListenBench tests whether that understanding holds at the level of individual musical attributes.</p>
<p>Every item asks one question: given two clips, are they the same or different in a named attribute? Each pair changes exactly one attribute and holds the rest fixed, so a correct answer requires hearing that attribute and ignoring the others.</p>
<p>Each item is one audio file: clip A, one second of silence, and clip B. One question follows, and the model answers with a single letter. Melody, harmony and timbre ask whether the two clips are the same; rhythm asks which clip is faster. When the clips differ, they differ in exactly the queried attribute.</p>
</div>
<ul class="aspect-list">
<li style="--c:var(--melody)"><b>Melody</b><span class="todo">TODO: how melody variants are constructed (e.g. which notes change, by how much).</span></li>
<li style="--c:var(--rhythm)"><b>Rhythm</b><span class="todo">TODO: how rhythm variants are constructed.</span></li>
<li style="--c:var(--timbre)"><b>Timbre</b><span class="todo">TODO: how timbre variants are constructed (instrument swaps, synthesis).</span></li>
<li style="--c:var(--harmony)"><b>Harmony</b><span class="todo">TODO: how harmony variants are constructed.</span></li>
<li style="--c:var(--melody)"><b>Melody</b><span>One note of a four-note motif moves by 100, 200, 400 or 700 cents. Clips are 2.5 s long.</span></li>
<li style="--c:var(--rhythm)"><b>Rhythm</b><span>The click rate of a hi-hat track differs by a ratio of 1.10, 1.20, 1.35 or 1.50, and the question is which clip is faster. Clips are 4.0 s long.</span></li>
<li style="--c:var(--timbre)"><b>Timbre</b><span>The instrument family changes (piano, guitar, violin or flute). Clips are 2.5 s long.</span></li>
<li style="--c:var(--harmony)"><b>Harmony</b><span>The third of the chord moves by 100, 200 or 300 cents. Clips are 2.0 s long.</span></li>
</ul>
</section>

<section class="block">
<h2>FLIP and STAY</h2>
<div class="prose">
<p>For each clean test item there are two more variants, made by replacing clip B. In <b>FLIP</b>, clip B is an edited clip that changes the queried attribute, so the correct answer switches. In <b>STAY</b>, clip B sounds different but keeps the attribute: room echo, a mild EQ change, background hiss, or a transposition of up to 200 cents. The correct answer stays.</p>
<p>The two variants of an item have opposite answers. A model that ignores the audio therefore gets exactly one of them right and scores 50%, whatever letter it prefers (on rhythm this holds on average). The miss rate (MR) is the share of FLIP items answered wrongly; the false-flip rate (FFR) is the share of STAY items answered wrongly. The FLIP/STAY accuracy is Acc<sub>FS</sub> = 100 − (MR + FFR) / 2.</p>
</div>
</section>

<section class="block">
<h2>Data</h2>
<div class="prose">
<p>The full dataset is on <a data-link="datasetUrl" href="#">Hugging Face</a>. A test split of 2,000 items (500 per attribute) is set aside for the leaderboard.</p>
<p><span class="todo">TODO: source of the musical material, clip length, sample rate, licence, and whether test labels are public or held out.</span></p>
<p>The benchmark has 10,000 training items (2,500 per task) and 3,000 test items: 1,000 clean items (250 per task) and their 2,000 FLIP/STAY variants. All audio is generated from symbolic music and rendered with FluidSynth and freely available soundfonts, so every answer is known exactly and no human labelling is needed. Test items use pitch ranges, a soundfont and sound effects that never appear in training. The item files, the generator and the scoring script are in the <a data-link="githubRepo" href="#">code repository</a>.</p>
</div>
</section>

<section class="block">
<h2>Evaluation</h2>
<div class="prose">
<p>Models answer a multiple-choice question for each pair. We report accuracy per attribute and overall.</p>
<p><span class="todo">TODO: exact prompt template, answer options, answer-extraction rule, how invalid answers are scored, and the chance level.</span></p>
<p>Each question ends with a fixed instruction that says which letter means which answer. Open models are scored by comparing the probabilities of the tokens ‘A’ and ‘B’; API models are scored by the letter they generate. Missing or invalid answers count as wrong. Zero-shot models and models that used the training split are listed separately.</p>
<h3>Findings from the paper</h3>
<p>We evaluate open and closed-source audio LLMs, and show that GRPO post-training improves scores while supervised fine-tuning does not. <span class="todo">TODO: one or two headline numbers.</span></p>
<p>Without training, seven of eight open audio LLMs score within 4 points of the 50% floor on FLIP/STAY. The best commercial model, Gemini 2.5 Pro, reaches 74.2%, and human listeners reach 87.0%. Post-training Qwen2.5-Omni with GRPO on the training split raises its clean accuracy from 61.2% to 99.0%, but its FLIP/STAY accuracy only reaches 86.0%: it misses almost no change (1.9%) and still false-flips on 26.1% of STAY items.</p>
</div>
</section>

Expand Down
Loading
Loading