Skip to content

Instantly share code, notes, and snippets.

@sleepingcat4
Last active July 29, 2026 11:37
Show Gist options
  • Select an option

  • Save sleepingcat4/14f5aad990b798dd7696add986450fad to your computer and use it in GitHub Desktop.

Select an option

Save sleepingcat4/14f5aad990b798dd7696add986450fad to your computer and use it in GitHub Desktop.
musicalitybench-explaination

Introducing MusicalityBench: A Living Benchmark for AI-Generated Songs

Official Announcement

We are excited to introduce MusicalityBench, a stateless, continuously updated benchmark for evaluating artificial intelligence-generated songs in real time.

Background

The concept of Musicality as a human evaluation metric was first introduced by Radford (https://arxiv.org/abs/2005.00341), where songs were judged by expert musicians against criteria for enjoyable, human-made music. MusicalityBench builds directly on this foundation, expanding the original definition of Musicality to keep pace with new models, new songs, and an evaluation landscape that never stops moving.

Why Existing Approaches Fall Short

Traditional Mean Opinion Score (MOS) evaluations have long been the default for judging songs, but being mainly used for songs does not mean MOS is sufficient. It carries real limitations:

  • No standardized question set. MOS studies are typically improvised on a per-project, per-author basis, making results difficult to compare across efforts.
  • Collapse into blind A/B testing. In many cases, MOS is reduced to a binary "AI vs. Human" detection task, which is useful for a single narrow question but says little about the actual quality or artistry of a piece.
  • Music vs. songs. MOS-style tests do not meaningfully distinguish between music, any creative instrumental composition, and songs, a deeper and more complex form of expression where spoken/sung words, delivery, and emotional tangibility all matter.
  • Cultural and temporal drift. How a song is perceived changes constantly. A song released in the early 2010s and a song released today can be judged completely differently based on shifting social, cultural, and ideological norms. Meanwhile, platforms like TikTok and Instagram Reels can send songs viral, even across languages, for reasons a static MOS or blind A/B test simply cannot capture.

How MusicalityBench Is Different

MusicalityBench is designed from the ground up to be dynamic and active:

  • The evaluation questions applied to any given song snippet change over time, evolving alongside current trends rather than staying fixed at launch.
  • The benchmark is continuously updated, with new songs and new models added on an ongoing, as-published basis rather than in periodic batches.
  • Users and creators can upload their own songs directly into the benchmark mix, keeping the pool of evaluated material perpetually fresh.

This stands in contrast to prior work such as MusicArena (https://arxiv.org/abs/2507.20900), which, while a valuable contribution in the spirit of LMSYS-style arenas, relies on a pre-defined song set curated by its authors. That approach inevitably grows stale within months as new models are released. MusicalityBench avoids this staleness by design: because the song pool is user-driven and always expanding, the benchmark stays current with both the models being evaluated and the songs the world is actually listening to.

A Living Benchmark

MusicalityBench is not a fixed leaderboard frozen in time. It is a moving target that reflects the state of AI-generated music as it actually exists today, and as it will continue to evolve tomorrow. As new models are published and new songs are uploaded, the benchmark will update to reflect them.

We're just getting started, and we're glad to have you along as MusicalityBench grows.


More details on evaluation methodology, submission guidelines, and leaderboard access will follow in subsequent updates.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment