YuE2 · Frontier Music with Symbolic Planning
YuE2, a 3.59B-parameter music generation model, scores 6.9632 on SongBench, beating Suno v5 via symbolic planning.
YuE2 is a music generation model of roughly 3.59B parameters and 28 layers supporting song creation, covering, and agentic editing through editable ABC symbolic scores. Its best-of-8 setting reaches 6.9632 on SongBench, the highest mean among 15 evaluated settings on WildSongBench (192 prompts), ahead of Suno v5 at 6.8721. The project also introduces MERT2, whose 632M-parameter encoders achieve state of the art on 14 of 15 MARBLE metrics, and SheetSage2, which transcribes beats, downbeats, key, chords, structure, and melody with SOTA on 10 of 13 benchmark metrics.
- YuE2 (3.59B parameters, 28 layers) generates, covers, and edits songs via semantic music tokens and ABC scores
- Best-of-8 YuE2 hits 6.9632 on SongBench, topping Suno v5 (6.8721) across 15 WildSongBench settings
- MERT2 encoders (632M parameters) achieve SOTA on 14 of 15 MARBLE music-representation metrics
- SheetSage2 reaches SOTA on 10 of 13 audio-to-score transcription benchmarks across six tasks
Full article975 words · extracted from map-yue2.github.io · click to collapse
Listen to a song, then explore the melody, rhythm, and chords in its symbolic plan.
Selected score
Loading selected song…
The generated song
The symbolic score
Original score recording
Interactive ABC score Red notes follow the score recording.
View original score pages
All selected scores
A familiar song can take a different shape. Listen to changes in melody, lyrics, tempo, and arrangement.
Agentic music editing
A song takes shape through a conversation. Loading the editing story…
The listening selection, gathered across genres and languages.
Search Language Genre Generation
Model & Results
YuE2 (best-of-8) reaches 6.9632 on SongBench, the highest observed mean among 15 evaluated settings on WildSongBench (192 prompts). Suno v5 scores 6.8721 in the same comparison.
Model architecture
Composing in symbols, performing in audio.
Explore benchmark scores WildSongBench · 15 settings · 7 metrics
WildSongBench192 prompts
Compare by
| Rank | System / setting | Prompts | SongBench ↑ |
|---|
How to read these results
WildSongBench. 192 prompts and 15 system settings. The table reports automatic evaluation scores. Best-of-8 selects one of eight generations by musicality, prompt control, and lyric accuracy.
Figure 1. Song quality combines SongBench and SongEval; text alignment combines MuLan, AllMusicCaps, and prompt control. Both axes show normalized comparison indices. Bubble area represents AudioBox production quality.
MERT2 · Music representations
Learning the structure behind the sound.
State of the art on MARBLE
MERT2-30s and MERT2-FS (full-song) achieve SOTA on 14 of 15 MARBLE metrics, leading across tagging, key, genre, and emotion recognition.
- 14 / 15
- 91.72
- 67.05
SOTA metricsMERT2-30s & MERT2-FS
Genre accuracy · GTZANMERT2-30s · score × 100
Key refined accuracy · GiantStepsMERT2-FS · score × 100
Explore MERT2 benchmark scores MARBLE · 15 metrics · 2 encoders
| Benchmark / metric | Best published baseline | MERT2-30s | MERT2-FS |
|---|---|---|---|
| MTT · TaggingROC-AUC | 91.70PupuJEPA-Large | 91.91 | 91.74 |
| MTT · TaggingAverage precision | 40.80PupuJEPA-Large | 41.29 | 41.20 |
| GiantSteps · KeyRefined accuracy | 66.10PupuJEPA-Large | 66.97 | 67.05 |
| GTZAN · GenreAccuracy | 86.90PupuJEPA-Large | 91.72 | 90.69 |
| GTZAN · BeatF1 | 91.00PupuJEPA-Large | 90.59 | 90.57 |
| EmoMusic · ValenceR² | 62.50PupuJEPA-Large | 63.23 | 63.52 |
| EmoMusic · ArousalR² | 78.50PupuJEPA-Huge | 80.01 | 78.14 |
| MTG-Jamendo · InstrumentROC-AUC | 78.40PupuJEPA-Large | 80.27 | 80.27 |
| MTG-Jamendo · InstrumentAverage precision | 21.20PupuJEPA-Large | 22.89 | 23.51 |
| MTG-Jamendo · Mood / themeROC-AUC | 76.20PupuJEPA-Large | 79.44 | 78.74 |
| MTG-Jamendo · Mood / themeAverage precision | 15.50Dasheng-1.2B | 16.68 | 15.74 |
| MTG-Jamendo · GenreROC-AUC | 86.30AudioMAE++ | 88.01 | 87.98 |
| MTG-Jamendo · GenreAverage precision | 20.10PupuJEPA-Large / PupuJEPA-Huge | 21.22 | 20.66 |
| MTG-Jamendo · Top 50ROC-AUC | 83.10AudioMAE++ / PupuJEPA-Huge | 84.18 | 84.13 |
| MTG-Jamendo · Top 50Average precision | 31.10AudioMAE++ | 32.17 | 31.62 |
SOTA counts use the best score across the two MERT2 encoders against the nine published baselines in this comparison. Both encoders have 632M parameters. MERT2-30s uses a 30-second training context; MERT2-FS uses 300 seconds. These are full-context representation benchmarks. MERT2 reports the best observed results across representations selected using test scores; each ROC-AUC / AP pair uses the same representation.
SheetSage2 · Audio to score
Hear a song. Read its composition.
Six transcription tasks, one model
SheetSage2 achieves SOTA on 10 of 13 benchmark metrics with one model for beat, downbeat, key, chord, structure, and melody transcription.
- 10 / 13
- 82.51
- 90.08
SOTA metricsOne model · six transcription tasks
Vocal melody · RWC-PopPitch-class note F1 · score × 100
Chord recognition · osu2017Maj/min · score × 100
Explore SheetSage2 benchmark scores 6 tasks · 13 metrics
| Benchmark / task | Metric | Best comparison | SheetSage2 |
|---|---|---|---|
| GTZANBeat | F1 | 88.75Beat This! | 85.65 |
| osu2017Beat | F1 | 91.55Madmom | 92.29 |
| GTZANDownbeat | F1 | 78.28Beat This! | 79.51 |
| osu2017Downbeat | F1 | 84.99Beat This! | 91.97 |
| GiantStepsKey | Weighted score | 74.62Madmom | 77.73 |
| GTZANKey | Weighted score | 74.43S-KEY | 75.77 |
| osu2017Chord | Maj/min | 84.59Jiang et al. 2019 | 90.08 |
| Chords1217Chord | Maj/min | 84.09ChordFormer | 83.81 |
| HarmonixSetStructure | Accuracy | 80.03SongFormer | 80.51 |
| HarmonixSetStructure | Boundary F1 · 0.5 s | 70.63SongFormer | 67.96 |
| HarmonixSetStructure | Boundary F1 · 3 s | 79.50SongFormer | 82.86 |
| RWC-PopMelody | Vocal pitch-class F1 | 62.71SheetSage1 | 82.51 |
| RWC-PopMelody | Full pitch-class F1 | 64.02SheetSage1 | 75.29 |
SOTA counts refer to the leading scores against SheetSage1, Madmom, and the task-specific systems in this comparison. Results use one model selected by validation loss. Melody F1 measures pitch-class notes; structure F1 measures section boundaries at the stated tolerance. On Chords1217, ChordFormer uses five-fold cross-validation, while SheetSage2 evaluates one fixed model on all 1,217 tracks.
Training data
Our models are trained primarily on CC0 music and synthetic data. Tokenwave.AI provides most of our synthetic training data under license. We are committed to the ethical and responsible use of data.
- 700K hours
- 28.4K hours
- 346K hours
MERT2
SheetSage2
YuE2
Text extracted automatically; images, tables and formatting may be missing. Original: https://map-yue2.github.io/