ZeroHour

Search: “Suno v5”

3 stories

YuE2 · Frontier Music with Symbolic Planning

YuE2, a 3.59B-parameter music generation model, scores 6.9632 on SongBench, beating Suno v5 via symbolic planning.

YuE2 is a music generation model of roughly 3.59B parameters and 28 layers supporting song creation, covering, and agentic editing through editable ABC symbolic scores. Its best-of-8 setting reaches 6.9632 on SongBench, the highest mean among 15 evaluated settings on WildSongBench (192 prompts), ahead of Suno v5 at 6.8721. The project also introduces MERT2, whose 632M-parameter encoders achieve state of the art on 14 of 15 MARBLE metrics, and SheetSage2, which transcribes beats, downbeats, key, chords, structure, and melody with SOTA on 10 of 13 benchmark metrics.

StepAudio 3 Music Technical Report

StepAudio 3 Music introduces long-form text-controlled music generation using ABC-notation planning and flow-matching diffusion, ranking near the top music arena.

StepAudio 3 Music generates long-form, text-controlled music using a 50-Hz single-codebook tokenizer with 65,536 entries and a flow-matching diffusion Transformer over VAE latents. A Mixture-of-Experts autoregressive model first plans an arrangement in ABC notation (ABC-CoT) before predicting music tokens. With DPO fine-tuning, it tops AudioBox content and production quality scores and reaches Quality Elo 1105 on the Artificial Analysis Music Arena, behind Suno V5.5 and Mureka. Generation covers songs, accompaniment from dry vocals, and cover synthesis up to 5 minutes 30 seconds at 48-kHz output.

Hugging Face daily papers · 5d agoAI research

Suno releases its first AI music model made with record industry help

Suno released its v6 music model family (v6, v6-wild, v6-mini), the first trained with licensed data from Warner Music Group, BMG, and Believe.

Suno's v6 comes in three variants: v6, the more unpredictable v6-wild, and resource-light v6-mini offered free to all users. The model was trained from the ground up on a new dataset including licensed content from Warner Music Group, BMG, and Believe, plus user data, though it is unclear if all dubiously obtained content was excluded. v6 shows dramatically improved genre fidelity, adds plain-language chat editing of individual song elements, multi-element mashups, and prompts based on images, video, or audio. The Verge notes it still cannot produce intentional imperfections like off-key vocals, and v6 starts rolling out now with older models eventually retired.

The Verge · AI · 6d agoModel release 3 sources