YuE2 open music model tops Suno v5 on SongBench; DeepSeek ships 552B DeepSeek-V4.1-Flash with 1M-token context
M-A-P's open-weights YuE2-3B (~3.59B parameters) scores 6.9632 best-of-8 on SongBench versus Suno v5's 6.8721 and runs 48 kHz stereo inference on a single 24GB GPU; separately, DeepSeek released DeepSeek-V4.1-Flash, a 552B-parameter multimodal MoE with…
Two open-weight model releases trended on Hugging Face in mid-September 2026. The M-A-P (multimodal-art-projection) team released YuE2, published as YuE2-3B and trending #30 as of 2026-09-09, an open-weights music generation model of roughly 3.59B parameters and 28 layers that turns lyrics and a style prompt into complete songs with vocals and accompaniment, and also supports song covering and agentic editing via editable symbolic scores (melody and chords, including ABC notation). Its AR-NAR Mixture-of-Transformers backbone uses symbolic planning and flow matching through a VAE. On 192 WildSongBench prompts, its best-of-8 setting reports a SongBench average of 6.9632 versus 6.8721 for Suno v5 — characterized one way as state of the art among evaluated open and proprietary models, and another way as the highest mean among 15 evaluated settings. It runs 48 kHz stereo inference locally on a single 24GB NVIDIA GPU without quantization. Companion releases include YuE2-Vae, the WildSongBench dataset, music encoders (named MERT-v2 in one report and MERT2 in the other; 632M parameters, state of the art on 14 of 15 MARBLE metrics), and SheetSage2, which transcribes beats, downbeats, key, chords, structure, and melody with SOTA results on 10 of 13 benchmark metrics across six audio-to-score tasks. Separately, DeepSeek released DeepSeek-V4.1-Flash, trending #28 on Hugging Face as of 2026-09-10. It is a multimodal Mixture-of-Experts model with a 552B-parameter backbone that activates 8B parameters per token during prefill and 16B during decode. It uses a Causal Encoder-Decoder architecture, Compressed Sparse Attention 2, and FP4 KV caching to cut the global KV cache footprint to 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash. It was trained from scratch on 45T multimodal tokens, with context extended to 1M tokens after sparse-attention training at 64K sequence length, includes a 196B-parameter Engram conditional-memory module, and is released under the MIT license. Post-training combines SFT, RL, and on-policy distillation with large-scale automated synthesis of agentic tasks and a controllable reasoning-effort setting from 1 to 100. The only source disagreement is the naming of the companion music encoders (MERT-v2 vs MERT2); all benchmark and capability figures are consistent across reports.
- YuE2 (published as m-a-p/YuE2-3B) is an open-weights music generation model with ~3.59B parameters and 28 layers; it was trending #30 on Hugging Face as of 2026-09-09.
- Best-of-8 YuE2 scores 6.9632 on SongBench versus Suno v5's 6.8721, on 192 WildSongBench prompts; described as the highest mean among 15 evaluated settings and as SOTA among evaluated open and proprietary models.
- YuE2 architecture: AR-NAR Mixture-of-Transformers with symbolic planning and flow matching through a VAE; supports editable melody/chord scores in ABC notation and agentic editing workflows.
- YuE2 runs 48 kHz stereo inference locally on a single 24GB NVIDIA GPU without quantization.
- Companion releases: YuE2-Vae, the WildSongBench dataset, music encoders (632M parameters, SOTA on 14 of 15 MARBLE metrics; named MERT-v2 in one report and MERT2 in the other), and SheetSage2 (SOTA on 10 of 13 audio-to-score transcription…
- DeepSeek-V4.1-Flash is a 552B-parameter multimodal Mixture-of-Experts model activating 8B parameters per token at prefill and 16B at decode; it was trending #28 on Hugging Face as of 2026-09-10.
- Its KV cache is reduced to 890 bytes per token, roughly 1/4 of DeepSeek-V4-Flash, via Compressed Sparse Attention 2 and FP4 KV caching within a Causal Encoder-Decoder architecture.
- It was trained from scratch on 45T multimodal tokens, with context extended to 1M tokens after sparse-attention training at 64K sequence length; it includes a 196B-parameter Engram conditional-memory module and is MIT-licensed.
Coverage timelineoldest first · each row is one article
- · 6d agom-a-p/YuE2-3B — new model trending #30 on Hugging Face
Hugging Face trending models· 28
M-A-P released YuE2-3B, an open music generation model that outperforms Suno v5 on WildSongBench and runs locally on a 24GB GPU.
- · 5d agodeepseek-ai/DeepSeek-V4.1-Flash — new model trending #28 on Hugging Face
Hugging Face trending models· 80
DeepSeek releases DeepSeek-V4.1-Flash, a 552B-parameter multimodal MoE model with 1M-token context and KV cache cut to 890 bytes per token.
- · 4d agoYuE2 · Frontier Music with Symbolic Planning
Hacker News · AI· 45
YuE2, a 3.59B-parameter music generation model, scores 6.9632 on SongBench, beating Suno v5 via symbolic planning.