ZeroHour
Hugging Face daily paperspublished ()ingested Haiwen Diao, Jiahao Wang, Chenjing Ding

SenseNova-U1.5: Towards Native Unified Visual Intelligence

infoModel releaseimportance 48
AI summary · glm-5.3-flash

SenseTime releases SenseNova-U1.5, an 8B-MoT encoder-free multimodal model unifying visual understanding, reasoning, and generation with native 4K resolution.

SenseNova-U1.5 is an 8B mixture-of-transformers multimodal model with an encoder-free, VAE-free architecture that understands, reasons about, and generates visual content at native resolutions up to 4K. Post-training optimizes specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, consolidated through multi-expert on-policy distillation. Evaluations report gains in image fidelity, text rendering, multi-reference editing, and instruction following. The team plans to open-source training code including supervised fine-tuning, reinforcement learning, and on-policy distillation.

  • 8B-MoT encoder-free, VAE-free architecture unifies multimodal understanding and generation
  • Specialized experts handle aesthetics, bilingual text rendering, infographics, and image editing
  • Multi-expert on-policy distillation consolidates expert capabilities
  • Training code for SFT, RL, and distillation will be open-sourced
VendorsSenseTime
OrganizationsSenseTime
Full article179 words · extracted from huggingface.co · click to collapse

We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions of up to 4K. For post-training, we optimize specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidate their capabilities through multi-expert on-policy distillation. Across extensive evaluations, SenseNova-U1.5 largely advances image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while improving instruction following and preserving subject identity, geometry, and unmodified regions. Despite limited exposure to structured formats in its generation data, SenseNova-U1.5 generalizes effectively to long, complex, and structured visual instructions, further proving that multimodal understanding can transfer to visual planning and creation. Together, these findings position native unified modelling as a promising path towards systems that perceive, reason and create within a fully end-to-end framework. We will open-source training code, including supervised fine-tuning, reinforcement learning, and on-policy distillation.

Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.11929