ZeroHour
Hugging Face daily paperspublished ()ingested Xingjian Ran, Xiaoye Mo, Sihao Liu

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

infoAI researchimportance 25
AI summary · glm-5.3-flash

SceneMosaic combines image-based 3D priors with VLM agent refinement to generate diverse, simulation-ready indoor scenes 24x faster than agentic baselines.

SceneMosaic is a hybrid framework that takes an initial candidate from a learned image-to-3D prior and evolves it with VLM agents for efficiency and physical validity. It decomposes scenes into independent local units, evolves each separately, and composes the global scene via Cartesian product. On SceneEval-100 it matches the strongest agentic baseline in semantic layout quality with a 24x speedup, substantially reduces physical violations, and receives the highest human ratings. Code is publicly available.

  • Hybrid design merges parametric image-to-3D priors with agentic VLM refinement
  • Local-unit decomposition with Cartesian product composition enables scene diversity
  • 24x speedup on SceneEval-100 with fewer physical violations
  • Targets embodied AI and interactive entertainment simulation environments
Full article202 words · extracted from huggingface.co · click to collapse

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placement and refinement. Another mainstream paradigm, parametric image-to-3D scene models, produces scenes efficiently from strong priors learned from 2D images but often leads to imprecise and physically invalid scenes. More importantly, both paradigms struggle to output diverse scenes for a single input, making it hard for them to reflect the dynamically changing nature of real scenes. In this paper we propose SceneMosaic, a framework that combines the merits of both paradigms. It obtains the initial candidate from the learned image-based prior, and subsequently evolves the result through VLM agents, ensuring both efficiency and physical validity. Within the evolution process, SceneMosaic exploits the locality of natural scenes and decomposes a scene into independent local units, allowing separate evolution within each unit before composing the global scene via Cartesian product. On SceneEval-100, SceneMosaic matches the strongest agentic baseline in semantic layout quality with a 24x speedup, substantially reduces physical violations, and receives the highest human ratings. Our code is publicly available at https://github.com/rxjfighting/SceneMosaic.

Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.05594