Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model
Physis-Lang improves physical video generation, and Cosmos3-Nano variants surpass proprietary Veo 3.1.
Physis-Lang represents physical processes in language covering entities, causes, interactions, principles, evolution, and effects, then optimizes that language across curation, training, and generation. PhysCapBench decomposes processes into atomic assertions and scores caption recall and precision, while an agentic loop fixes assertion errors and retrieves videos for missing processes. On four physical-video benchmarks, Wan and Cosmos backbones improve, and Cosmos3-Nano variants enhanced by Physis-Lang surpass proprietary Veo 3.1.
- Physis-Lang treats physical language as a shared, optimizable representation.
- PhysCapBench scores captions using assertion-level recall and precision.
- An agentic loop revises caption instructions from assertion errors.
- Enhanced Cosmos3-Nano models surpass proprietary Veo 3.1 on physical plausibility.
Full article193 words · extracted from huggingface.co · click to collapse
Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.40358