World Embedding Benchmark
World Embedding Benchmark offers 8,000 simulation cases with physical annotations to test how video embeddings encode physics.
The World Embedding Benchmark comprises 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism, pairing rendered videos with simulation-derived physical annotations. Pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific pairs improves retrieval but degrades physical-property regression, exposing an alignment-versus-quantitative-information trade-off. Using embeddings to retrieve references for retrieval-augmented generation with MiniMax-H3 improves physical fidelity of generated videos.
- 8,000 controlled simulation cases across 80 physical-domain families
- Omnimodal embedding models show weak cross-modal physical alignment
- Contrastive training improves retrieval but degrades property regression
- Retrieved references improve physical fidelity of MiniMax-H3 video generation
Full article192 words · extracted from arxiv.org · click to collapse
Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics, dynamics, and optics & electromagnetism. Each case pairs a rendered video with simulation-derived physical annotations, supporting three complementary tasks: text-video retrieval, physical-property regression, and multiple-choice video-description pair classification. We use these tasks to distinguish cross-modal physical alignment from the recoverability of quantitative physical information. Evaluated pre-trained omnimodal embedding models show weak retrieval and near-chance within-family pair classification, while lightweight probes recover useful physical information from frozen video embeddings. Continual contrastive training with physics-specific video-text pairs improves retrieval and pair classification but degrades physical-property regression, revealing a trade-off between alignment and quantitative information recoverability. Finally, we use the embeddings to retrieve reference videos for retrieval-augmented generation with MiniMax-H3. Retrieved references improve the physical fidelity of generated videos, with stronger retrieval models yielding larger gains in our experiments. Together, these findings highlight the need to evaluate physical alignment and property recoverability jointly, and demonstrate the utility of physical representations for improving video generation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03632