DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training
DreamTrue cuts AgiBot robot interaction defects from 48.12% to 6.25% and wins the 2026 world-model track.
DreamTrue is a multi-view, cross-embodiment robot world model for action-faithful, physically plausible video prediction. It aligns rendered action trajectories with videos via offline geometric calibration and uses counterfactual post-training plus a human-annotated embodied video reward model. On AgiBot, interaction defect rates fell from 48.12% to 6.25%, and the model ranked first in the AgiBot World Challenge 2026 world-model track.
- Multi-view world model predicts action-faithful robot video across embodiments.
- Offline geometric calibration aligns rendered action trajectories with target videos.
- Counterfactual post-training covers failed contacts using an embodied video reward model.
- Human-assessed interaction defects on AgiBot fell from 48.12% to 6.25%.
- Ranked first in the AgiBot World Challenge 2026 world-model track.
Full article182 words · extracted from huggingface.co · click to collapse
We present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstacles: imprecise calibration can impair action following, while limited coverage of unsuccessful interactions can bias predictions toward successful outcomes. To improve action following across embodiments, we render action trajectories into image-space conditions and introduce offline geometric calibration to align these conditions with the target videos. To broaden interaction coverage, we introduce counterfactual post-training, modifying recorded action trajectories and generating future videos under a wider range of actions and contact configurations. To provide feedback on these predictions without paired ground-truth futures, we construct a human-annotated video dataset covering robot, object, and interaction defects and use it to train an embodied video reward model. Its scores guide reinforcement-learning post-training toward more physically plausible interaction outcomes. On AgiBot, DreamTrue attains state-of-the-art action following, while reducing the human-assessed interaction defect rate from from 48.12% to 6.25%. Notably, our model ranks first in the world model track of the AgiBot World Challenge 2026. The project page can be found at https://brave-eai.github.io/DreamTrue.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12468