Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
Dream4ACT unifies multi-robot video-action modeling with visual action views, scoring 88.98% on RoboTwin 2.0.
Dream4ACT is a multi-embodiment world model that represents robot actions as shared visual action views rendered from four virtual cameras using URDF forward kinematics. Observation and action sequences share a video autoencoder and diffusion transformer, with masked flow-matching supporting forward dynamics, inverse dynamics, and joint generation. Executable actions are recovered without a learned embodiment-specific decoder. It reports 88.98% average success on RoboTwin 2.0 and an overall score of 65.66 on TriWorldBench.
- Action views render joint targets from four virtual cameras via URDF forward kinematics.
- One diffusion model handles forward dynamics, inverse dynamics, and joint generation.
- Training-free multiview recovery turns predicted views into executable joint sequences.
- Average success is 88.98% on RoboTwin 2.0 and 65.66 overall on TriWorldBench.
Full article203 words · extracted from huggingface.co · click to collapse
Video generation models (VGMs) offer strong spatiotemporal priors for embodied observation--action modeling. However, joint-space action vectors lack explicit image-space structure and vary in dimensionality and semantics across embodiments, making it challenging to directly leverage the rich spatiotemporal priors of VGMs. End-effector visualizations provide an alternative but do not specify the full articulated configuration needed for robot execution. We present Dream4ACT, a world model built for joint video-action modeling across embodiments. To unify action representations across embodiments, we introduce a shared visual action interface, called action views, which render target joint configurations from four prescribed virtual cameras using URDF-based forward kinematics. This shared visual representation preserves embodiment-specific articulated geometry while allowing observation and action sequences to share a video autoencoder and diffusion transformer. Through masked flow-matching, our model supports forward dynamics, inverse dynamics, and joint observation--action generation within a single jointly trained model by varying which future sequences are corrupted. To recover executable action sequences from predicted action views, we propose a training-free, URDF-constrained multiview recovery mechanism, without a learned embodiment-specific decoder. Dream4ACT achieves an average success rate of 88.98\% on RoboTwin~2.0 and an overall score of 65.66 on TriWorldBench, supporting effective closed-loop manipulation and competitive action-conditioned multiview prediction through the visual action interface.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.40153