EVO-WAM: Evolving World Action Models through Video-Action Verification
EVO-WAM self-improves world action models on unseen robot tasks without new expert demonstrations.
EVO-WAM adapts world action models to unseen robot tasks by learning from their own generated video-action trajectories, without new expert demonstrations or external execution of candidate actions. It adds state prediction and anchored multi-frame context, then keeps task-completing prefixes verified by a vision-language model and an inverse dynamics model. Iterative retraining raises average success on seven unseen RoboTwin 2.0 tasks from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero. On three real long-horizon tasks, Cosmos3 improves from 20.0% to 76.7%.