AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
AnyStep-WAM cuts world-action denoising steps by up to 85% while holding RoboTwin success rates.
World-action models typically use a fixed denoising budget even when action chunks differ in error sensitivity. AnyStep-WAM distills interval-conditioned flow maps from a frozen teacher and uses a risk-benefit scheduler to choose the smallest budget that meets fidelity needs. On Motus, FastWAM, and LingBotVA with RoboTwin 2.0, average denoising steps drop 60.2%, 49.8%, and 85.28% while baseline success holds. One-step success rises by 7.07%, 12.08%, and 8.94%, and six real-world manipulation tasks also validate the method.
- Scheduler picks the smallest denoising budget from a one-step preview
- Steps fall 60.2%, 49.8%, and 85.28% on three WAMs
- One-step success rises 7.07%, 12.08%, and 8.94%
- Tested on RoboTwin 2.0 and six real manipulation tasks
Full article192 words · extracted from huggingface.co · click to collapse
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33748