FastOPD: On-Policy Distillation for Lightweight VLA Deployment
FastOPD distills large vision-language-action models into few-step students that keep most performance at much lower latency.
FastOPD is a foundation-to-lightweight framework that distills large vision-language-action policies via on-policy distillation, using a flow map for single-state teacher supervision plus a self-consistency objective. On LIBERO it retains 84% of pi0.5 performance with two inference steps and cuts latency by 78.1%, beating prior few-step distillation baselines. With LingBot-VLA as teacher, single-step success rises 15.9 percentage points on RoboTwin 2.0, and a MolmoAct2 student was deployed on a real robot.
- FastOPD distills large VLAs with a flow map and a self-consistency objective.
- On LIBERO it retains 84% of pi0.5 success in two steps, cutting latency 78.1%.
- LingBot-VLA teacher lifts single-step success 15.9 points on RoboTwin 2.0.
- A compact student distilled from MolmoAct2 was deployed on a real robot.
Full article208 words · extracted from huggingface.co · click to collapse
Vision-Language-Action (VLA) foundation models have scaled rapidly to enhance manipulation performance and generalizability, but this scaling incurs high computational costs that render real-world deployment increasingly challenging. Existing approaches typically mitigate this issue by designing smaller architectures or reducing the iterative denoising steps in flow-based policies. In this work, we propose FastOPD, a foundation-to-lightweight VLA framework that enables the practical deployment of large-scale VLAs through efficient on-policy distillation. Specifically, FastOPD adapts a flow map for single-state teacher supervision and combines it with a self-consistency objective to construct a compact student that learns the teacher dynamics. Furthermore, we theoretically demonstrate that minimizing this objective allows the distilled student to recover a distribution on par with that induced by an ideal few-step teacher model. We evaluate FastOPD across diverse foundation policies in simulation and real-world experiments. On LIBERO, FastOPD retains 84% of the performance of π_{0.5} with only two inference steps, reducing inference latency by 78.1% while outperforming existing few-step distillation baselines in average success rate. With LingBot-VLA as the teacher, FastOPD improves the single-step success rate over the base student by 15.9 percentage points on RoboTwin 2.0. We further demonstrate its applicability to a World Action Model (WAM) and deploy a compact student distilled from MolmoAct2 on a real robot.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.02832