Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models
Kinematic MeanFlow enables one-step robot actions and cuts GR00T-N1.6 action-head latency by up to 74%.
Kinematic MeanFlow (K-MF) enables one-step action generation for robotic foundation models, avoiding the latency of multi-step flow matching. The method splits MeanFlow's time-derivative term at an intermediate point so early and late denoising dynamics are modeled separately, preventing the collapse seen when MeanFlow is applied directly. Across scratch training and fine-tuning, K-MF outperformed multi-step flow matching in most settings and cut GR00T-N1.6 action-head latency by 67.5% to 74.4% on L40 and Jetson Orin, with end-to-end reductions of 30.3% to 54.9%.
- Direct MeanFlow collapses on robotic foundation models because of late velocity-field spikes.
- K-MF splits the time derivative across an intermediate point using a kinematic identity.
- One-step policies work for training from scratch and fine-tuning and beat multi-step flow matching in most settings.
- GR00T-N1.6 action-head latency fell 67.5% to 74.4% on L40 and Jetson Orin.
Full article220 words · extracted from huggingface.co · click to collapse
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.00864