Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
One-word instruction edits swing VLA success sharply; rule-based rephrasing improves frozen policies without retraining.
The paper shows vision-language-action models remain highly sensitive to instruction wording even when built on vision-language models. Pi0.5 turns on a LIBERO stove 100% of the time for "switch on the stove" but only 2% for "switch on the hot plate," and a rephrase-augmented pi0 checkpoint still swings by up to 61 points. An oracle phrase search nearly closes a 21-point gap between in-distribution and out-of-distribution tasks. Distilling ten to twenty rephrasing rules and rewriting each instruction once improves frozen pi0 by 16–27% relative on twelve held-out tasks and lifts pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining.