Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
One-word instruction edits swing VLA success sharply; rule-based rephrasing improves frozen policies without retraining.
The paper shows vision-language-action models remain highly sensitive to instruction wording even when built on vision-language models. Pi0.5 turns on a LIBERO stove 100% of the time for "switch on the stove" but only 2% for "switch on the hot plate," and a rephrase-augmented pi0 checkpoint still swings by up to 61 points. An oracle phrase search nearly closes a 21-point gap between in-distribution and out-of-distribution tasks. Distilling ten to twenty rephrasing rules and rewriting each instruction once improves frozen pi0 by 16–27% relative on twelve held-out tasks and lifts pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining.
- One-word edits can swing VLA success by tens of points.
- Pi0.5 stove success falls from 100% to 2% under a synonym.
- An LLM distills 10–20 rephrasing rules from scored training phrases.
- Frozen pi0 improves 16–27% relative; pi0.5 LIBERO success rises to 97.8%.
Full article217 words · extracted from arxiv.org · click to collapse
Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $π_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $π_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen $π_0$ by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on $π_{0.5}$ and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: https://sttawm.github.io/rephrase-before-you-act
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.10526