VLA papers certify speedups and cut wording sensitivity
CARE certifies 9.0-10.8x VLA speedups under a failure budget, while rephrasing rules lift frozen policies after wording swings success.
Two independent vision-language-action papers, dated 2026-10-05 and 2026-10-07, address different problems and do not contradict each other. CARE selects the fastest accelerator that stays inside a user-specified failure budget, using paired closed-loop rollouts for finite-sample guarantees. On four LIBERO suites with OpenVLA-OFT it certifies 9.0-10.8x speedups while preserving at least 85.8% of reference-solved episodes at 95% confidence; sequential testing uses 78.9% fewer rollouts, and the method also covers flow-step reduction for pi0.5 plus Qwen3.5-9B and Llama-3.1-8B agents in Crafter. Selectors without guarantees exceed tight budgets in up to 75% of trials. Separately, one-word instruction edits swing success sharply: pi0.5 turns on a LIBERO stove 100% of the time for "switch on the stove" but only 2% for "switch on the hot plate," and a rephrase-augmented pi0 checkpoint still swings by up to 61 points. Distilling ten to twenty rephrasing rules and rewriting each instruction once improves frozen pi0 by 16–27% relative on twelve held-out tasks and lifts pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining; an oracle phrase search nearly closes a 21-point gap between in-distribution and out-of-distribution tasks.
- CARE certifies 9.0-10.8x speedups on four LIBERO suites with OpenVLA-OFT while preserving at least 85.8% of reference-solved episodes at 95% confidence.
- CARE uses paired closed-loop rollouts for finite-sample guarantees; its sequential form uses 78.9% fewer rollouts than exhaustive evaluation.
- Selectors without guarantees exceed tight failure budgets in up to 75% of trials.
- CARE also covers flow-step reduction for pi0.5 and Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
- Pi0.5 turns a LIBERO stove on 100% of the time for "switch on the stove" but only 2% for "switch on the hot plate"; a rephrase-augmented pi0 checkpoint still swings by up to 61 points.
- An oracle phrase search nearly closes a 21-point in-distribution versus out-of-distribution gap.
- Distilling 10–20 rephrasing rules and rewriting each instruction once improves frozen pi0 by 16–27% relative on twelve held-out tasks and lifts pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining.
Coverage timelineoldest first · each row is one article
- · 5d agoCARE: Certifying Acceleration for Vision-Language-Action Inference
Hugging Face daily papers· 48
CARE certifies VLA inference accelerators so speedups stay inside a user-specified failure budget.
- · 3d agoRephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
arXiv cs.AI / cs.LG / cs.CL· 46
One-word instruction edits swing VLA success sharply; rule-based rephrasing improves frozen policies without retraining.