Three papers report separate vision-language-action generalization gains
October 2026 papers separately report AE-VLA, DRIVE, and instruction-rephrasing gains for vision-language-action policies on different benchmarks.
Three October 2026 papers describe distinct ways to improve vision-language-action policies and do not contradict one another, so their success rates are not directly comparable. A Hugging Face daily-papers note dated 2026-10-04 describes AE-VLA, which pairs SkillLoRA adapters with arm-wise attention on a shared pi0.5 backbone. On ACG-Bench—23 task-condition pairs across eight families, 17 of them unseen compositions—AE-VLA reached 21.53% simulation success versus 2.94% for a single pi0.5 policy, 3.06% for MA-VLA, and 5.53% for two independent pi0.5 policies, and averaged 39% on SO101 robots across five unseen conditions versus 10% for the strongest baseline. An arXiv paper dated 2026-10-07 introduces DRIVE, which rewards diversity among successful behaviors during reinforcement-learning fine-tuning; across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, average out-of-domain performance rose by 5.3 points on pi0 and 2.0 points on pi0.5 versus vanilla RL fine-tuning, and dual-arm AgileX PiPER-X out-of-domain success rose from 64.1% to 73.3%. A second 2026-10-07 arXiv paper finds sharp language sensitivity: pi0.5 turned a LIBERO stove on 100% of the time for "switch on the stove" but only 2% for "switch on the hot plate," and a rephrase-augmented pi0 checkpoint still swung by up to 61 points. Distilling ten to twenty rephrasing rules and rewriting each instruction once improved frozen pi0 by 16–27% relative on twelve held-out tasks and lifted pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining; an oracle phrase search nearly closed a 21-point in-distribution versus out-of-distribution gap.
- A 2026-10-04 note describes AE-VLA (SkillLoRA plus arm-wise attention on pi0.5); on ACG-Bench (23 conditions, eight families, 17 unseen) simulation success was 21.53% versus 2.94% for one pi0.5 policy, 3.06% for MA-VLA, and 5.53% for two…
- On physical SO101 robots, AE-VLA averaged 39% success across five unseen conditions versus 10% for the strongest baseline.
- A 2026-10-07 paper introduces DRIVE, which rewards diversity among successful behaviors during RL fine-tuning; average out-of-domain gains versus vanilla RL were 5.3 points on pi0 and 2.0 points on pi0.5 across LIBERO-Plus, ManiSkill3, and…
- On a dual-arm AgileX PiPER-X platform, DRIVE raised average out-of-domain success from 64.1% to 73.3%.
- A second 2026-10-07 paper finds wording sensitivity: pi0.5 stove success was 100% for "switch on the stove" and 2% for "switch on the hot plate," and a rephrase-augmented pi0 checkpoint still swung by up to 61 points.
- Distilling 10–20 rephrasing rules and rewriting each instruction once improved frozen pi0 by 16–27% relative on twelve held-out tasks and lifted pi0.5 in-finetune LIBERO success from 93.6% to 97.8% without retraining; an oracle phrase…
- The three sources describe different methods and benchmarks and do not contradict one another, so the success rates are not directly comparable.
Coverage timelineoldest first · each row is one article
- · 4d agoArm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models
Hugging Face daily papers· 41
AE-VLA lifts dual-arm generalization to 21.53% in simulation and 39% on SO101 robots.
- · 1d agoMany Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
arXiv cs.AI / cs.LG / cs.CL· 34
DRIVE, a diversity-based RL fine-tune, lifts vision-language-action out-of-domain success, including +9.2 points on real dual-arm robots.
- · 1d agoRephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models
arXiv cs.AI / cs.LG / cs.CL· 46