ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Vivek Chavan

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

infoAI researchimportance 18
AI summary · glm-5.3-flash

Researchers diagnose conditional visual grounding failures in visuomotor imitation policies and show targeted interventions substantially improve distractor robustness.

The paper studies why ACT-based visuomotor imitation policies fail when visually similar distractor objects or receptacles are introduced, finding sensitivity depends on both distractor type and manipulation stage. Interventions including distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting improve target selection while preserving spatial control information, with gains in simulation and on a physical UR3e. The same failure pattern is confirmed in a pretrained vision-language-action policy on a state-conditioned medical instrument-handling task.

  • Distractor sensitivity varies with visual similarity type and manipulation stage
  • Three complementary interventions improve target selection under distractors
  • Failure pattern reproduced in a pretrained vision-language-action policy
  • Robustness gains validated in simulation and on a physical UR3e robot
ProductsUR3e
AI modelsACT
Full article205 words · extracted from arxiv.org · click to collapse

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.05376