Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults
Camera blackouts and freezes cause distinct hazardous motions in pi0.5 and GR00T robot policies.
Researchers compare vision-language-action models pi0.5 and GR00T under camera blackouts and frozen frames. Freezing produces more extreme joint behavior, while blackout after gripper closure causes more object drops, especially without proprioception. Proprioception reduces non-target contact but does not restore success when wrist-view object information is removed. Camera-blackout training and training-free embedding replacement improve selected success rates yet can increase unintended contact, including in real-robot trials.
- Blackout and freeze faults produce distinct physical failure modes.
- Freezing causes more extreme joint motion than blackout.
- Proprioception reduces non-target contact but not task failure.
- Mitigations raise success yet can disturb nearby objects.
- Real-robot trials still showed unintended physical interactions.
Full article183 words · extracted from arxiv.org · click to collapse
Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $π0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39145