TAPDreamer: Transferable Adversarial Patches for World Action Models
TAPDreamer crafts transferable adversarial patches that drop world-action robot success to near zero without querying policies.
Researchers introduce TAPDreamer, an attack that builds a fixed adversarial patch from a public visual encoder without querying the victim policy. The patch covers about 6.5% of the input and is trained on six frames from one source task so it transfers across tasks and action architectures. In closed-loop tests it cuts FastWAM success from 97.7% to 0.0% on 40 LIBERO tasks and from 90.8% to 0.0% on 50 RoboTwin tasks, while matched random patches retain about 80% success. The same patches reduce two DreamWAM configurations to 2.1% and 0.8% and Motus to 10.0%.
- Patch is built from a public encoder with no target-policy queries.
- One frozen patch covering about 6.5% of the image transfers across tasks.
- FastWAM success falls from 97.7% to 0% on 40 LIBERO tasks.
- The same patches drop DreamWAM to 2.1% and 0.8%, and Motus to 10%.
- Authors say shared visual encoders need defenses, not only action policies.
Full article242 words · extracted from arxiv.org · click to collapse
World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06814