ZeroHour
arXiv cs.CRpublished ()ingested Sizhe Chen

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

infoAI safety & securityimportance 55
AI summary · glm-5.3-flash

Researchers unveil Repeat-After-Me, a black-box visual prompt injection achieving over 80% success on Qwen3.6-27B and 47% on GPT-5.5.

Researchers present Repeat-After-Me, a black-box adaptive visual prompt injection that induces frontier VLMs to reveal PII or make malicious tool calls via injected images. It exceeds 80% attack success rate on Qwen3.6-27B and 47% on GPT-5.5 even when the benign user prompt is unrelated and does not authorize the injected task. In a real-world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling later remote code execution and secret exfiltration.

  • Black-box visual injection reaches 80%+ ASR on open-weight and 47% on commercial frontier VLMs
  • Attack triggers PII disclosure or precise, parseable malicious tool calls
  • Optimized injections retain 43-46% ASR across commercial victims; cross-sample transfer retains 64-66%
  • Demonstrated on OpenClaw: injected image overwrites TOOLS.md, enabling RCE and secret exfiltration
  • Works where adaptive textual prompt injection fails; potential defenses discussed
Full article247 words · extracted from arxiv.org · click to collapse

Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success rates (ASRs). In the image domain, however, existing visual prompt injection methods are substantially less effective in attacking frontier commercial VLMs for materially harmful behavior. Achieving such outputs is hard because it requires a long and/or format-compliant target string, such as a precise, parseable native tool call with exact function names and arguments. We present Repeat-After-Me, a black-box adaptive visual prompt injection attack that can reveal personally identifiable information or make malicious tool calls. Across both open-weight and commercial frontier VLMs, including Qwen3.6-27B and GPT-5.5, our method achieves ASRs exceeding 80% and 47%, respectively, under a realistic setting in which the benign user prompt is semantically unrelated to the injected task and does not verbally authorize it. In our evaluation, injections optimized on one surrogate retain 43-46% of the original ASR on two commercial victims, and cross-sample transferability retains 64-66% of the original ASR on those two models. We test our attack in a real-world OpenClaw agent: in a default OpenClaw Discord deployment, an untrusted user can use a minimally injected image to overwrite TOOLS.md, enabling future sensitive behaviors like remote code execution and secret exfiltration. We show our new attack vector works in cases where adaptive textual prompt injection fails. We discuss potential defenses.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.04533