AdvSim2Real trains web agents against adaptive prompt injection
AdvSim2Real co-trains a 4B web agent against adaptive prompt injection, lifting completion 33.6% under an unseen adversary.
AdvSim2Real jointly trains a web agent against adaptive prompt injection by co-evolving a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum emphasizes tasks the agent solves about half the time, and the adversary is rewarded only when an injection flips a judged success into a failure. A 4B agent trained this way improved task completion both with and without attacks and held up against a frontier-model adversary it never saw during training. On 150 web tasks, completion under that unseen adversary rose 33.6% relative to the base agent, and capability gains carried over to a real browser. Hugging Face daily papers and arXiv agree on these figures; the papers also note that fine-tuning against fixed injections can be bypassed once attackers adapt.
- AdvSim2Real co-evolves a task curriculum, an adaptive prompt-injection adversary, and a web agent inside a frozen web world model so defenses are not limited to injections fixed before training.
- The curriculum favors tasks the agent solves about half the time; the adversary is rewarded only when an injection turns a judged success into a failure.
- A 4B agent trained this way improved completion both with and without attacks.
- On 150 web tasks, completion under an unseen frontier-model adversary rose 33.6% relative to the base agent.
- Capability and robustness gains from simulator training transferred to a real browser.
- Sources note that fixed-injection fine-tuning can be bypassed when attackers adapt to the trained model.
- Both Hugging Face daily papers and arXiv (cs.AI / cs.LG / cs.CL) carried the same result on 2026-10-06.
Coverage timelineoldest first · each row is one article
- · 2d agoAdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Hugging Face daily papers· 55
AdvSim2Real co-trains web agents against adaptive prompt injection, lifting completion 33.6% under an unseen adversary.
- · 2d agoAdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
arXiv cs.AI / cs.LG / cs.CL· 58
AdvSim2Real co-trains a 4B web agent against adaptive prompt injection, lifting completion 33.6% under an unseen adversary.