AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
AdvSim2Real co-trains web agents against adaptive prompt injection, lifting completion 33.6% under an unseen adversary.
AdvSim2Real jointly evolves a task curriculum, an injection adversary, and a web agent inside a frozen web world model so defenses are not limited to injections fixed before training. The curriculum favors tasks the agent solves about half the time, and the adversary is rewarded only when an injection turns a judged success into a failure. A 4B agent trained this way improves completion with and without attacks, holds up against a frontier-model adversary it never trained against, and transfers capability gains to a real browser. On 150 web tasks, completion under that unseen adversary rises 33.6% relative to the base agent.
- Fixed-injection fine-tuning is bypassed when attackers adapt to the trained model.
- A task curriculum, injection adversary, and agent co-evolve in a frozen web simulator.
- The adversary is rewarded only for flipping a judged success into failure.
- On 150 web tasks, completion under an unseen adversary rises 33.6% relatively.
- Capability gains from simulator training carry over to a real browser.
Full article201 words · extracted from huggingface.co · click to collapse
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.08773