AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination
AgentLSD benchmark shows deceptive CTF artifacts like fake flags and decoy endpoints steer AI security agents wrong, inflating turns and tokens.
The paper defines adversarial task contamination, where deceptive artifacts in agent environments, including non-instructional evidence beyond prompt injection, influence AI security agents. AgentLSD injects trap artifacts such as fake flags, misleading hints, decoy endpoints, and hidden cues into 11 web CTF challenges, evaluating six models with paired clean and trap-augmented runs. Clean-condition agents capture 41% of flags, and even successful captures see roughly +20 turns and +2k reasoning tokens, with heterogeneous solve-rate effects. The framework, configurations, and traces are released.