Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Re-evaluation shows prompt-injection detector benchmark scores transfer poorly to real LLM agents; rankings on BIPIA, AgentDojo and tau-bench diverge sharply.
The authors replay ground-truth tool calls from AgentDojo and tau-bench without an LLM to build benign-by-construction tool outputs, then label injections via differential replay to evaluate fifteen detectors plus two task-aware LLM judges, including Meta's Prompt Guard 2. Detection rankings transfer poorly: the best BIPIA detector catches only 2% of AgentDojo injections at a 1% false-positive rate, and a detector catching 72% of AgentDojo injections catches just 15% on tau-bench. False-positive rates on tool outputs range from none to over 90% and do transfer, and training-data form explains results: the detector best on both agent benchmarks was trained on agent-style inputs sharing no benchmark data. They recommend deployment evaluations use the agent's own tool outputs, report detection at low false-positive rates, and audit detector training data.