Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
Re-evaluation shows prompt-injection detector benchmark scores transfer poorly to real LLM agents; rankings on BIPIA, AgentDojo and tau-bench diverge sharply.
The authors replay ground-truth tool calls from AgentDojo and tau-bench without an LLM to build benign-by-construction tool outputs, then label injections via differential replay to evaluate fifteen detectors plus two task-aware LLM judges, including Meta's Prompt Guard 2. Detection rankings transfer poorly: the best BIPIA detector catches only 2% of AgentDojo injections at a 1% false-positive rate, and a detector catching 72% of AgentDojo injections catches just 15% on tau-bench. False-positive rates on tool outputs range from none to over 90% and do transfer, and training-data form explains results: the detector best on both agent benchmarks was trained on agent-style inputs sharing no benchmark data. They recommend deployment evaluations use the agent's own tool outputs, report detection at low false-positive rates, and audit detector training data.
- Detector rankings on BIPIIA do not predict performance inside AgentDojo or tau-bench agents
- Best BIPIA detector catches only 2% of AgentDojo injections at 1% false-positive rate
- False-positive rates on tool outputs range from 0% to over 90% and transfer across agent benchmarks
- Detectors trained on agent-style inputs generalize best; benchmark-trained detectors overfit
- Deployment evaluations should use the agent's own tool outputs and audit detector training data
Full article227 words · extracted from arxiv.org · click to collapse
LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a 1% false-positive rate, and a detector that catches 72% of AgentDojo injections catches 15% on tau-bench. False-positive rates on tool outputs, which range from none to over 90%, do transfer between the two agent benchmarks. Where training data is public, the form of the training inputs explains the results. The BIPIA leader was trained on full BIPIA inputs, but having seen InjecAgent's attack strings as short prompts does not help it find them inside tool outputs; the best detector on both agent benchmarks shares no data with any benchmark and was trained on agent-style inputs. Evaluations meant to inform deployment should use the agent's own tool outputs, report detection at a low false-positive rate, and audit what the detector was trained on.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03448