Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Agent-security benchmark ASR scores shift up to 13 percentage points when tool names change, showing single-representation robustness scores may not generalize.
The paper introduces threat-preserving representation sensitivity (TPRS), measuring how attack success rate changes when only the agent-visible representation changes. On Agent Security Bench, replacing threat-related tool names with neutral ones raises committed ASR by 11.67 points on GPT-5-mini and 13.21 points on Claude Haiku 4.5; on MCPTox the opposite substitution lowers ASR by 11.00 and 4.11 points respectively. A token-count-matched neutral name reproduces most of the MCPTox shift (8.54 of 11.00 points on GPT-5-mini), suggesting models respond to surface cues rather than threat semantics. The authors argue robustness claims should be validated across a controlled set of threat-preserving representations rather than a single score.
- TPRS quantifies ASR change under threat-preserving changes to agent-visible representation
- ASR shifts of 11-13 points on GPT-5-mini and Claude Haiku 4.5 from tool renaming alone
- Token-count-matched neutral names reproduce most of the effect, implicating surface features
- AgentDojo shows tiny ASR change but 5.36-point benign utility drop from threat wording
- Single-representation benchmark scores may not generalize across equivalent security problems
Full article268 words · extracted from arxiv.org · click to collapse
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03585