Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Agent-security benchmark ASR scores shift up to 13 percentage points when tool names change, showing single-representation robustness scores may not generalize.
The paper introduces threat-preserving representation sensitivity (TPRS), measuring how attack success rate changes when only the agent-visible representation changes. On Agent Security Bench, replacing threat-related tool names with neutral ones raises committed ASR by 11.67 points on GPT-5-mini and 13.21 points on Claude Haiku 4.5; on MCPTox the opposite substitution lowers ASR by 11.00 and 4.11 points respectively. A token-count-matched neutral name reproduces most of the MCPTox shift (8.54 of 11.00 points on GPT-5-mini), suggesting models respond to surface cues rather than threat semantics. The authors argue robustness claims should be validated across a controlled set of threat-preserving representations rather than a single score.