Agent benchmark scores shift; scoped auth contains tool harm
Tool renaming moved agent-security ASR by up to about 13 points, while scoped authorization blocked harmful executions that broad tokens still allowed.
Two October 2026 papers examine different parts of LLM-agent security. The 2 October study introduces threat-preserving representation sensitivity and finds that renaming tools alone moves benchmark attack success: on Agent Security Bench, neutral names raised committed ASR by 11.67 points on GPT-5-mini and 13.21 points on Claude Haiku 4.5, while on MCPTox the reverse substitution lowered ASR by 11.00 and 4.11 points. A token-count-matched neutral name accounted for 8.54 of the 11.00-point GPT-5-mini MCPTox shift, which the authors read as sensitivity to surface cues; on AgentDojo they report little ASR change but a 5.36-point benign-utility drop from threat wording, and they argue single-representation robustness scores may not generalize. The 5 October study uses paired replay of the same action and arguments across 128 scenarios and five local model configurations. Among valid attacked decisions, harmful execution under broad bearer tokens ranged from 8.9% to 37.8%, while scoped JWT, sender-constrained, and Open Policy Agent conditions recorded zero, and an AgentDojo extension showed 11 of 24 broad-policy flags versus 0 of 24 scoped flags. The papers do not disagree: one questions benchmark score stability under equivalent wording, and the other treats scoped authorization as containment of tool execution rather than prevention of prompt injection.
- A 2026-10-02 paper defines threat-preserving representation sensitivity (TPRS) as the change in attack success rate when only the agent-visible representation changes.
- On Agent Security Bench, replacing threat-related tool names with neutral ones raised committed ASR by 11.67 points on GPT-5-mini and 13.21 points on Claude Haiku 4.5.
- On MCPTox the opposite substitution lowered ASR by 11.00 points (GPT-5-mini) and 4.11 points (Claude Haiku 4.5); a token-count-matched neutral name reproduced 8.54 of the 11.00-point GPT-5-mini shift.
- On AgentDojo, threat wording produced little ASR change but a 5.36-point drop in benign utility.
- A 2026-10-05 paired-replay study covered 128 scenarios and five local model configurations under broad bearer, scoped JWT, sender-constrained, and Open Policy Agent checks.
- Among valid attacked decisions, harmful execution was 8.9% to 37.8% under broad bearer tokens and zero under all three scoped conditions; policy changed execution, not the frozen model decision.
- A bounded AgentDojo extension recorded 11 of 24 broad-policy attack flags versus 0 of 24 scoped flags; authors frame the result as containment, not prevention of prompt injection.
Coverage timelineoldest first · each row is one article
- · 6d agoThreat-Preserving Representation Sensitivity in Agent-Security Benchmarks
arXiv cs.CR· 50
Agent-security benchmark ASR scores shift up to 13 percentage points when tool names change, showing single-representation robustness scores may not generalize.
- · 3d agoCompromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
arXiv cs.CR· 56
Scoped authorization stopped harmful LLM-agent tool execution that broad bearer tokens still allowed.