Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
Scoped authorization stopped harmful LLM-agent tool execution that broad bearer tokens still allowed.
The paper tests whether task-scoped authorization contains tool use after a model follows a malicious instruction. A paired-replay testbed resubmits the same action and arguments under broad bearer, scoped JWT, sender-constrained, and Open Policy Agent conditions across 128 scenarios and five local model configurations. Among valid attacked decisions, harmful execution under broad bearer tokens ranged from 8.9% to 37.8%, while all three scoped conditions recorded zero. A bounded AgentDojo extension recorded 11 of 24 broad-policy attack flags versus 0 of 24 scoped flags. The authors frame this as containment, not prevention of prompt injection.
- Paired replay compared broad bearer, scoped JWT, sender-constrained, and OPA checks.
- Broad-bearer harmful execution ranged from 8.9% to 37.8%; scoped conditions were zero.
- Policy changed execution, not the frozen model decision.
- AgentDojo extension recorded 11/24 broad-policy flags versus 0/24 scoped flags.
Full article198 words · extracted from arxiv.org · click to collapse
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-constrained, and Open Policy Agent conditions. The frozen main experiment uses 128 scenarios across four tool domains and five local model configurations. Among valid attacked post-exposure decisions, broad-bearer harmful execution ranges from 8.9% to 37.8% across models; all three scoped conditions record zero. The policy changes execution, not the frozen model decision. These results support containment of the tested cross-action and cross-resource consequences under researcher-supplied task authority, not prevention of prompt injection. A bounded six-task AgentDojo extension preserves native scoring and records 11/24 injected broad-policy attack flags versus 0/24 scoped flags, with utility of 5/24 and 6/24. Low exposure, invalid decisions, and purposive task selection limit that comparison. The primary experiment measures safe continuation rather than final-answer correctness. An empirical-bank fresh-sampling comparison finds no estimation advantage from pairing when scoped outcomes are constant zero. The contribution is a controlled measurement of compromise, enforcement, and continuation, with explicit limits on what each observation establishes.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.05840