ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
ToolFence authorizes tool-using LLM agents before execution, driving AgentDojo attack success near zero on Qwen3-max.
ToolFence defends tool-using LLM agents against indirect prompt injection, including within-tool attacks that keep the intended tool but manipulate its arguments. It compiles a typed authorization blueprint before execution, enforces it with a deterministic monitor, and asks a judge for new capabilities only when the blueprint is incomplete. Provenance-aware checks separate user-authorized values from untrusted observations. On AgentDojo with Qwen3-max, overall attack success falls to near zero, clean utility drops 3.80 percentage points, and runtime overhead stays practical compared with slower data-flow controls such as CaMeL.
- Typed blueprint and deterministic monitor authorize tool effects
- Addresses within-tool attacks that alter arguments but keep the tool
- AgentDojo attack success near zero with Qwen3-max
- Clean utility falls 3.80 percentage points with limited judge calls
Full article181 words · extracted from arxiv.org · click to collapse
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.37196