Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC
Offline MFOTL monitoring flags about 70% of AgentDojo and STAC agent attacks but also 29% of benign runs.
Researchers evaluate metric first-order temporal logic as a reusable guardrail for tool-using LLM agents by replaying AgentDojo, STAC, and R-Judge trajectories through unmodified MonPoly, offline. Five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, while firing on 29.3% of benign runs. Missing approvals and timestamps reduce history-dependent rules to risky action types. Binding provenance to the lookup that produced it stops a one-line evasion that defeated naive checks on 94-99% of otherwise flagged runs, and the authors propose a twelve-field enforcement-ready trace schema.
- Five generic MFOTL obligations catch 71.8% of STAC attack chains.
- The same rules flag 70.1% of successful AgentDojo attacks.
- Benign runs fire 29.3% of the time because traces lack approvals and timestamps.
- Binding provenance to its lookup closes a 94-99% evasion.
- Authors propose a twelve-field enforcement-ready agent trace schema.
Full article197 words · extracted from arxiv.org · click to collapse
Guardrails for tool-using LLM agents are usually application-specific rules, which makes multi-step, data-dependent safety policies hard to specify, audit and reuse. As a declarative alternative, we evaluate metric first-order temporal logic (MFOTL), replaying the recorded trajectories that AgentDojo, STAC and R-Judge already ship through the unmodified MonPoly monitor, offline and without running an agent. On these corpora, five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs. This imprecision stems from the corpora rather than the logic: they rarely record approvals and never record timestamps, so history-dependent obligations reduce to detecting risky action types. Where the trace does carry relational context, provenance-aware policies discriminate better; that context, however, is itself attackable, and one planted line defeats a naive provenance check on 94-99% of the runs it would otherwise flag. Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. Taken together, these results show that formal temporal monitoring adds value exactly when the trace exposes trustworthy history. We therefore quantify how far current benchmarks are from that point and propose a twelve-field enforcement-ready trace schema.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.09793