Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC
Offline MFOTL monitoring flags about 70% of AgentDojo and STAC agent attacks but also 29% of benign runs.
Researchers evaluate metric first-order temporal logic as a reusable guardrail for tool-using LLM agents by replaying AgentDojo, STAC, and R-Judge trajectories through unmodified MonPoly, offline. Five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, while firing on 29.3% of benign runs. Missing approvals and timestamps reduce history-dependent rules to risky action types. Binding provenance to the lookup that produced it stops a one-line evasion that defeated naive checks on 94-99% of otherwise flagged runs, and the authors propose a twelve-field enforcement-ready trace schema.