AuthGuard-R: Safety-Compliant Mission Hijacking and Dual-Gate Defense for LLM-Controlled Robots
Researchers define safety-compliant mission hijacking of LLM robot planners and show AuthGuard-R blocked all 109 unauthorized actions.
The paper defines safety-compliant mission hijacking, in which an LLM robot planner is steered into a physically safe action that still violates the authorized mission, such as changing a delivery target, object, region, sensor use, or timing. MissionPAIR searches for executable plans that pass a safety gate while breaking mission constraints. AuthGuard-R deterministically binds each action to a signed mission, robot identity, scope, state, time, and input provenance, alongside an independent safety gate. In 240 trials with Claude Haiku 4.5 and Qwen2.5-7B, planners followed injected deviations in 109 cases, and AuthGuard-R rejected all 109; eleven protocol and policy attacks were also blocked.
- Physically safe robot actions can still violate the authorized mission.
- MissionPAIR seeks plans that pass a safety gate while breaking mission constraints.
- AuthGuard-R binds each action to a signed mission, identity, scope, state, time, and provenance.
- In 240 trials, 109 deviations occurred and AuthGuard-R rejected every unauthorized action.
- Eleven protocol and policy attacks were also blocked completely.
Full article248 words · extracted from arxiv.org · click to collapse
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most defenses ask whether a proposed action is physically safe. This paper studies a different problem: an action may be physically safe and still violate the mission authorized by the user. An attacker may redirect a delivery robot, replace an approved object, extend a robot's operating region, activate an unnecessary sensor, or delay a mission without creating an immediate physical hazard. We call this attack \emph{safety-compliant mission hijacking}. We propose MissionPAIR, an adaptive attack framework that searches for executable plans that pass a safety gate while violating an authenticated mission. We also propose AuthGuard-R, a deterministic authorization layer that binds every executable action to a signed mission, robot identity, object and region scope, current state, time, and input provenance. AuthGuard-R operates with an independent safety gate, giving a dual-gate architecture. We formalize mission policies over robot traces, define security games, and prove authorization soundness, mission non-escalation, replay resistance, robot binding, provenance separation, threshold-approval security, audit-log tamper evidence, and trace-level composition. We report a preliminary cross-model evaluation with Claude Haiku~4.5 and the open-source Qwen2.5~7B planner. Across 240 live attack trials, the planners followed an injected mission deviation in 109 trials; AuthGuard-R rejected all 109 resulting unauthorized actions. A separate hand-constructed suite of eleven protocol- and policy-level attacks was also blocked completely.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31110