AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
AgentXploit audits AI-agent repos into confirmed runtime attacks, reaching 59.3% success on 72 benchmark bugs.
AgentXploit is a two-role white-box auditor that separates repository attack-path discovery from runtime exploitation of AI agents, with attacks required to use the task-defined interface and pass an external verifier. AgentXploit-Bench contains 72 reproducible vulnerabilities across 12 open-source agent systems. Across three runs it reaches 59.3% end-to-end success, compared with 38.4% for Codex and 46.3% under a matched token budget. On AgentDojo, where injection points are provided, the Exploiter reaches 79.2% versus 52.7% for AgentVigil.
- Analyzer discovers code-supported paths; Exploiter confirms attacks at runtime.
- Benchmark covers 72 vulnerabilities in 12 open-source AI-agent systems.
- End-to-end success is 59.3%, versus 38.4% for Codex.
- Token-budget-matched Codex reaches 46.3%.
- On AgentDojo, Exploiter hits 79.2% versus 52.7% for AgentVigil.
Full article196 words · extracted from arxiv.org · click to collapse
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31318