Trajectory-Level Security Debt in LLM Coding Agents
A new metric tracks static-analysis security debt across intermediate states of LLM coding agents, not just final code.
The authors define the Security Debt Line Integral (SDLI), which accumulates static-analysis risk as an LLM coding agent reaches a new best test-pass ratio. They apply four SAST tools to 830 passing SWE-bench runs, 712 ProgramBench workspaces, and 13 public MirrorCode trajectories. Two-tool CWE agreement is only 3.9% of SWE-bench runs and 26.2% of 80 ProgramBench runs that pass at least 90% of tests; excluding three advisory-heavy classes drops the latter to 6.2%. The authors stress these are scanner findings, not validated vulnerabilities, and that SDLI's use for steering agents remains unproven.
- SDLI accumulates SAST risk when an agent reaches a new best test-pass ratio.
- Study covers 830 SWE-bench runs, 712 ProgramBench workspaces, and 13 MirrorCode trajectories.
- Two-tool CWE agreement is 3.9% on SWE-bench and 26.2% on strong ProgramBench runs.
- Reported findings are scanner signals, not confirmed exploitable vulnerabilities.
- A repair case lowers scanner signal but is sensitive to equivalent API rewrites.
Full article196 words · extracted from arxiv.org · click to collapse
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral (SDLI), which accumulates static-analysis risk when an agent reaches a new best test pass ratio. We instantiate it with four static application security testing (SAST) tools and study artifacts from 830 passing SWE-bench runs, 712 ProgramBench final workspaces, and 13 public MirrorCode trajectories. The two large populations use the final-state special case of SDLI. Two-tool Common Weakness Enumeration (CWE) class agreement occurs in 3.9% of SWE-bench runs and 26.2% of the 80 ProgramBench runs passing at least 90% of official tests. These are scanner findings, not validated vulnerability rates. Excluding three advisory-heavy classes reduces the latter rate to 6.2%. Same-task runs differ in their measured scores, while one reconstructed ProgramBench run exposes persistent findings from its first implementation write. A repair case study reduces the scanner signal while preserving tested behavior, but also reveals sensitivity to equivalent API rewrites. SDLI offers a way to study progress and security findings together. Its value for steering agents and confirming exploitable vulnerabilities remains to be established.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35199