SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities
SecProbe adaptively tests coding agents on synthesized vulnerability repairs; frontier models peak at 28.33% success.
SecProbe combines item response theory with on-demand synthesis of repository-scale vulnerability-repair tasks to estimate coding-agent ability and choose the most informative next task. A case study builds 353 tasks spanning six programming languages and 151 CWE types and evaluates nine frontier models with two agent harnesses. Peak success is 28.33%. Compared with random and one-shot baselines, ability estimates stay comparable while agents solve up to 29.5% fewer tasks.
- Constructs 353 repository-scale tasks across six languages and 151 CWE types.
- Nine frontier models with two harnesses peak at 28.33% success.
- Matches baseline ability estimates while solving up to 29.5% fewer tasks.
Full article159 words · extracted from arxiv.org · click to collapse
Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce SecProbe, a framework for adaptive evaluation that combines Item Response Theory (IRT) with on-demand synthesis of repository-scale vulnerability-repair tasks. From observed performance, \textsc{SecProbe} estimates agent ability and identifies where additional evidence is most informative, selecting existing tasks or synthesizing new ones accordingly. As one use case, we construct 353 tasks spanning six programming languages and 151 CWE types and evaluate nine frontier models with two agent harnesses. Success rates peak at 28.33\%, highlighting substantial gaps in vulnerability recognition and repair. Compared with random and one-shot baselines, \textsc{SecProbe} achieves comparable agent ability estimates while requiring agents to solve up to 29.5\% fewer tasks. These results support adaptive evaluation as an efficient and discriminative approach to assessing cybersecurity vulnerability awareness.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33763