60
30
55
57
42
55
42
The Next Model Won’t Fix This: Three AppSec Imperatives From Black Hat USA 2026
Checkmarx-funded study discussed around Black Hat USA 2026 found frontier models produce working code 83-95% of the time but only 24-36% of solutions meet the study's requirements.
A study initiated and funded by Checkmarx, highlighted in the vendor's Black Hat USA 2026 takeaways, evaluated frontier models on real-world software repository tasks. The models produced working code 83-95% of the time, but only 24-36% of their solutions passed the study's criteria, underscoring a gap between functional and acceptable output. Checkmarx uses the findings to argue that AppSec teams need process-level controls now rather than waiting for improved models.
30
30
30
55
30
30
60
55
55
55
30
55
30
30
30
55
55
55
60
57
55
55
55
55
30
30
55
60
57
57
30
30
57