The Next Model Won’t Fix This: Three AppSec Imperatives From Black Hat USA 2026
Checkmarx-funded study discussed around Black Hat USA 2026 found frontier models produce working code 83-95% of the time but only 24-36% of solutions meet the study's requirements.
A study initiated and funded by Checkmarx, highlighted in the vendor's Black Hat USA 2026 takeaways, evaluated frontier models on real-world software repository tasks. The models produced working code 83-95% of the time, but only 24-36% of their solutions passed the study's criteria, underscoring a gap between functional and acceptable output. Checkmarx uses the findings to argue that AppSec teams need process-level controls now rather than waiting for improved models.
- Frontier models succeeded functionally on 83-95% of real-world repository coding tasks
- Only 24-36% of solutions met the study's evaluation criteria, per Checkmarx
- Findings framed as AppSec imperatives: better models alone will not close the AI code risk gap
- Piece reflects how AI security concerns became operational topics at Black Hat USA 2026
Black Hat USA 2026 didn’t introduce the AI security problem, but made it clear how fast that problem is becoming operational. An independent study initiated and funded by Checkmarx helps explain why. On real-world repository tasks, frontier models produced working code 83% to 95% of the time, but only 24% to 36% of their solutions […]
This source does not provide full text. Read it at checkmarx.com.