Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies
Researchers release SecTB-RTL, an auditable framework showing that AI-generated RTL verification plans passing provider schema checks can still fail execution validity.
The paper presents SecTB-RTL, an auditing framework covering 31 tasks and 124 authored hardware-security regressions for AI-generated RTL verification plans. A confirmatory run of 1,860 calls was accepted by the provider (1,857 responses) but only nine passed the production semantic validator, leading the authors to record it as an instrument-validation incident with no prompt-effect estimate. They release the benchmark, failure-preserving contract, incident provenance, and governance controls separating infrastructure behavior from model behavior.
- Framework spans 31 tasks and 124 authored hardware-security regressions
- Provider schema acceptance did not match production execution rules
- 1,860 calls made; only 9 responses passed the semantic validator
- Benchmark, incident provenance, and governance controls are released
Full article185 words · extracted from arxiv.org · click to collapse
AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardware-security regressions. A deterministic non-AI baseline killed 36, 75, and 78 mutants at increasing resource limits. The first confirmatory run (C1-R2) failed before model execution because the provider rejected its response schema. After a schema-only repair made without viewing outcomes, a separately frozen follow-up run (C1-R3) completed 1,860 calls. The provider accepted 1,857 responses, but only nine passed the production semantic validator. The generation and execution rules did not match. We therefore preserve the run as an instrument-validation incident and report no prompt-effect estimate. This incident shows that provider or schema acceptance does not establish execution validity. Compilation and coverage are only diagnostics; the exact saved artifact must pass the full production path. A subsequent follow-up is excluded because it did not satisfy the preregistered evidence-completeness gate and is treated only as future work. We release the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.19844