Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates
Researchers propose TAIP to continuously recompute assurance for LLM repository auditors used as CI merge gates.
LLM repository auditors increasingly admit or block changes in CI, but nondeterministic behavior and shifting policy make point-in-time audits insufficient. The authors propose Policy-Evidence-Execution Separation, implemented by the Trustworthy AI Posture (TAIP) engine as Continuous Control Posture Assurance. Using unmodified RepoAudit on a fixed Python null-pointer benchmark, they retain 80 executions across gpt-4o-mini and gpt-4.1. Peak policy-to-posture latency was 1.1 ms, and refreshing 1,000 decision-gateway contexts took at most 1.62 seconds, excluding RepoAudit execution and provider inference.
- TAIP separates policy from execution and binds evidence to a versioned Posture Tree.
- Evaluation retained 80 RepoAudit runs across gpt-4o-mini and gpt-4.1.
- Maximum policy-to-posture latency was 1.1 ms in one execution.
- Refreshing 1,000 contexts took at most 1.62 seconds against a 5-second budget.
Full article221 words · extracted from arxiv.org · click to collapse
Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls, their non-deterministic behaviour changes the evidence, while organisational risk appetite and jurisdictional or data-sovereignty policy change its interpretation. Point-in-time audits therefore cannot maintain current assurance for merge decisions. We propose the Policy-Evidence-Execution Separation Pattern, implemented by the Trustworthy AI Posture (TAIP) Assurance Engine and operated as Continuous Control Posture Assurance (CCPA). By separating policy from stable execution and binding admitted evidence to a versioned Posture Tree, the same assurance logic operates across models, environments, and policy profiles. We evaluate the approach using unmodified RepoAudit on a fixed Python Null Pointer Dereference benchmark. The retained evidence repository contains 80 RepoAudit executions across two OpenAI model configurations, gpt-4o-mini and gpt-4.1. TAIP recomputes assurance posture after policy, evidence, and model-context changes and is evaluated across increasing numbers of independent Decision Gateway contexts. The maximum observed policy-to-posture latency was 1.1 ms across three policy-class cycles in one execution. At 1,000 independent assurance contexts, full policy-triggered recomputation with one worker recorded a maximum aggregate refresh of 1.62 s, below the predeclared 5 s Decision Gateway budget. These single-host measurements concern assurance over retained evidence and exclude RepoAudit execution and provider inference.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35266