Approval Integrity and Recovery in LLM Answer Publication
Study measures approval integrity in Lightcap LLM answer publication, finding the 14B response-act checker accepts 291 of 302 unsupported answers.
The study evaluates exact-content binding, authorization freshness, and checkpoint recovery in Lightcap's publication enforcement using 3,600 assessments over 900 human-annotated RAGTruth responses from three Ministral models. The production 14B response-act checker accepts 291 of 302 unsupported answers versus 41 for a direct-grounding baseline, with supported-answer retention of 95.2% versus 66.9%. A stateful recheck-recovery policy increases exact-match error by 9.23 percentage points relative to initial checkpoints, and controlled evidence-fingerprint changes expose asymmetric freshness enforcement between publication and recovery. A separate BIPIA prompt-injection experiment records zero target insertions among 266 valid editor outputs.
- Production 14B checker accepts 291 of 302 unsupported answers; direct baseline accepts 41
- Supported-answer retention 95.2% production versus 66.9% direct baseline
- Recheck-recovery policy adds 9.23 percentage points exact-match error
- Asymmetric authorization freshness exposed between publication and recovery paths
- BIPIA prompt-injection test: zero target insertions in 266 editor outputs
Full article182 words · extracted from arxiv.org · click to collapse
Publication integrity in LLM systems requires binding approved content to its current authorization context. We examine exact-content binding, authorization freshness and checkpoint recovery in Lightcap's publication enforcement mechanism. On 900 independently human-annotated RAGTruth responses from 150 source tasks, three dated Ministral models and a same-model direct-grounding baseline yield 3,600 assessments. The production response-act checker instantiated with 14B accepts 291 of 302 unsupported-labelled answers; the direct baseline accepts 41. Supported-answer retention is 95.2% and 66.9%, respectively. An exact promotion-correction identity tracks error through 100 chronological 3B-14B-8B-14B answer trajectories. Among 65 initially approved answers, the final stateful recheck-recovery policy increases exact-match error by 9.23 percentage points relative to the initial checkpoint (95% article-clustered interval [-1.72, 19.61]). Controlled evidence-fingerprint changes expose asymmetric freshness enforcement between publication and recovery. A separate BIPIA prompt-injection experiment records zero target insertions among 266 valid editor outputs. External Hugging Face calibration experiments transfer retrieval models from ArguAna to SciFact and NFCorpus, and diagnostic decision rules from Thunderbird to BGL, distinguishing probability calibration from ranking changes. The measurements separate semantic false approval, stale authorization and recovery-induced error at executable publication boundaries.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.15576