COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints
COGNIT-Guard cascades a CPU gatekeeper to a 322M model for low-latency prompt safety screening.
COGNIT-Guard couples a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to Laya-322M, a 322M bidirectional direct-decision safety model, under an asymmetric false-positive penalty. On the unseen DUCS-Bench split (N=607) it reports 98.85% accuracy, 0.42% benign false-positive rate (1/238), 1.12% ECE, and a 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference averages 21.77 ms, while the deployed CPU-NPU cascade averages 41.63 ms at 99.23% accuracy and 0.00% FPR. Experience replay restores SafetyBench-ZH out-of-domain accuracy to 64.10-65.05%.
- CPU gatekeeper escalates uncertain prompts to NPU-resident Laya-322M.
- Unseen DUCS-Bench: 98.85% accuracy, 0.42% benign FPR, 1.12% ECE.
- Live cascade mean latency is 41.63 ms at 99.23% accuracy and 0% FPR.
- Experience replay lifts SafetyBench-ZH out-of-domain accuracy to 64-65%.
Full article235 words · extracted from arxiv.org · click to collapse
When must a foundation-model safety gateway generate tokens, and when should it directly output a calibrated decision? We study calibrated standalone direct-decision foundation models for real-time pre-ingestion safety guardrails, jointly addressing probability calibration, dual-use false-positive control, and heterogeneous CPU-NPU routing under explicit latency SLOs. Pre-ingestion guardrails must screen prompts prior to target-LLM prefill with low false alarms on benign compliance inquiries; however, shallow classifiers are brittle to phrasing shifts, hidden-state probes require coupling to a target LLM, and generative guards incur high decoding latency and dual-use false positives. We present COGNIT-Guard, coupling a validation-calibrated CPU fast gatekeeper with confidence-gated escalation to an NPU-resident 322M bidirectional direct-decision model (Laya-322M) under an asymmetric false-positive penalty. On the clean unseen DUCS-Bench test split ($N=607$), COGNIT-Guard achieves 98.85% accuracy (McNemar $p = 1.19 \times 10^{-4}$ vs. ML), reduces benign FPR to 0.42% ($1/238$; Fisher's exact $p = 8.23 \times 10^{-4}$ vs. ML), and attains 1.12% ECE and 0.0104 Brier score. On Huawei Ascend 910C NPUs, pure NPU inference runs in 21.77 ms mean latency (45.90 QPS), while the live serial CPU-NPU cascade ($θ^*_{\mathrm{deploy}}=0.70$) achieves 41.63 ms mean latency (P50: 39.47 ms, 99.23% accuracy, 0.00% FPR). Evaluation on SafetyBench-ZH ($N=2,100$) and comparison against a bi-encoder direct-decision baseline (CLM-8B) disentangle in-domain gains, OOD alignment tax (60.33% $\to$ 56.81% on Laya; 55.10% on domain CLM-8B), and experience replay recovery, restoring OOD accuracy to 64.10%-65.05% and reaching 99.67%-99.84% in-domain accuracy with 0.00%-0.42% FPR.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33671