Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
RLCDAlignBench benchmarks Jev, a calibrated-decision model detecting ten AI alignment failure types zero-shot at 0.886 median AUROC, 63x cheaper than LLM judges.
The paper introduces RLCDAlignBench, evaluating Jev — a model trained with reinforcement learning for calibrated decisions — on ten alignment failures including jailbreaks, prompt injection, deception, sycophancy, and power seeking, across 44 benchmarks and five target models. By varying question wording and input fields separately, a single generic question reaches a median AUROC of 0.886 zero-shot, beating supervised baselines on most benchmarks. Jev matches reference scorers' agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers.