Jev detects ten alignment failures zero-shot at 0.886 AUROC
Jev flags ten AI alignment failures zero-shot at a 0.886 median AUROC on 44 benchmarks, matching human-label agreement at 63 times lower cost than LLM judges.
Researchers introduce Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), which answers many typed questions about one input and returns calibrated probabilities in a single call. RLCDAlignBench evaluates it on ten alignment failures across 44 benchmarks and five target models. Both reports name jailbreaks, prompt injection, deception, sycophancy, and power seeking, while only the Hugging Face report also lists hallucination. A single generic question reaches a median AUROC of 0.886 zero-shot, beats supervised baselines on most benchmarks, matches reference scorers' agreement with human labels, and costs 63 times less than LLM-judge scorers. The arXiv report adds that human labels exist on two benchmarks and that Jev surfaces label defects in existing benchmarks. The sources agree that input context matters more than question wording; the Hugging Face report specifies label-encoding fields, and they do not dispute the shared figures.
- Jev is trained with reinforcement learning for calibrated decisions (RLCD) and scores many typed questions about one input with calibrated probabilities in a single call.
- RLCDAlignBench covers ten alignment-failure types across 44 benchmarks and five target models.
- Both reports name jailbreaks, prompt injection, deception, sycophancy, and power seeking; only the Hugging Face report also lists hallucination.
- A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.
- Jev matches reference scorers' agreement with human labels; the arXiv report says human labels exist on two benchmarks and that Jev surfaces label defects.
- The method costs 63 times less than LLM-judge scorers.
- Question wording matters little; input context, especially label-encoding fields, matters more.
Coverage timelineoldest first · each row is one article
- · 3d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
Hugging Face daily papers· 56
Jev detects ten alignment failures zero-shot, with a median AUROC of 0.886 on RLCDAlignBench.
- · 2d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
arXiv cs.CR· 45
RLCDAlignBench benchmarks Jev, a calibrated-decision model detecting ten AI alignment failure types zero-shot at 0.886 median AUROC, 63x cheaper than LLM judges.