JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
A decision-only JEV judge nearly matches a top LLM judge at a fraction of the cost.
JEV-as-a-judge is compared with sixteen generative and reward-model judges under blinded human adjudication. It stays within three percentage points of the strongest LLM judge on ordinary preference and evidence-grounded factuality at 0.36% of that judge's fee. Larger gaps appear when checking derivations or resisting elaborate wrong answers. A frozen cascade that escalates low-confidence cases retains 99% of comparator accuracy at lower cost.
- Within three points of the strongest comparator on preference and factuality.
- Its fee is 0.36% of the comparator's cost.
- Larger gaps appear on derivations and elaborate wrong answers.
- A confidence cascade retains 99% of comparator accuracy.
Full article123 words · extracted from huggingface.co · click to collapse
LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.26550