Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
Sys1Cal-v1 finds Jev hides an unknown probability in binary choices, and recovering it improves calibration.
Sys1Cal-v1 is a calibration dataset of true/false questions whose proposition probability is known by construction, scored by total variation from the ground-truth distribution. The authors query Jev through its Noul, Choice, and Score primitives and also evaluate SemIf, an open-source Choice-style baseline. They argue Jev's Choice answers force P(A)+P(not A)=1 while suppressing a nonzero unknown mass P(U). Recovering that term raises median soft accuracy on Choice answers from 0.771 to 0.978.
- Sys1Cal-v1 uses true/false items with known ground-truth probabilities.
- Items are queried via Jev's Noul, Choice, and Score primitives.
- SemIf is evaluated as an open-source Choice-style baseline.
- Choice outputs omit a third unknown mass P(U).
- Recovering P(U) lifts median soft accuracy from 0.771 to 0.978.
Full article253 words · extracted from huggingface.co · click to collapse
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition A for which the exact probability P(A) is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in Choice answers, P(A) and P(neg A) are presented as if P(A)+P(neg A)=1, while a term P(U)neq0 is missing in the sum. Recovering P(U) leads to an improvement of median soft accuracy in Choice answers from 0.771 to 0.978, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35342