SearchJev: A Fast and Calibrated System-1 Model for Search Agents
SearchJev makes calibrated search decisions 5.2× faster than Qwen3.5 and lifts BrowseComp-Plus accuracy from 45% to 54%.
SearchJev is a calibrated System-1 model that scores legal search options directly rather than generating them autoregressively, trained with Soft-Label Learning for Calibrated Decisions. On SearchDecision-Bench it beat same-size Qwen3.5 models, made decisions 5.2–5.3 times faster, and reduced expected calibration error by 41–74%. Uncertain cases are delegated to a System-2 model that keeps planning, query generation, and answer writing. On BrowseComp-Plus, dual-system agents improved answer accuracy from 45% to as high as 54% while speeding active search 3.7–4.7 times.
- SearchJev scores legal options without autoregressive text generation.
- SLCD learns decision probabilities from uncertain supervision.
- 5.2–5.3× faster than same-size Qwen3.5 with 41–74% lower ECE.
- BrowseComp-Plus accuracy rose from 45% to 54% at 3.7–4.7× speed.
Full article167 words · extracted from huggingface.co · click to collapse
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05107