ZeroHour
Story · 1 source · 2 articlesfirst updated ()1

TypeSafe's Jev 'System One' decision model claims 20-200x faster, 40-400x cheaper classification than frontier LLMs; open reproduction and critique follow within days

infoModel releaseimportance 75
What's new: New this window (first merged summary; sources dated 2026-09-16 and 2026-09-18): (1) Product launches - TypeSafe's Jev decision model, Google's Gemini 3.8 Live and 3.8 Live Extended Thinking, Gemini managed-agent Credentials and Files APIs, Anthropic's Claude Code Projects, and OpenAI's Astra for Law via Trusted Access. (2) Benchmark results - Gemini 3.8 Live debuts #1 on Artificial Analysis'…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

TypeSafe launched Jev, an RLCD-trained decision model that only decides, classifies, routes, or scores rather than generating text, claiming 20-200x faster and 40-400x cheaper classification/routing than frontier LLMs (headline framing: >100x faster, >200x…

Per Latent Space digests dated 2026-09-16 and 2026-09-18 (the latter covering 9/16-9/17/2026 news), the central shared story is TypeSafe's Jev, a 'System One' decision model framed as a discriminative routing/classification primitive rather than a chatbot. It is trained with RLCD, targets only decisions, classification, routing, judgment, and structured decisions, and produces constrained output with free output tokens and no hallucinated text. The 9/16 report's headline claims >100x faster and >200x cheaper than small frontier LLMs, while its body claims a broader 20-200x faster and 40-400x cheaper classification and routing than frontier LLMs; the two framings differ in scope (headline compares against small frontier LLMs, the range against frontier LLMs generally). Reception within roughly a day: an open reproduction, openjev-s, was built on Qwen3.6-35B-A3B, and Theo argued Jev-based history compaction can degrade reasoning traces and invalidate cached prefixes. Other items across the two digests: Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking, supporting 97 languages and async tool calls while speaking, debuting #1 on Artificial Analysis' speech-to-speech index at 82.6, and its Gemini managed agents added a Credentials API, Files API, egress proxying, and a claimed 30% cost reduction; Periodic Labs' Neon, a ~1T-parameter XRD analysis model trained with RL on proprietary lab data using 1,300 H200s, lifted FrontierXRD success from 2.7% to 55.3%, surpassing GPT-6 Astra at lower inference cost; Anthropic's Claude Code Projects enables one conversation to spawn parallel cloud sessions sharing context under one coordinator Claude; OpenAI's Astra for Law ships 26 partner-built and 47 community plugins via Trusted Access in ChatGPT and Codex, with reports it beats generic GPT-6 Astra plus web search on Vals' legal benchmark; DeepMind's Stellar Colosseum multi-agent math harness scored 4263 on Codeforces and 71.0% on TCS-Bench; and NVIDIA-associated Agora uses Git commits as shared memory. No figures conflict across the two reports beyond the differing claim scopes on Jev's speed/cost noted above.

  • TypeSafe launched Jev, a 'System One' decision model trained with RLCD that only decides, classifies, routes, or scores - not a text generator or chatbot.
  • Jev's claimed performance: 20-200x faster and 40-400x cheaper classification and routing than frontier LLMs, with free output tokens and no hallucinated text; the 9/16 headline states >100x faster and >200x cheaper than small frontier LLMs…
  • Open reproduction openjev-s exists, built on Qwen3.6-35B-A3B.
  • Theo argues Jev-based history compaction can degrade reasoning traces and invalidate cached prefixes.
  • Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking: 97 languages, async tool calls while speaking, #1 on Artificial Analysis' speech-to-speech index at 82.6.
  • Google's Gemini managed agents add a Credentials API, Files API, and egress proxying, with a claimed 30% cost reduction.
  • Periodic Labs' Neon is a ~1T-parameter XRD analysis model trained with RL on proprietary lab data using 1,300 H200s; it lifted FrontierXRD success from 2.7% to 55.3%, surpassing GPT-6 Astra at lower inference cost.
  • Anthropic's Claude Code Projects lets one conversation spawn parallel cloud sessions sharing context under one coordinator Claude.

Coverage timeline

  1. · 2d ago
    Latent Space· 75
    [AINews] Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs

    TypeSafe launches Jev, an RLCD-trained decision model claiming 20-200x faster, 40-400x cheaper classification than frontier LLMs, alongside Gemini 3.8 Live and Neon.

  2. · 8h ago
    Latent Space· 35
    [AINews] not much happened today

    Latent Space AI news digest covers Anthropic's Claude Code Projects, Google's managed agent APIs, TypeSafe's Jev classifier, and OpenAI's Astra for Law launch.