ZeroHour
Story · 3 sources · 3 articlesfirst updated ()

Google DeepMind Launches Gemini 3.8 Live and 3.8 Live Extended Thinking Speech Models

infoModel releaseimportance 78
What's new: The MarkTechPost report adds new details not in the previous summary: per-minute pricing ($0.005/min audio input, $0.018/min audio output), asynchronous function calling that streams audio while tools execute, alphanumeric precision, and the note that the models are hosted via the Gemini Live API and AI Studio. All previously reported benchmark figures, 97-language support, SynthID watermarking,…
Merged summary · glm-5.3 · rewritten as coverage arrives

Google DeepMind released Gemini 3.8 Live and 3.8 Live Extended Thinking near-real-time speech-to-speech models for voice agents; Extended Thinking tops Artificial Analysis' Speech-to-Speech Quality Index at 82.6, with 97-language support, background tool…

Google DeepMind announced Gemini 3.8 Live, a cost-efficient model built for near-real-time dialogue with real-time visual grounding, and Gemini 3.8 Live Extended Thinking, aimed at high-complexity multi-step reasoning in voice agents. Extended Thinking ranks #1 on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6, and scores 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice-banking agentic benchmark, and 97.7% on Big Bench Audio. Both models detect and switch among 97 languages mid-conversation, process near-real-time visual input, and execute background tool and API calls, including asynchronous function calling that streams audio while tools execute; MarkTechPost also credits them with alphanumeric precision. Per MarkTechPost, the models are hosted via the Gemini Live API and AI Studio at $0.005/min for audio input and $0.018/min for audio output. Availability spans the Gemini API, AI Studio, Gemini Enterprise private preview, and Search Live; the Hacker News report additionally lists Workspace as part of the rollout. All generated audio carries Google DeepMind's imperceptible SynthID watermark.

  • Extended Thinking ranks #1 on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6
  • Benchmark scores: 68.6% on τ-Voice, 35.1% on Sierra's τ-Voice-banking, 97.7% on Big Bench Audio
  • Models detect and switch among 97 languages mid-conversation
  • Real-time visual grounding plus background tool and API calls, including asynchronous function calling (MarkTechPost)
  • Pricing reported by MarkTechPost: $0.005/min audio input, $0.018/min audio output, hosted via Gemini Live API and AI Studio
  • Availability: Gemini API, AI Studio, Gemini Enterprise private preview, and Search Live; Hacker News additionally lists Workspace
  • All generated audio is watermarked with Google DeepMind's imperceptible SynthID

Coverage timeline

  1. · 4h ago
    Google DeepMind· 74
    Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

    Google DeepMind launched Gemini 3.8 Live and 3.8 Live Extended Thinking speech models, topping Artificial Analysis' Speech-to-Speech Quality Index at 82.6.

  2. · 3h ago
    Hacker News · AI· 72
    Gemini 3.8 Live and 3.8 Live Extended Thinking

    Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking speech models, topping speech-to-speech benchmarks with parallel reasoning for voice agents.

  3. · 34m ago
    MarkTechPost· 78
    Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

    Google launches Gemini 3.8 Live and Extended Thinking speech-to-speech models for production voice agents, topping speech-to-speech benchmarks.