ZeroHour
Story · 1 source · 1 articlefirst updated ()

GPT-6 Astra tops ARC-AGI-3, ErdosBench and robotics benchmarks; OpenAI ships GPT-Live-1 voice API as Cognition counters with SWE-2

infoModel releaseimportance 92
What's new: Beyond the previously covered GPT-6 Astra benchmark results (ARC-AGI-3 99.9%, ErdosBench 3.23, StationeryBench 7/100) and the GPT-Live-1 launch: (1) Added Cognition's SWE-2 launch with full benchmark scores (50.0 FrontierCode 1.1 Main, 73.0 DeepSWE 1.1, 92.8 Terminal-Bench 2.1, 27.3 Terminal-Bench 4.0), cost claims (64% cheaper than Claude Fable 5.1; 58% fewer turns and 81% cheaper than SWE-1.7),…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

OpenAI's new flagship GPT-6 Astra scores 99.9% on ARC-AGI-3 (vs 7.8% for GPT-5.6 Sol), 3.23 on ulam.ai's ErdosBench (106 of 226 problems solved) and completes 7 of 100 StationeryBench robotics tasks; OpenAI also shipped the $0.05/min full-duplex GPT-Live-1…

Coverage from September 9-12, 2026 centers on OpenAI's release of GPT-6 Astra, its strongest model to date. Reviewer Sebastian Raschka calls it the best model he has used, with disproportionate gains in 3D rendering, animation, and computer use through the Codex/ChatGPT harness. Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol and leads the Artificial Analysis Coding Agent Index, though independent aggregate indices show more incremental gains. The coverage also discusses unconfirmed looped-transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning. On ulam.ai's ErdosBench, Astra scores 3.23, solving 106 of 226 open math problems (43 fully, 27 disproved), ahead of GPT-5.6 Sol's 78 solved at 3.12 with maximum reasoning. Chief scientist Jakub Pachocki says OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research; benchmark developer Przemek Chojecki estimates the gain at 5-10% across tested math-research skills; and Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload. Early StationeryBench robotics results show Astra fully completing 7 of 100 dual-arm manipulation tasks with a median progress score of 46/100 over 200 trials, versus zero tasks and 12/100 for Ai2's MolmoAct2, using identical dual-arm YAM robots on five desk-object tasks. Cornell and Google DeepMind researcher Yoav Artzi calls it a 'step change in spatial reasoning' and says Astra approaches human-level accuracy on the unpublished REMAP benchmark; he speculates OpenAI trained on large 3D datasets such as Blender scenes, and OpenAI reportedly plans consumer robots. OpenAI simultaneously shipped GPT-Live-1 in the API, a single-model full-duplex voice system that listens and speaks simultaneously, replacing chained STT-LLM-TTS architectures and delegating reasoning and tool calls to backend models such as GPT-6 Astra; it already powers ChatGPT voice. OpenAI cites a +30 percentage-point gain on Full Duplex Bench over GPT-Realtime-2.1, while The Decoder's reported figures (80.1% versus 45.4% on full-duplex interactivity) imply a 34.7-point gap; the two sources' metrics may differ. The Decoder also reports 0.8-second versus 1.4-second turn-taking latency, 87% versus 60% tool-calling accuracy, and a 32% versus 12.4% pass rate on a banking voice-support…

  • GPT-6 Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol and leads the Artificial Analysis Coding Agent Index; independent aggregate indices show more incremental gains.
  • ErdosBench (ulam.ai): Astra scores 3.23 with 106 of 226 open math problems solved (43 fully, 27 disproved); GPT-5.6 Sol trails at 3.12 with 78 solved at maximum reasoning.
  • Chief scientist Jakub Pachocki says OpenAI skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research; benchmark developer Przemek Chojecki estimates a 5-10% gain across tested…
  • StationeryBench: Astra fully completed 7 of 100 tasks with median progress 46/100 over 200 trials; Ai2's MolmoAct2 completed zero with 12/100; identical dual-arm YAM robots on five desk-object tasks.
  • Yoav Artzi (Cornell/Google DeepMind) calls Astra a 'step change in spatial reasoning' nearing human-level on the unpublished REMAP benchmark; he speculates training on large 3D data such as Blender scenes, and OpenAI reportedly plans…
  • GPT-Live-1 is a single-model full-duplex voice API that replaces chained STT-LLM-TTS architectures, already powers ChatGPT voice, and delegates reasoning/tool calls to backend models such as GPT-6 Astra.
  • Voice benchmarks: OpenAI cites +30 percentage points on Full Duplex Bench over GPT-Realtime-2.1; The Decoder reports 80.1% vs 45.4% full-duplex interactivity (implied 34.7-point gap), 0.8s vs 1.4s turn-taking latency, 87% vs 60%…
  • GPT-Live-1 ranks #1 on Tau3 when paired with GPT-6 Astra at medium reasoning effort; costs $0.05 per minute for the front-end voice layer; ships twelve new voices with telephony, native ASR transcripts, and keyword biasing.

Coverage timeline

  1. · 7d ago
    Hacker News · AI· 92
    GPT-6 Astra, Looped Transformers, and Hidden Reasoning

    OpenAI released GPT-6 Astra, its strongest model to date, with standout 3D rendering and computer-use performance and 99.9% on ARC-AGI-3.