GPT-6 Astra tops ARC-AGI-3, ErdosBench and robotics benchmarks; OpenAI ships GPT-Live-1 voice API as Cognition counters with SWE-2
OpenAI's new flagship GPT-6 Astra scores 99.9% on ARC-AGI-3 (vs 7.8% for GPT-5.6 Sol), 3.23 on ulam.ai's ErdosBench (106 of 226 problems solved) and completes 7 of 100 StationeryBench robotics tasks; OpenAI also shipped the $0.05/min full-duplex GPT-Live-1…
Coverage from September 9-12, 2026 centers on OpenAI's release of GPT-6 Astra, its strongest model to date. Reviewer Sebastian Raschka calls it the best model he has used, with disproportionate gains in 3D rendering, animation, and computer use through the Codex/ChatGPT harness. Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol and leads the Artificial Analysis Coding Agent Index, though independent aggregate indices show more incremental gains. The coverage also discusses unconfirmed looped-transformer/recurrent-depth architecture rumors and speculation that Astra hides its chain-of-thought reasoning. On ulam.ai's ErdosBench, Astra scores 3.23, solving 106 of 226 open math problems (43 fully, 27 disproved), ahead of GPT-5.6 Sol's 78 solved at 3.12 with maximum reasoning. Chief scientist Jakub Pachocki says OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research; benchmark developer Przemek Chojecki estimates the gain at 5-10% across tested math-research skills; and Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload. Early StationeryBench robotics results show Astra fully completing 7 of 100 dual-arm manipulation tasks with a median progress score of 46/100 over 200 trials, versus zero tasks and 12/100 for Ai2's MolmoAct2, using identical dual-arm YAM robots on five desk-object tasks. Cornell and Google DeepMind researcher Yoav Artzi calls it a 'step change in spatial reasoning' and says Astra approaches human-level accuracy on the unpublished REMAP benchmark; he speculates OpenAI trained on large 3D datasets such as Blender scenes, and OpenAI reportedly plans consumer robots. OpenAI simultaneously shipped GPT-Live-1 in the API, a single-model full-duplex voice system that listens and speaks simultaneously, replacing chained STT-LLM-TTS architectures and delegating reasoning and tool calls to backend models such as GPT-6 Astra; it already powers ChatGPT voice. OpenAI cites a +30 percentage-point gain on Full Duplex Bench over GPT-Realtime-2.1, while The Decoder's reported figures (80.1% versus 45.4% on full-duplex interactivity) imply a 34.7-point gap; the two sources' metrics may differ. The Decoder also reports 0.8-second versus 1.4-second turn-taking latency, 87% versus 60% tool-calling accuracy, and a 32% versus 12.4% pass rate on a banking voice-support…
- GPT-6 Astra scores 99.9% on ARC-AGI-3 versus 7.8% for GPT-5.6 Sol and leads the Artificial Analysis Coding Agent Index; independent aggregate indices show more incremental gains.
- ErdosBench (ulam.ai): Astra scores 3.23 with 106 of 226 open math problems solved (43 fully, 27 disproved); GPT-5.6 Sol trails at 3.12 with 78 solved at maximum reasoning.
- Chief scientist Jakub Pachocki says OpenAI skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research; benchmark developer Przemek Chojecki estimates a 5-10% gain across tested…
- StationeryBench: Astra fully completed 7 of 100 tasks with median progress 46/100 over 200 trials; Ai2's MolmoAct2 completed zero with 12/100; identical dual-arm YAM robots on five desk-object tasks.
- Yoav Artzi (Cornell/Google DeepMind) calls Astra a 'step change in spatial reasoning' nearing human-level on the unpublished REMAP benchmark; he speculates training on large 3D data such as Blender scenes, and OpenAI reportedly plans…
- GPT-Live-1 is a single-model full-duplex voice API that replaces chained STT-LLM-TTS architectures, already powers ChatGPT voice, and delegates reasoning/tool calls to backend models such as GPT-6 Astra.
- Voice benchmarks: OpenAI cites +30 percentage points on Full Duplex Bench over GPT-Realtime-2.1; The Decoder reports 80.1% vs 45.4% full-duplex interactivity (implied 34.7-point gap), 0.8s vs 1.4s turn-taking latency, 87% vs 60%…
- GPT-Live-1 ranks #1 on Tau3 when paired with GPT-6 Astra at medium reasoning effort; costs $0.05 per minute for the front-end voice layer; ships twelve new voices with telephony, native ASR transcripts, and keyword biasing.
Coverage timelineoldest first · each row is one article
- · 7d agoGPT-6 Astra, Looped Transformers, and Hidden Reasoning
Hacker News · AI· 92
OpenAI released GPT-6 Astra, its strongest model to date, with standout 3D rendering and computer-use performance and 99.9% on ARC-AGI-3.