ZeroHour
Story · 6 sources · 11 articlesfirst updated ()

OpenAI's GPT-6 Astra posts benchmark and adoption wins as Meta's Muse and Cognition's SWE-2 expand the agent landscape

infoAI industryimportance 78
What's new: OpenAI's GPT-6 Astra posted benchmark leadership on ErdosBench (106/226 problems) and StationeryBench robotics (7/100 tasks vs zero for Ai2's MolmoAct2) while gaining its first published enterprise adopters, with OpenAI case studies detailing Astra-driven automated testing at Cognition/Devin and end-to-end use at Perplexity, plus new official prompt-migration guidance for developers. Meta…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

OpenAI's GPT-6 Astra led ulam.ai's ErdosBench math benchmark (106 of 226 problems) and showed a 'step change' in spatial reasoning on StationeryBench, while new OpenAI case studies detail production use at Cognition and Perplexity. Meta launched its…

OpenAI's GPT-6 Astra dominated this week's news. On ulam.ai's ErdosBench it scored 3.23, solving 106 of 226 open math problems (43 fully, 27 disproven), ahead of GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning; chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research. Early StationeryBench robotics results over 200 trials showed Astra fully completing 7 of 100 dual-arm YAM robot tasks (median progress 46/100) versus zero for Ai2's MolmoAct2 (12/100), which researcher Yoav Artzi called a 'step change in spatial reasoning.' OpenAI also published its first Astra customer case studies — Cognition uses it across Devin's cloud agent, CLI, and desktop for automated testing, and Perplexity uses it via API for communications, production edits, and end-to-end tests — plus developer guidance to shorten skills, trim AGENTS.md requirements, and define 'done,' since Astra may stop earlier than GPT-5.6 Sol. Meta launched Muse, a WhatsApp-controlled agent running on an isolated cloud VM with a Sentinel credential gatekeeper and Stripe Link one-time-card payments; its model reportedly scored 44-48 on Artificial Analysis Intelligence Index v4.3, near GPT-5.6 Sol's 47, but The Verge's hands-on found it surfacing Instagram/Facebook API-derived interests beyond users' visible ad-topic settings, which Meta disputes is a sharing issue. Cognition's SWE-2, post-trained from the 2.8T-parameter Kimi K3 MoE, posted vendor-reported scores of 50.0% on FrontierCode 1.1 Main and 92.8 on Terminal-Bench 2.1, within one point of Claude Fable 5.1 at a claimed 64% lower cost, though it trails Fable 5.1 and GPT-6 Astra on long-horizon Terminal-Bench 4.0 tasks. Separately, a Hugging Face paper found six frontier models from four labs converging on one imagined architecture under school-audience framing, coining 'epistemic jailbreak.'

  • GPT-6 Astra scored 3.23 on ulam.ai's ErdosBench, solving 106 of 226 open math problems (43 fully, 27 disproven); GPT-5.6 Sol trails with 3.12 and 78 solved at maximum reasoning.
  • Benchmark developer Przemek Chojecki estimated Astra's gain at 5-10% across tested math-research skills; OpenAI chief scientist Jakub Pachocki said targeted math optimization was skipped in favor of recursive self-improvement and automated…
  • Mathematician Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload.
  • On StationeryBench (five desk-object tasks, identical dual-arm YAM robots, 200 trials), Astra fully completed 7 of 100 tasks with median progress 46/100; Ai2's MolmoAct2 completed zero with 12/100.
  • Cornell/Google DeepMind researcher Yoav Artzi called Astra's result a 'step change in spatial reasoning,' said it approaches human-level accuracy on the unpublished REMAP benchmark, and speculated training on large 3D datasets such as…
  • Cognition uses GPT-6 Astra across Devin's cloud agent, CLI, and desktop: Devin tested the iPhone game Otter Run (returning a simulator recording plus coverage report) and fixes customer-reported bugs from screenshots with verification;…
  • Perplexity cofounder and Chief Strategy Officer Johnny Ho says GPT-6 Astra crafts communications, edits production systems, and builds mock services to test applications end to end, with far less supervision than previous models.
  • OpenAI's Eric Provencher recommends shorter skill descriptions (too many or conflicting ones truncate in Codex's context), selective AGENTS.md document reads, explicit permissions for safe operations like local test runs, and upfront…

Coverage timeline

  1. · 6d ago
    The Decoder· 72
    Muse can shop, write emails, and negotiate prices for users, all through WhatsApp

    Meta launched Muse, a WhatsApp-controlled agent running on an isolated VM with a Sentinel gatekeeper, able to shop, email, book travel, and negotiate.

  2. · 6d ago
    The Decoder· 68
    GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design

    OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.

  3. · 6d ago
    The Verge · AI· 78
    Sure, Meta’s AI Muse works, but it sure creeps me out

    Hands-on review finds Meta's Muse AI agent completes shopping and email tasks but surfaces personal Instagram API data beyond user-visible ad-topic settings.

  4. · 6d ago
    Hacker News · AI· 72
    Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

    Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.

  5. · 6d ago
    Hacker News · AI· 62
    Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

    Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.

  6. · 6d ago
    TechCrunch · AI· 32
    Meta’s AI agent Muse is now the No. 2 app in the US

    Meta's agentic AI app Muse ranked No. 2 on the US iOS charts with 83,000+ downloads, trailing Threads and ChatGPT launch pace.

  7. · 5d ago
    OpenAI News· 36
    Cognition helps Devin test its own work with GPT‑6 Astra

    Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.

  8. · 5d ago
    OpenAI News· 38
    Perplexity trusts GPT-6 Astra with end-to-end systems

    Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.

  9. · 4d ago
    The Decoder· 32
    GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends

    OpenAI's Eric Provencher advises developers using GPT-6 Astra to shorten skill descriptions, trim AGENTS.md reading requirements, relax approval rules, and define clear completion goals.

  10. · 4d ago
    The Decoder· 48
    GPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks

    Early StationeryBench robotics results show OpenAI's GPT-6 Astra far ahead of Ai2's MolmoAct2 at dual-arm manipulation, completing 7 of 100 tasks versus zero.

  11. · 4d ago
    Hugging Face daily papers· 35
    Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?

    Six frontier models from OpenAI, Anthropic, xAI, and Google DeepMind converge on one imagined successor architecture when asked under a school-audience framing.