OpenAI's GPT-6 Astra posts benchmark and adoption wins as Meta's Muse and Cognition's SWE-2 expand the agent landscape
OpenAI's GPT-6 Astra led ulam.ai's ErdosBench math benchmark (106 of 226 problems) and showed a 'step change' in spatial reasoning on StationeryBench, while new OpenAI case studies detail production use at Cognition and Perplexity. Meta launched its…
OpenAI's GPT-6 Astra dominated this week's news. On ulam.ai's ErdosBench it scored 3.23, solving 106 of 226 open math problems (43 fully, 27 disproven), ahead of GPT-5.6 Sol's 3.12 and 78 solved at maximum reasoning; chief scientist Jakub Pachocki said OpenAI deliberately skipped targeted math optimization to prioritize recursive self-improvement and automated alignment research. Early StationeryBench robotics results over 200 trials showed Astra fully completing 7 of 100 dual-arm YAM robot tasks (median progress 46/100) versus zero for Ai2's MolmoAct2 (12/100), which researcher Yoav Artzi called a 'step change in spatial reasoning.' OpenAI also published its first Astra customer case studies — Cognition uses it across Devin's cloud agent, CLI, and desktop for automated testing, and Perplexity uses it via API for communications, production edits, and end-to-end tests — plus developer guidance to shorten skills, trim AGENTS.md requirements, and define 'done,' since Astra may stop earlier than GPT-5.6 Sol. Meta launched Muse, a WhatsApp-controlled agent running on an isolated cloud VM with a Sentinel credential gatekeeper and Stripe Link one-time-card payments; its model reportedly scored 44-48 on Artificial Analysis Intelligence Index v4.3, near GPT-5.6 Sol's 47, but The Verge's hands-on found it surfacing Instagram/Facebook API-derived interests beyond users' visible ad-topic settings, which Meta disputes is a sharing issue. Cognition's SWE-2, post-trained from the 2.8T-parameter Kimi K3 MoE, posted vendor-reported scores of 50.0% on FrontierCode 1.1 Main and 92.8 on Terminal-Bench 2.1, within one point of Claude Fable 5.1 at a claimed 64% lower cost, though it trails Fable 5.1 and GPT-6 Astra on long-horizon Terminal-Bench 4.0 tasks. Separately, a Hugging Face paper found six frontier models from four labs converging on one imagined architecture under school-audience framing, coining 'epistemic jailbreak.'
- GPT-6 Astra scored 3.23 on ulam.ai's ErdosBench, solving 106 of 226 open math problems (43 fully, 27 disproven); GPT-5.6 Sol trails with 3.12 and 78 solved at maximum reasoning.
- Benchmark developer Przemek Chojecki estimated Astra's gain at 5-10% across tested math-research skills; OpenAI chief scientist Jakub Pachocki said targeted math optimization was skipped in favor of recursive self-improvement and automated…
- Mathematician Terence Tao warned at the 2026 International Congress of Mathematicians that AI-generated proofs could shift mathematics from proof scarcity to proof overload.
- On StationeryBench (five desk-object tasks, identical dual-arm YAM robots, 200 trials), Astra fully completed 7 of 100 tasks with median progress 46/100; Ai2's MolmoAct2 completed zero with 12/100.
- Cornell/Google DeepMind researcher Yoav Artzi called Astra's result a 'step change in spatial reasoning,' said it approaches human-level accuracy on the unpublished REMAP benchmark, and speculated training on large 3D datasets such as…
- Cognition uses GPT-6 Astra across Devin's cloud agent, CLI, and desktop: Devin tested the iPhone game Otter Run (returning a simulator recording plus coverage report) and fixes customer-reported bugs from screenshots with verification;…
- Perplexity cofounder and Chief Strategy Officer Johnny Ho says GPT-6 Astra crafts communications, edits production systems, and builds mock services to test applications end to end, with far less supervision than previous models.
- OpenAI's Eric Provencher recommends shorter skill descriptions (too many or conflicting ones truncate in Codex's context), selective AGENTS.md document reads, explicit permissions for safe operations like local test runs, and upfront…
Coverage timelineoldest first · each row is one article
- · 6d agoMuse can shop, write emails, and negotiate prices for users, all through WhatsApp
The Decoder· 72
Meta launched Muse, a WhatsApp-controlled agent running on an isolated VM with a Sentinel gatekeeper, able to shop, email, book travel, and negotiate.
- · 6d agoGPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design
The Decoder· 68
OpenAI's GPT-6 Astra tops ulam.ai's ErdosBench math benchmark with 106 of 226 problems solved, while the company prioritizes recursive self-improvement over math optimization.
- · 6d agoSure, Meta’s AI Muse works, but it sure creeps me out
The Verge · AI· 78
Hands-on review finds Meta's Muse AI agent completes shopping and email tasks but surfaces personal Instagram API data beyond user-visible ad-topic settings.
- · 6d agoCognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Hacker News · AI· 72
Cognition released SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, near Fable 5.1 at 64% lower cost.
- · 6d agoCognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
Hacker News · AI· 62
Cognition releases SWE-2, a 2.8T-parameter MoE coding model post-trained from Kimi K3, scoring 92.8 on Terminal-Bench 2.1.
- · 6d agoMeta’s AI agent Muse is now the No. 2 app in the US
TechCrunch · AI· 32
Meta's agentic AI app Muse ranked No. 2 on the US iOS charts with 83,000+ downloads, trailing Threads and ChatGPT launch pace.
- · 5d agoCognition helps Devin test its own work with GPT‑6 Astra
OpenAI News· 36
Cognition integrates GPT-6 Astra into Devin, its CLI, and desktop products to automate testing and provide evidence for code review.
- · 5d agoPerplexity trusts GPT-6 Astra with end-to-end systems
OpenAI News· 38
Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.
- · 4d agoGPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends
The Decoder· 32
OpenAI's Eric Provencher advises developers using GPT-6 Astra to shorten skill descriptions, trim AGENTS.md reading requirements, relax approval rules, and define clear completion goals.
- · 4d agoGPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks
The Decoder· 48
Early StationeryBench robotics results show OpenAI's GPT-6 Astra far ahead of Ai2's MolmoAct2 at dual-arm manipulation, completing 7 of 100 tasks versus zero.
- · 4d agoAnother Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
Hugging Face daily papers· 35
Six frontier models from OpenAI, Anthropic, xAI, and Google DeepMind converge on one imagined successor architecture when asked under a school-audience framing.