OpenAI's GPT-6 Astra runs production systems at Perplexity and posts a strong early lead on a robotics spatial-reasoning benchmark
A new OpenAI customer case study says Perplexity now trusts GPT-6 Astra with communications, production system edits, and end-to-end automated testing, while early StationeryBench robotics results reported by The Decoder show Astra completing 7 of 100…
Two reports published on September 12, 2026 describe OpenAI's GPT-6 Astra as a model capable of operating real-world systems with less human oversight. In an OpenAI case study, Perplexity cofounder and Chief Strategy Officer Johnny Ho says the company uses GPT-6 Astra via API to craft communications, edit real-world systems, and monitor production software for its AI-powered answer engine, and that the team checks the model's work much less frequently than with previous models. Perplexity also has Astra build small test programs that stand in for external services, such as language model APIs and connectors, to verify applications end to end, with better code generation improving web and internal information search summarization. Separately, The Decoder reports that StationeryBench, a new robotics benchmark, tested GPT-6 Astra against Ai2's MolmoAct2 on five desk-object tasks using identical dual-arm YAM robots over 200 trials: Astra fully completed 7 of 100 tasks with a median progress score of 46 out of 100, while MolmoAct2 completed zero with a median score of 12. Cornell and Google DeepMind researcher Yoav Artzi called the result a "step change in spatial reasoning" and said Astra approaches human-level accuracy on the unpublished REMAP benchmark. He speculated OpenAI trained the model on large amounts of 3D data such as Blender scenes, and OpenAI reportedly plans consumer robots. The Perplexity claims come from an OpenAI-published case study, while the benchmark figures come from third-party reporting.
- Perplexity uses OpenAI's GPT-6 Astra via API to craft communications, edit real-world production systems, and monitor production software (OpenAI case study).
- Perplexity's Johnny Ho says the team checks the model's work much less frequently than with previous models.
- Astra generates mock service responses, standing in for external services like language model APIs and connectors, to test application workflows end to end.
- Better code generation directly improves Perplexity's web and internal information search summarization.
- On the StationeryBench robotics benchmark, GPT-6 Astra fully completed 7 of 100 tasks; Ai2's MolmoAct2 completed zero.
- Median progress scores across 200 trials on five desk-object tasks with identical dual-arm YAM robots: Astra 46/100 vs MolmoAct2 12/100.
- Researcher Yoav Artzi (Cornell and Google DeepMind) called Astra a "step change in spatial reasoning" and said it approaches human-level accuracy on the unpublished REMAP benchmark.
- Artzi speculated OpenAI trained the model on large amounts of 3D data such as Blender scenes; OpenAI reportedly plans consumer robots.
Coverage timelineoldest first · each row is one article
- · 4d agoPerplexity trusts GPT-6 Astra with end-to-end systems
OpenAI News· 38
Perplexity uses OpenAI's GPT-6 Astra to craft communications, edit production systems, and generate end-to-end automated tests for its search engine.
- · 4d agoGPT-6 Astra appears to show a "step change" in spatial reasoning based on early benchmarks
The Decoder· 48
Early StationeryBench robotics results show OpenAI's GPT-6 Astra far ahead of Ai2's MolmoAct2 at dual-arm manipulation, completing 7 of 100 tasks versus zero.