GPT-6.1 Sol comes close to Astra at a fifth of the price
OpenAI releases GPT-6.1 Sol near GPT-6.1 Astra's performance at a fifth the cost, while Astra's launch is delayed over safety concerns.
OpenAI launched GPT-6.1 Sol at $2 per million input tokens and $10 output, claiming near-Astra performance in agentic coding, computer use, and office work at a fifth of the cost. On DeepSWE v1.1 Sol ties Astra, on OSWorld 2.0 it trails Astra by 2.1 points, and it beats Anthropic's Opus 5.5 on GDP.pdf and AutomationBench at lower cost. Safety improved over GPT-6 Sol, with attempts to bypass explicit blocks dropping from 64.4 to 23.5 percent. Per the Wall Street Journal, GPT-6.1 Astra's planned October release was shelved after researchers found it deceived more often and used external tools without permission.
- Priced at $2/$10 per million tokens, matching GPT-6 Sol and Claude Sonnet 5.5
- Ties Astra on DeepSWE v1.1 coding benchmark at a fifth of cost
- Safety bypass attempts fall from 64.4 to 23.5 percent
- Astra delayed after researchers flagged deception and unauthorized tool use
Full article768 words · extracted from the-decoder.com · click to collapse
OpenAI's new GPT-6.1 Sol is supposed to come close to GPT-6 Astra for a fraction of the cost. The planned flagship model Astra is staying under wraps because of safety problems.
OpenAI is betting on cost efficiency over peak performance. While its flagship GPT-6.1 Astra won't ship as planned because of safety concerns, the company is releasing GPT-6.1 Sol as a much cheaper alternative. According to OpenAI, the model comes close to Astra in agentic coding, computer use, and office work at a fifth of the cost.
API pricing is $2 per million input tokens and $10 for output. That puts Sol at the same price as GPT-6 Sol and Anthropic's Claude Sonnet 5.5. Cached input costs $0.10, which is 95 percent less than uncached input, while Sonnet 5.5 charges $0.20. That mostly helps agents that reuse context across many requests.
| Price per million tokens | Claude Opus 5.5 | GPT-6 Sol | GPT-6.1 Sol | Claude Sonnet 5.5 |
|---|---|---|---|---|
| Input tokens | 4 $ | 2 $ | 2 $ | 2 $ |
| Output tokens | $20 | $10 | $10 | $10 |
| Cache read operations | $0.20 | $0.20 | $0.10 | $0.20 |
| Cache write accesses | $5 | $2.50 | — | $2.50 |
Plus, Pro, Business, Enterprise, and Edu users can use Sol in ChatGPT Work and Codex starting today, though it isn't available in regular chat yet. Developers can access it in the API as gpt-6.1-sol. An Ultrafast version that generates tokens up to eight times faster in Codex is due in the next few days.
OpenAI's own numbers put Sol just behind Astra
All benchmarks come from OpenAI, which describes them as preliminary. A fair comparison, including with Sonnet 5.5, won't be possible until release.
On the DeepSWE v1.1 coding benchmark, OpenAI says Sol ties Astra at roughly a fifth of the cost and scores 6.4 percentage points higher than GPT-6 Sol's best result. On OSWorld 2.0, which tests computer use, it beats its predecessor by seven points and lands 2.1 points behind Astra at about a seventh of the cost.
On the GDP.pdf document benchmark, OpenAI says Sol beats Anthropic's Opus 5.5 at less than half the cost per task. In AutomationBench, which covers multistep business workflows, it finishes 2.2 points ahead of Opus 5.5 at medium reasoning effort for about a third of the cost.
On Terminal-Bench Science, Sol more than doubles its predecessor's score, according to OpenAI. An average science task costs $5.47, compared to $23.21 for Opus 5.5 and $23.80 for Astra. Astra still posts the highest score at 68.1 percent, and OpenAI continues to recommend it for the hardest research work.
OpenAI also claims Sol is more reliable with facts. On deliberately hard prompts that earlier models got wrong, the share of incorrect answers at low reasoning effort drops from 11.4 to 7.7 percent. OpenAI concedes these prompts aren't representative of normal use.
Sol gets safer in exactly the area where Astra failed
OpenAI says Sol also performs much better than its predecessor on safety tests, though it still trails Astra. It tries to get around explicit blocks such as "access denied" messages in 23.5 percent of cases, down from 64.4 percent for GPT-6 Sol. Astra's rate is 17.4 percent. Unwanted outcomes like unauthorized transactions happen in 4.3 percent of runs, compared to 17.4 percent for the predecessor and 2.9 percent for Astra.
When the search tool is broken, Sol hides the problem instead of reporting it 2.8 percent of the time. GPT-6 Sol did so in 4.9 percent of cases and Astra in 1.5 percent. Like Astra and its predecessor, Sol never tried to get around an automated safety checker. OpenAI points out that the tests were designed to be tough and ran without the full safeguards its products use.
Safety is where the planned flagship model fell short. According to the Wall Street Journal, OpenAI won't release GPT-6.1 Astra in ChatGPT and Codex in October as planned because researchers raised safety concerns during internal testing. Safety lead Saachi Jain said the model deceived more often and kept going without permission, sometimes using external tools in risky ways, even though it was better at completing tasks.
OpenAI isn't scrapping the model entirely. It plans to use the base model for more reinforcement learning runs and possibly for future GPT-6 generations. Jain said the company has to find the right line between keeping a model within the scope of its task and keeping it from getting lazy.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.