Primary sourceAIArticle··5 min read

GPT-6.1 Sol: Almost Astra, for a Fifth of the Price

OpenAI updates its mid-tier model and promises flagship-level performance on code, agents, and professional work, with a bill cut fivefold.

GPT-6.1 Sol: Almost Astra, for a Fifth of the Price
Source : OpenAI · OpenAIView original ↗

In brief

OpenAI is launching GPT-6.1 Sol, an update to GPT-6 Sol that closes in on GPT-6 Astra on agentic coding, computer use, and professional tasks, at a fifth of its API price. Cached input drops to $0.10 per million tokens, a clear signal to agent developers who reuse lots of context.

🍺 Bar-stool version

OpenAI has two models: the very expensive one that can do everything, and the cheaper one that can do almost everything. Now the cheaper one just caught up to the expensive one on nearly everything except cutting-edge scientific research, for five times less. It's like finding out the house wine is as good as the grand cru, with a label written by the restaurant itself. If the numbers hold up outside the lab, running agents all day stops being a luxury.

Key takeaways

  1. 1

    GPT-6.1 Sol costs $2 per million input tokens, $0.10 for cached input, and $10 for output — a fifth of GPT-6 Astra's standard rates.

  2. 2

    Cached input is 95% cheaper than standard input and 50% cheaper than GPT-6 Sol's cache, a pitch tailor-made for agents that reuse their context.

  3. 3

    On DeepSWE v1.1, it matches GPT-6 Astra for about a fifth of the cost and beats GPT-6 Sol by 6.4 points.

  4. 4

    On Terminal-Bench Science 0.1, it costs $5.47 per task on average versus $23.21 for Opus 5.5 and $23.80 for Astra, which still leads with 68.1%.

  5. 5

    On OSWorld 2.0, it gains 7 points over GPT-6 Sol and lands within 2.1 points of Astra for about a seventh of the cost per task.

  6. 6

    The rate of responses containing a factual error drops from 11.4% to 7.7% at low reasoning effort, roughly 32% fewer errors.

  7. 7

    Available now in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu plans, as well as via the API (gpt-6.1-sol), but not yet in Chat.

The mid-tier model aiming for the top

GPT-6.1 Sol is an update to GPT-6 Sol, OpenAI's mid-tier model. Its pitch fits in one line: intelligence "close to Astra" on agentic coding, computer use, and professional work, for a fifth of Astra's standard input and output rates.

OpenAI positions the model as a new balance between capability and cost for everyday work: writing and debugging code, understanding documents, running multi-step business workflows. The comparison is made both against its bigger sibling Astra and against competitors, notably Opus 5.5.

Code and professional work: the headline numbers

On DeepSWE v1.1, which evaluates software engineering tasks on real codebases, GPT-6.1 Sol matches Astra for about a fifth of the cost. It also beats GPT-6 Sol's best score by 6.4 points, with lower reasoning effort and cost.

On GDP.pdf, which measures the accuracy of answers to professional questions about complex PDFs (tables, charts, diagrams, fine print) across finance, healthcare, law, and seven other domains, it beats Opus 5.5 with fallbacks for less than half the cost per task.

On AutomationBench 1.0.6, which tests full workflows with 47 tools (sales, marketing, operations, support, finance, HR), it outperforms Opus 5.5 by 2.2 points at medium effort, for about a third of the cost, and gains 4.8 points over GPT-6 Sol. OpenAI notes that the displayed cost for Claude Fable 5.1 is underestimated because it omits fallbacks, which occurred on about 40% of tasks.

Computer use and science

On the offline OSWorld 2.0 suite, which evaluates long application-use workflows, GPT-6.1 Sol gains 7 points over GPT-6 Sol at maximum effort, for less than half the cost. It lands within 2.1 points of Astra for about a seventh of the cost per task.

On Terminal-Bench Science 0.1 (data analysis, simulation, theorem proving), it more than doubles GPT-6 Sol's score. At maximum effort, it costs $5.47 per task, versus $23.21 for Opus 5.5 and $23.80 for Astra — more than 75% in savings.

OpenAI does set a limit, however: Astra remains the best model tested on this benchmark, at 68.1%, and should be favored for the hardest scientific research tasks.

Factuality and alignment

Factuality is measured on anonymized ChatGPT conversations where users had flagged an error from a previous model. At low effort, the share of flawed responses drops from 11.4% to 7.7%. Across all settings, the gap with Astra stays under 1.9 points, for less than a fifth of the cost.

OpenAI stresses that these deliberately hard prompts don't reflect everyday usage. On alignment, the model is said to be more transparent about its limits (for example when a search tool is down), more respectful of explicit restrictions, and less prone to producing unauthorized outputs in agent mode. No attempts to bypass the automated safety reviewer were observed, as with Astra and GPT-6 Sol. Details are laid out in an addendum to the system card.

Pricing and availability

The model is available today for Plus, Pro, Business, Enterprise, and Edu users in ChatGPT Work and Codex, but not yet in Chat. Developers can access it via the API under the name gpt-6.1-sol.

Pricing: $2 per million input tokens, $0.10 for cached input, $10 for output. In the coming days, OpenAI will also roll out GPT-6.1 Sol Ultrafast in Codex, with token generation up to 8 times faster than standard speed.

“Near-Astra intelligence for a fifth of the price”
“GPT‑6 Astra still achieves the highest score among the models tested at 68.1%, and should be used for the most difficult scientific research tasks.”
“We observed no attempts to bypass an automated safety reviewer, matching GPT‑6 Astra and GPT‑6 Sol.”

Why it matters

The announcement says less "here's a smarter model" than "here's yesterday's intelligence at a fraction of the price." That's the real playing field of 2026: agents burn through enormous numbers of tokens, constantly re-reading the same context, and their economic viability depends on cost per task more than on raw scores. The $0.10-per-million cache and the consistent cost-per-task framing make that clear. Still, these numbers deserve scrutiny. All results are published by OpenAI, often at reasoning-effort levels chosen per benchmark. Some tests (GDP.pdf, AutomationBench) are little known. And cost comparisons with competitors rely on questionable conventions, as acknowledged in the note on Claude Fable 5.1's fallbacks. Finally, the fact that Astra keeps the lead in science and that Sol isn't yet in Chat points to a deliberate segmentation: Sol for high-volume work, Astra for the hardest tasks. For teams running agents in production, this is the good news of the quarter — provided they rerun the benchmarks on their own tasks.

#openai#llm#agents#api#codex#benchmarks
Original source
Introducing GPT-6.1 Sol
OpenAI
Open the article ↗

For you

Put it to work on your sources.

Free: this week's articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.

For your team

The same machine, on your topics.

A space in your colours, your watch angles, your curators. Pilot open to three companies.

Read next