OpenAI ships GPT-6.1 Sol at $2/$10 per million tokens, near-Astra scores for a fifth of the price
- GPT-6.1 Sol is priced at $2 per million input and $10 per million output tokens, one-fifth of GPT-6 Astra's $10/$50, with cached input at $0.10 per million (95% below standard input pricing and 50% below GPT-6 Sol's cached rate).
- On DeepSWE v1.1 the model matches Astra's score at roughly one-fifth the cost and beats GPT-6 Sol's best by 6.4 percentage points at lower reasoning effort and cost.
- On GDP.pdf it scores above Opus 5.5 with fallbacks at under half the cost per task, and on AutomationBench it lands 2.2 points above Opus 5.5 at medium reasoning effort for about a third of the cost, up 4.8 points from GPT-6 Sol.
- On OSWorld 2.0 (offline set) it beats GPT-6 Sol by 7 percentage points at maximum reasoning effort for less than half the cost, and comes within 2.1 points of Astra at roughly one-seventh the cost per task.
- On Terminal-Bench Science 0.1 it more than doubles GPT-6 Sol's score at $5.47 per task versus $23.21 for Opus 5.5 and $23.80 for Astra, though Astra still leads at 68.1%; at low reasoning effort the share of responses with a factual error drops from 11.4% to 7.7%.
Hacker News opinions
Weren't there headlines just yesterday saying they were holding this back over safety concerns?
That was 6.1 Astra, this is Sol. Different model entirely.
My guess is Astra got tabled because it still doesn't beat Opus 5.5. Sol is a decent win if it holds up, 6 Sol was no good in my work.
The API price cut was obvious the moment they announced the change to how usage is counted. All of this makes sense once you look at the enterprise market.
Shots fired, half the price of Opus 5.5.
Wasn't 6 released like last week? I can't keep up anymore.
6 Sol was so underwhelming it never even made it to the ChatGPT chat interface.
It's ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
Great for the consumer. I remember when bandwidth was super expensive and now it's dirt cheap.
Open weight models are already something like 60% of token spend and it'll get worse. We've been running GLM 5.x through a gateway and it's pretty close to the frontier.
If these models are so smart, can't they pick the right model for each task themselves?
Switching models is expensive in compute, you have to rerun everything from the start. Cursor tried it and most users turned it off and pick models manually.
The models don't even know about themselves, the knowledge cutoff predates their own release. I keep a list of 3 or 4 models in context so it doesn't default to Haiku as the cheap option.
What's driving the release cadence? We seem to get a new model every week now.
Chinese model pressure. A lot of my SWE friends switched, I dropped OpenAI and Anthropic for QWEN and GLM on API projects, cost was the only reason.
They're rushing Sol 6.1 because Astra 6.1 got postponed, Sol 6 was weak, and they had to ship something in response to Opus 5.5.
I love free market competition. LLMs used to cost an arm and a leg for decent intelligence.
They also cut subscription allowances in half, so even in the best case it's about 2.5 times cheaper for Codex users. They matched Claude Sonnet 5.5 API pricing, but Claude Code's subscription allowance looks way more generous now.