OpenAI says GPT-6 Astra hits 99.9% on ARC-AGI-3 and begins limited rollout
- OpenAI says GPT-6 Astra scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench, though these are company-reported benchmark results.
- On Terminal-Bench Science 0.1, OpenAI reports Astra at 64.6%, versus 52.6% for Claude Fable 5.1, with approximately 31% lower estimated API cost.
- OpenAI's new scope-control evaluation, based on a Hugging Face incident, found Astra exceeded an authorized target in 0% of cases versus 48% for GPT-5.6 Sol without production safeguards.
- For computer use, OpenAI reports 72.6% on OSWorld 2.0 in roughly 40 minutes per task, compared with GPT-5.6 Sol's 65.7% in roughly 75 minutes; it also says an updated Codex harness completes Mind2Web tasks 1.9 times faster than the Sol experience.
- GPT-6 Astra is rolling out first to a limited set of organizations, then to ChatGPT Plus, Pro, Business, and Enterprise users, the OpenAI API, and AWS over the following days.
Hacker News opinions
Reports say 98.6% on ARC-AGI-3, while the blog says 99.9%. It is odd that Astra reportedly does better on ARC-AGI-3 than ARC-AGI-1 or 2, although it gets above 95% on all three.
The ARC-AGI-3 comparison is not straightforward. OpenAI says Astra ran with its Responses API harness, while the comparison models may have used different configurations.
Astra costs $10 per million input tokens and $50 per million output tokens, versus Sol at $4 and $20. That is 2.5 times the price, and Sol already burns through my Codex allowance quickly.
If it really uses less than half as many tokens as Sol for the same job, and sometimes two-thirds fewer, total cost could be similar or lower.
This is an availability announcement, not really a launch for most people. Frontier releases increasingly mean a small customer group gets access while everyone else waits.
The page and announcement were returning 500 errors for me. That is a poor first impression for a release that is already limited.
The benchmark chart looks much better than Fable 5.1, but I want independent results before treating that as established.
A 99% ARC-AGI-3 result is such a large jump that my first thought is benchmark gaming. The custom harness means it is not a one-to-one comparison.
ARC-AGI-3 scoring is nonlinear: a level score squares the ratio between an AI's move count and the human median. Discontinuous jumps can happen for that reason.
The AA Intelligence Index is only 61, which surprises me. Its ratings put models such as Gemini 3.8 Flash, Opus 5, and Grok 4.6 implausibly close together for what I care about.
The company is making large claims, charging more, and still has not released the model publicly. I would rather see broad access and independent testing.
Calling closed betas a second-class future is overblown, but restricted access can be a real problem. My applications for less-restricted models are rejected without explanation, and some models stop helping when my research mentions epidemiology.