GPT-6 Astra Scores 99.9% on ARC-AGI-3 With Persistent Reasoning State
- OpenAI's GPT-6 Astra scored 99.9% on ARC-AGI-3 Semi-Private at high reasoning effort with the Provider Adapter harness, at a reported cost of $18,817.
- With ARC Prize's Standard harness, Astra at maximum reasoning effort scored 62.7% for $26,098; the harness lets the model retain notes it chooses across the environment.
- The Provider Adapter harness preserves opaque reasoning state between requests and compacts long conversations, letting Astra reuse prior work rather than reconstruct it on each request.
- Astra used fewer actions than the median tested human on 96% of levels, surpassing the human baseline for action efficiency.
- ARC Prize reports that Astra built compact symbolic world models, expressing game mechanics as logical rules and creating its own domain-specific shorthand for state tracking and planning.
Hacker News opinions
That score and cost curve is wild. DeepSeek v4 Flash recently showed the same counterintuitive pattern where more reasoning costs less.
I do not think raw brain energy is a fair comparison. A human who can do competent work on demand requires considerable investment beyond the electricity their brain consumes.
A person's time is not worthless. It is probably the most valuable resource we have, so pricing a human task at brain-energy cost misses the point.
99.9% with the right harness sounds like AGI to me. I expect people will move the target to human cost or efficiency, then invent another human-friendly test when that gap closes.
AGI has no fixed meaning beyond what each person projects onto it. That is why arguments about whether the goalposts moved never end.
We already saw this with Opus 5 and Nvidia's harness hitting 100% on ARC. The official ARC setup discards prior context, while OpenAI's adapter keeps reasoning state with compaction, which is much closer to a normal deployment.
The article says Astra clarifies what remains out of reach, but I do not see it identify those capabilities. It mostly shows a model becoming more superhuman on ARC-AGI-3.
Solving a snake-like puzzle in few moves is not my definition of intelligence. Basic inferential logic with memory still feels far short of general human capability.
If the model has never encountered the game, figures out that it is snake-like from a general input domain, and then solves it efficiently, that is evidence of intelligence.
ARC keeps getting new versions because earlier versions saturate. A year and a half ago, Gemini Pro 2.5 reportedly needed a page of reasoning for every tic-tac-toe move, so the progress is real even if the task alone does not prove AGI.
Anything with a verifiable answer will eventually be saturated by a model. The remaining hard evaluations may be subjective taste.
Verifiability is not the whole issue. Predicting a coin flip, earning $100, or increasing paid subscriptions in an A/B test are easy to verify but still hard to learn, especially under a tight budget.
The $19K to $50K figures add up to hundreds of thousands over the tests. I want to know who paid for all those model calls and whether OpenAI supplied an effectively unlimited API key.
The no-reasoning result is strange: 35.2% on the Standard harness beats Opus 5 on high, while low reasoning is much worse. I wonder whether the API's 'none' setting defaults to something like medium.
I do not care about benchmark puzzle games until a system can spontaneously fix my leaky faucet. If it cannot do that, it does not solve the problem I actually have.
AGI is usually capability across cognitive tasks, not consciousness, emotions, or sentience. I prefer recursive self-improvement as a more measurable target.