Artificial Analysis v4.2 adds private agentic and PDF tests, putting Claude Fable 5.1 first
- Claude Fable 5.1 leads Artificial Analysis's Intelligence Index v4.2, followed by GPT-6 Astra, which scores 4 points above GPT-5.6 Sol; Meta ranks third among labs.
- The update adds AA-Briefcase, a private agentic knowledge-work evaluation built around multi-week projects with linked tasks and thousands of source files, graded by rubrics and pairwise comparisons.
- Surge AI's GDP.pdf tests single-turn reasoning over 100 PDFs across 10 domains and 4,592 pages; GPT-6 Astra leads its all-pass rate at 33.2%, ahead of GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.
- Private held-out tests now account for 40% of Index weighting, double v4.1's share; the sets include AA-Briefcase, AA-Omniscience, and CritPt solutions to reduce benchmark gaming.
- The index removes GPQA Diamond because Artificial Analysis says models have saturated it, and changes grading systems for AA-LCR, GDPval-AA, AA-Briefcase, and SciCode. The changes include corrected answer keys, Elo re-anchoring, and sandbox fixes for slow correct code.
Hacker News opinions
ARC-AGI-3 is missing from this index, and I think it would materially change the rankings.
I want to compare v4.2 with the previous version, but I can only find images posted on Artificial Analysis's X account.
I am not sure which index version the intelligence-versus-cost graph uses, since they apparently did not run every model under v4.2. It changed substantially earlier today, so I suspect it is v4.2.
The old index putting Astra level with Sol looked silly, but rushing an update that makes Astra look better also feels unscientific.
I do not see evidence that Artificial Analysis lets anyone peer-review its process. They cite a 95% confidence interval under plus or minus 1% from more than 10 repeats, but do not publish the outcome data, methodology, variables, or calculation behind that claim.
Updating a methodology after seeing bad results is part of real science, but it creates bias when fixes happen only after results look wrong. Artificial Analysis should commit in advance to a fixed update cadence regardless of rankings.
With so many private benchmarks, I worry they tried combinations until they got the desired result. That would discredit the index if changes track social-media expectations.
Benchmark addiction is a problem here. Model families have different strengths and weaknesses, and aggregate scores rarely explain those differences.
Why call an update unscientific when you agree the previous score was wrong? This is a benchmark for very new technology.
I find AA-Omniscience more useful than the aggregate index because it rewards correct answers, penalizes hallucinations, and does not punish refusals. A model I can trust to say it does not know is more useful than one trained to answer everything.
That matches my production experience with narrow structured LLM tasks. A confident wrong answer on unusual inputs breaks pipelines, while a model that admits uncertainty on edge cases is much easier to use safely.
I disagree that their scoring captures model differences well enough. Fable, Opus, Sol, and Astra differ substantially in practice, and Muse and 3.8 ranking highly recently makes this feel like an LMArena-style PR metric.
Hallucinations matter, but Omniscience appears to test memorized knowledge. I care much more about whether a model faithfully processes supplied context and tool-call results.
Astra's token efficiency is a real achievement. At maximum effort it has the second-highest score and the third-lowest output-token count among the models shown.
The advantage may be overstated because Artificial Analysis ran Astra at six effort levels while most other models ran only at maximum effort. Token counts also depend on tokenizers, so cost-per-task is a more meaningful comparison.
The timing gives OpenAI a boost, even if nothing improper happened. It would have looked better if the update had arrived before the Fable 5.1 and GPT-6 releases.