Fable 5 reaches a 2,726 NanoGPT speedrun record, closing 81.7% of the human gap
- Fable 5 achieved the best listed NanoGPT optimizer speedrun record at 2,726, reported as closing 81.7% of the gap between the human record of 2,600 and the 3,290 baseline.
- The Fable 5 run used claude-code at high effort for 8.7 days, consuming 800M total tokens across 811 experiments and about 3,000 calls; its 24-hour score was 3,010.
- Opus 5 reached 2,920 in 2.9 days and Kimi K3 reached 2,930 in 3.6 days, the next two best listed records after Fable 5.
- The page lists 41 model, harness, and seed trajectories; results include harnesses such as claude-code, codex, prime-agent, kimi-code, grok-cli, qwen-code, pi, and muse-code.
- Several entries remain running, including Qwen3.8 Max, DeepSeek V4 Pro, Grok 4.6, Muse Spark 1.2, and GLM 5.3; GLM 5.3 has no record yet.
Hacker News opinions
I want to know how much run-to-run variation each model has. The graph only shows the best validated result per model.
I still do not understand what an optimizer run is or why it measures research ability. The post should explain that instead of sending readers to Anthropic's evaluation.
My reading is that each model gets eight attempts to reduce nanoGPT below 3.28 loss under time and token limits. But 18 models times eight runs is 144, not the claimed 153 runs.
They are asking models to improve a small model's training, repeatedly test changes, and use results to choose the next change. Different seeds repeat the autonomous session to measure variance.
I do not see why Fable 5 ran at high effort while Opus 5 ran at max. Different effort settings make this look like an experimental error, even if models can now change effort dynamically.
The claim that good traces preserve weak signals makes me wonder whether a harness with a better history log would change the result. Small prompt changes such as asking the model to track weak signals or consider novel ideas might also matter.
I am watching the still-running Grok 4.6 run. I think it is a good fit for this benchmark and plan to check again in a few days.
I co-authored METR's NanoGPT speed-run post, and I worry about contamination if these models start from the original baseline. The token-scaling plots also show some models have not plateaued, so this is a performance-at-cost bound, not a full capability ceiling.
I am confused by GPT-5.6 Sol because its note says it spent much of the run waiting. That may stretch its time-axis result, and notes about older serial versions of program.md mean the comparisons are not fully apples-to-apples.
Grok's result looks poor to me. I cannot tell how much is model weakness versus a bad harness, but building a better harness is much cheaper than building a model.
Models that have spent long periods in multi-day LLM training environments may have an advantage on long-horizon benchmarks like this. Others may simply lack that experience.
GPT-5.6 Luna looks unusually good for a cheap model. It is near Sonnet-level performance here.