xAI releases Grok 4.6, claims parity with GPT-5.6 Sol on composite benchmark
- xAI launched Grok 4.6, built on Grok 4.5, targeting long running agents and more ambitious visual and interactive project work.
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index (a nine benchmark composite), matching GPT-5.6 Sol's 61 and trailing Fable 5 Max's 62.
- Training added a longer supplemental run with curated model generated data, an improved optimizer, and SFT trajectories regenerated by Grok 4.5 across reasoning efforts and domains like STEM and software engineering.
- Pricing starts at $2 per million input tokens and $6 per million output tokens, with a fast variant at double that price, and the model is available today in Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare.
- xAI is offering 2x included usage in Grok Build and Cursor for the first week, and says safeguards were recalibrated with its widest ever pre-deployment testing suite.
Hacker News opinions
This dropped right after DeepSeek-V4-Pro-0813, feels timed on purpose. Both probably came right after Qwen 3.8 max too, everyone's racing to match each other.
Cursor's blog post confuses me. The xAI acquisition of Cursor hasn't even closed yet, what are they doing shipping this alongside Composer?
Fable level performance but faster and way cheaper, that's genuinely impressive if it holds up.
Beats GPT-5.6-Sol on most benchmarks and it's cheaper than Kimi K3 on the API, plus generous usage on Cursor subscriptions. I'm tempted to switch off Opus for cost reasons alone, but Opus 4.8 has been so good it's hard to leave.
In my testing Grok 4.5 wasn't Opus level, more like between Sonnet and Opus, closer to Sonnet honestly. We'll see if 4.6 changes that.
I won't touch xAI models because of the Nazi stuff. Musk tinkering with the RL to make it sound like him is wild, don't care how cheap or smart it is.
People downvoting the Nazi comment, are they wrong though? It's literally generated Nazi content and CSAM and xAI is in court over it.
Pricing page still shows 4.5, they haven't even updated it yet.
Anyone else notice that within two months of Fable releasing, every major lab suddenly has a Fable-level model? Doesn't add up unless it's benchmark hacking or there's no real moat between these labs.
Fable is basically Mythos which was previewed back in April, so it's really been four months not two, that changes the math a bit.
I'm solidly in the benchmaxxing camp. I gave GPT-5.6-Sol an auth ticket, walked away, came back to 25,000+ lines of code that another instance of GPT-5.6 called 98% garbage. Fable's theory was it lost context after compaction and kept working blindly.
Maybe there's just no moat at all between these labs anymore. GPU access, data, and a handful of published techniques get you to the same place eventually.
This is the smartphone Snapdragon cycle all over again. Whoever gets first crack at new compute looks like a genius for a few weeks until everyone else catches up on the same hardware generation.