xAI ships Grok 4.7 at Grok 4.6 pricing, claiming frontier price-performance on long coding tasks
- Grok 4.7 is available today in Cursor and Grok Build, plus the Grok API, third-party coding harnesses, and cloud model routers, priced at $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the output speed and twice the price. Pricing and speed match Grok 4.6 despite a larger base model.
- The model uses a larger base model than Grok 4.6 and a longer reinforcement learning run weighted toward tasks that take many hours, and was trained to natively understand the Grok Bot harness. xAI says it verifies its own work better and handles longer context.
- Benchmark scores: CursorBench 4.0 46.3% (4.6: 40.4%), DeepSWE v1.1 71.0% at high effort, EEBench 64.0%, Terminal-Bench 4.0 38.0%, Harvey Legal Agent 19.6%, HealthBench Professional 56.7%, and GDPval Elo 1695. Fable 5.1 Max still beats it on CursorBench at 51.8% and Terminal-Bench at 57.9%.
- A new safeguard stack tops LatchBio's biosafety benchmark at 62.4%, and on HackerBench v0.3 only 3.3% of risky dual-use prompts get through while legitimate security work is rarely blocked. Select cybersecurity partners get invite-only access to the model's red-team capabilities.
- The launch chart omits GPT-6 Astra, which commenters flag as a gap; OpenAI pulled out of Cursor before Astra shipped, so it was never scored on CursorBench 4.0, the benchmark that runs Cursor as the harness.
Hacker News opinions
After a few months in Cursor and trae.ai, Grok in Cursor is clearly the better of the two for me.
Maybe it's just me, but for personal chat Grok is the worst of the bunch next to Claude, ChatGPT, DeepSeek and Gemini. The personality is bland and it won't put in the work.
I ran the same prompt through Qwen, DeepSeek, Gemini and Grok on OpenRouter. Grok did the best research and produced less bullshit, especially when I asked it to be critical of an idea.
Same experience here. Grok ends the task almost immediately and says "Done!". It's the laziest and most dishonest model I use, I can't trust it with serious coding.
Do you actually want your LLM to have a personality? Personality is exactly the thing people complain about with Claude.
Grok 4.7 reportedly has 40% more weights than 4.6 at the same $2 in, $6 out pricing, and they pushed the release back almost two weeks, so xAI probably isn't thrilled with it. They also shipped the day before the rumored Opus 5.5 launch. Grok 4.5 fixed buildroot issues Fable 5 couldn't touch, and post-Cursor Groks are phenomenal at frontend web dev, though Claude is better at backend Ruby. My favorite part is that they speak plain English instead of Claudish.
Token price alone doesn't tell you much unless you know token efficiency.
What's with the graph at the top leaving out Astra? That can't be an oversight. Ah, OpenAI pulled out of Cursor before Astra existed, and CursorBench runs Cursor as the harness, so it never got the score.
Grok inside Grok Build is a solid, no-nonsense pick that stays on course. Grok bot is another surface I actually enjoy using.
I've had good runs with 4.6 and Grok Build on tscircuit, writing code with spatial reasoning and importing CAD components from other formats into tsx. Claude came in to error check and found nothing to improve across three projects. Excited for 4.7, though I share the skepticism since the price didn't move.
I tried it in Omp and it's problematic. It loops in thinking mode counting fixes up to 80, ignores AGENTS.md, and corrupts plan files. I keep 5.6 Sol as a watchdog and it blocks every turn. 4.6 wasn't this bad.
Artificial Analysis says 4.7 is a lot less token efficient than Grok 4.6.
Why are they comparing 4.7 x-high against 4.6 high? Because 4.6 never had an x-high level, xhigh is new to Grok.
Grok is really expensive. I get amazing results from DeepSeek 4.1 Flash for a fraction of the price.
DeepSeek 4.1 Flash is garbage, it produces almost nothing but trash code. It's only fine if you're doing extremely dumb things.