DeepSeek V4 Pro 0813 launches, undercutting Opus and Sol on price by up to 60x
- DeepSeek V4 Pro 0813 is a mixture-of-experts model priced at $0.435 input / $0.87 output per 1M tokens with a 1M token context window, listed on OpenRouter as GA released Aug 12, 2026.
- HN commenters calculate it as roughly 20x to 60x cheaper than Opus 4.8 once cache-read discounts on agentic coding workloads are factored in ($0.000875 vs $0.052 per request).
- Benchmark tables shared from r/LocalLLaMA show V4 Pro scoring 87.9 on Terminal Bench 2.1 and 83.3 on Cybergym, beating its predecessor V4 Flash 0731 but trailing Opus-4.8 and Fable 5 on most tasks.
- A geometric mean across all listed benchmarks ranks it mid-pack: Sol 65.5, Fable 5 64.5, Opus 5 64.0, DeepSeek V4 Pro 62.5, Kimi-K3 62.3, versus GLM-5.2 at just 47.3.
- DeepSeek is raising its official API prices the same day, and commenters note the release timing coincides with Qwen's launch of Qwen3.8-max weights.
Hacker News opinions
Pricing wise it's competitive with Opus 4.8 but weaker than Sol or Fable, and roughly 20x cheaper overall.
If you factor in the typical cache-read split for agentic coding (750 in, 290 out, 82k cached), it's actually more like 60x cheaper than Opus. $0.000875 per request vs $0.052 for Opus.
And unlike Sol or Fable, you can actually self-host this one or rent a GPU and run it yourself.
It's interesting how much model adoption is just momentum. These Chinese models are legit capable but devs default to whatever's already the 'industry standard'.
HN is Bay Area centric where a few hundred bucks a month on AI is nothing. Cheaper models with sketchier data policies are way more popular in developing countries.
Enterprises avoid Chinese models mostly for political risk, not capability. Nobody wants to build deep into a stack that a government ban could force them to rip out in six months.
Honestly a known model beats an unknown one even if the unknown is better, because learning a new model's quirks is mentally exhausting.
I've been running Deepseek Flash full time this week to evaluate dropping Anthropic entirely, results have been pretty positive, can't wait to try Pro tomorrow.
My company only lets me use Claude models so that's what I stick with, but I miss the old Sol model, it paired with Codex was great at first-shot understanding.
I keep using Claude and Codex anyway because the subscription rates are so much cheaper than paying per token, even with cheaper models available.
There's no room in people's attention for models that are neither SOTA nor absurdly cheap. Either be name brand like Anthropic/OpenAI/Google or be like 500x cheaper, anything in between gets ignored.
Looking at the benchmark table, geometric mean across all of them: Sol 65.5, Fable 5 64.5, Opus 5 64.0, V4 Pro 62.5, Kimi-K3 62.3, V4 Flash 55.8, GLM-5.2 47.3.
The timing looks deliberate, releasing the same day Qwen dropped Qwen3.8-max weights. Comparing claimed benchmarks, V4 Pro is better on average and way cheaper, so no real reason to pick Qwen3.8-max unless you need vision.
HLE scores without tools seem to track real world performance better than the tool-augmented numbers, it's like RL-tuned benchmark performance vs actual base model knowledge.
The wildest part is how close Flash scores to Pro across the board, I haven't tried the new ones yet but the gap must be bigger in practice than these numbers suggest.
Deepseek is burning money on their official API and raising prices starting today. V4 Flash 0731 is still probably the best model to come out in the last few months.