DeepSeek V4 Flash scores 50 on Intelligence Index, undercuts OpenAI Luna on price per task by about 2x
- DeepSeek V4 Flash 0731 (reasoning, max effort) scores 50 on the Artificial Analysis Intelligence Index, ranking #3 of 101 comparable models and well above the median of 25.
- Pricing is $0.14 per 1M input tokens and $0.28 per 1M output tokens, both far below the median ($0.43 and $1.20), with a cache hit price of $0.003 per 1M tokens, a 98% discount.
- The model has 284B total parameters but only 13B active parameters per token, and supports a 1M token context window under an MIT license with weights on Hugging Face.
- It cost $72.02 to run the full Intelligence Index evaluation on this model, and it generated 210M tokens during testing, twice the median verbosity of 100M tokens.
- HN commenters say DeepSeek V4 Flash beats OpenAI Luna by roughly 2x on price per task at similar intelligence scores, though Luna runs 2 to 5x faster depending on effort setting.
Hacker News 의견들
Nice catch on the URL, the one you posted 404s but the model page itself works fine.
Deepseek Flash already beats OpenAI Luna on price per task by roughly 2x, high effort Luna is $0.03 at index 46 while Flash max effort is about $0.03 at index 50.
A fairer read is Luna costs 2 to 3x more for similar or slightly better performance, but Luna also runs 2 to 5x faster depending on effort level.
If only Deepseek let you opt out of training use, it might actually be viable for me.
The output tokens per Intelligence Index chart looks off, it shows Kimi K3 Max reasoning less than Flash and even gpt-oss-120b, which does not match my experience of K3 being the most verbose reasoner I use.
I tried the preview on my own codebase and it used over double the tokens of the prior version on a simple prompt, felt a lot more eager too.
If v4 Flash is already beating V4 Pro, are we getting a new V4 Pro that matches Opus 5 soon?
I've been building an app on v4 Flash and it's shockingly cost effective, good enough that I can offer a generous free tier.
Deepseek said in the v4 Flash announcement that an updated V4 Pro is coming soon.
Too bad Flash still has no multimodal support, otherwise it would beat GPT 5.6 SOL on value.
Does it dodge questions about Tiananmen Square or answer straight now?
It's open weight so you can uncensor it yourself, don't blame the researchers for what their government mandates.
Western models censor plenty too, just different topics, so this whole line of complaint feels lazy at this point.
No tokens per second benchmarks on this page? On OpenRouter I'm seeing 93 TPS.
New Deepseek drops feel like Christmas, nobody beats them on cost, and API pricing feels like the more sustainable business model versus subsidized subscriptions.
My coworkers burn through their Claude quota in an hour while I run Deepseek Flash all day, only pulling out Claude for the genuinely tricky stuff.
Looking at the numbers myself, Flash needs about 5x the tokens to match GPT 5.6 Luna's results, so it's not a clean win despite the low price.
I've burned tens of millions of tokens on Deepseek for pennies, it crushes some tasks and just isn't great for others.
Now let's see Anthropic cut its prices, honestly I think they can't survive this kind of competition without reacting fast.
This website is unbelievably heavy and slow, we need a benchmark for the benchmark site at this point.