Anthropic's Claude Haiku 5.5 runs 75% cheaper and hits 1620 on GDPval-AA, with a 100K-token price cliff
- Anthropic released Claude Haiku 5.5, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, rising to $0.50 and $2.50 above that cutoff. On average it costs about 75% less to run than Haiku 4.5.
- Haiku 5.5 scores 1620 on GDPval-AA v2.1 and 72.4% on OSWorld 2.1, against 735 and 15.7% for Haiku 4.5, and reaches 57.4% on Humanity's Last Exam with tools versus 18.7% for its predecessor.
- It is the first Haiku-class model with an adjustable effort setting (Low, Med, High, Xhigh, Max), letting users trade cost against accuracy on computer use, knowledge work, and reasoning benchmarks.
- Anthropic also halved the price of Sonnet 5.5 cache reads, which makes Sonnet 5.5 about 20% cheaper on most agentic work, and added a monthly API credit for Claude Max and Team subscribers.
- In customer testing Anthropic cites, Asana measured over 30% lower task latency and up to 2.5x faster inference per agent turn, while HubSpot's CRM suite scored 92.8% averaged over three runs, its best result on that suite.
Hacker News opinions
Do we have an LLM equivalent of Moore's Law? How often should we expect improvement in this tech, and on what timeline?
Epoch AI has the number: the cost of hitting a given performance level has fallen about 47% per quarter since 2023, roughly 13x per year.
My take on why smaller models keep working: they can now forget useless information because they can derive it in reasoning, so the model gets smaller at the cost of burning more reasoning tokens.
For about five years they've gotten 90% more efficient at the same quality every 18 months, and there's little sign of it slowing. You'll know it's over when the gap between 7B and 32B stops shrinking.
Glad about the Sonnet cache read cut. That was basically forced to match GPT-6.1 Sol, costs and caching are now equal. Under the old cache prices Sonnet 5.5 made zero sense next to Opus 5.5.
The pricing is weird though. 100k tokens is an absurdly low cutoff and it only applies to Haiku, not Sonnet or Opus. Any real agent workflow blows past it fast. Still way cheaper than Haiku 4.5's $1 input and $5 output.
Luna has a threshold too, just higher. OpenAI's site says prompts over 272K input tokens get charged 2x input and cache rates and 1.5x output for the whole request.
The announcement says 90% of Haiku 4.5 requests were under 100,000 tokens, so the cutoff covers most of the actual traffic. Classification and summarization work fits easily.
Watch the tokenizer, not just the raw numbers. Modern Claude's 100K tokens are about 60-65K modern GPT tokens, so in practice Luna's cutoff is much further away than Haiku's.
Flat per-token pricing was always the strange part. Neither encode nor decode is linear in compute, so providers price for the average expected length. This is just getting closer to true cost.
If you can't get coding done with 100K context that's a broken model, a broken harness, or a skill issue. I run Haiku in task and explorer subagents and plenty of my sessions cap out well below that.
Who uses Haiku when Mimo or GLM cost 10% of this with smarter models?
Haiku 5.5 is noticeably smarter than GPT-6 Luna, so the pricing strategy makes sense. In some coding benchmarks it even beats Sonnet 5. Intelligence per dollar has grown a lot in a few months.
Worth noting they list classification as a use case, which is Jev's territory, and the $0.10/M matches GPT-6 Luna behind OpenAI's Decisions API. Both are still 2.5x Jev's $0.04/M.
GDPval-AA going from 735 on Haiku 4.5 to 1620 here is staggering. The 100k pricing looks like a hedge against OpenAI's Decisions API and the open source Jev alternatives.
It's about time Haiku got an update. And these benchmarks make it look like it smokes Luna, so I'm excited to put it through its paces.