Anthropic launches Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0, 30% faster, up to 30% cheaper per task
- Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, on September 28, 2026, running 30%+ faster than Sonnet 5 and costing up to 30% less per task in Anthropic's testing.
- Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 versus Sonnet 5's 10.3%, and lands two points below Opus 5.5 on GDPval-AA (1844 vs 1846), though Anthropic says Opus 5.5 remains clearly stronger at complex open-ended work.
- List price is unchanged at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads; the savings come from needing far fewer tokens for the same task.
- Sonnet 5.5 is the first Sonnet model to ship with cyber safeguards and fallbacks like those on Opus 5.5, because its cybersecurity capabilities are comparable to Opus 5's; biology safeguards match Sonnet 5's, and Anthropic says routine software development is unaffected.
- Claude Haiku 5.5, aimed at high-volume and cost-sensitive applications, joins the Claude 5.5 family in the coming weeks.
Hacker News opinions
Benchmarks for agentic coding basically stack up 1:1 with Opus 5.5. Terminal-Bench 70.6 vs 66.4, FrontierCode 52.1 vs 54.4, CursorBench 55.5 vs 57.8. Opus 5.5 is the best model I've ever used and Sonnet 5.5 matches it on some of these.
This is long overdue. Sonnet 5 was terrible API value for agentic coding, open models like GLM-5.3 Flash beat it at 1/20th the price. The OpenAI and Anthropic lead is vanishingly small now.
I pay $200 a month and I'm in their Cyber Verification Program, but I can't use Opus 5.5 or Sonnet 5.5 for authorized bounty work. I get flagged for Cyber immediately. The safeguards are shit.
Yeah, as far as I can tell the Cyber Verification Program does absolutely nothing. Read the docs though, it says the program doesn't cover Opus 5.5 yet and they'll expand it to 5.5 soon.
I got flagged just for using the word fuzz. It was a parser, so security adjacent, but still. What's the point of a verification program if it doesn't skip those checks?
I had it look at 30-year-old C code I wrote in college and it tripped a guard rail. The code was full of buffer overflows, but I already knew that.
Wanted to do ESP32 bluetooth presence detection for my smarthome and Claude flagged it and degraded me to Sonnet 4.6. Went to Codex, no issues.
I was verifying a write-ahead log implementation with Opus 5.5 and got flagged, forced down to Opus 4.8. Switched to OpenCode plus OpenRouter and kept working. My trust in Anthropic evaporated the moment I started getting blocked.
Notice they omit Fable numbers from every chart and only show Opus, Sonnet and OpenAI models. Maybe Fable is out the door.
Fable is off the price/performance pareto frontier now. They'll probably drop an updated Fable that's frontier intelligence until the next Opus.
Models are getting more efficient way faster than they're getting more intelligent. From a marketing angle that's the impressive story, and Fable would just look orders of magnitude more expensive for marginal gain.
Loving the tit for tat cost charts. A few days ago OpenAI ruled the cost pareto frontier, now Anthropic takes it back. See you all same time next week?
Can't wait until this becomes like an internet subscription: unlimited tokens 24/7 for a low fixed monthly price.
Sonnet 5 felt benchmaxxed to me, and so did Opus 5. I'll give 5.5 a long eval period before deploying it with enthusiasm the way I did with 4.6. 5.5 seems a lot better so far, but still not close to Fable on quality.
Sounds like for Anthropic models we hit peak cyber capability with Opus 4.8. Everything after that just falls back to worse models for anything security related.
Daybreak Blue isn't bad and the bar to get into OpenAI's program is reasonable if you want an alternative.