Anthropic's Claude Opus 5.5 prompt guide: 30% faster output tokens, medium effort matches Opus 5 at high effort
- Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and usually finishes the same task in fewer tokens, so existing Opus 5 prompts run without changes.
- At its default medium effort, Opus 5.5 matched or beat Claude Opus 5 at high effort on agentic coding in a real repository, in fewer steps and with fewer tokens, and early testers reported more bugs caught in code review with fewer false alarms.
- Even at its lowest effort setting, Opus 5.5 read values off dense charts more accurately than Opus 5 at its highest effort, using a small fraction of the output tokens, and at default effort it matched Opus 5's computer-use success rate at a much higher setting.
- Anthropic warns that a general instruction like "avoid a generic AI look" mostly swaps one default style for another; the guide says to name specific patterns to avoid and to check the first result iteratively.
- The guide organizes fixes by symptom: calibrate effort when turns cost more than on Opus 5, handle prompts written for thinking disabled, restart unattended agents that stop partway, and mark pasted text so the model does not follow instructions embedded in it.
Hacker News opinions
Opus 5.5 is a good model, but the social media hype about its 2D work is way overblown. Both this and Astra's 3D stuff leaned on third party APIs for assets, and a lot of what the model did was just coordinate everything.
Not true, Opus 5.5 needs nothing beyond some javascript and typescript libraries to make very detailed 2D and 3D visualizations. I burned a week of tokens feeling it out and the step change in visual work genuinely shocked me, I haven't been surprised like this by an LLM in a long time.
I don't even use Claude Code or any agent that touches my machine. I paste the whole codebase as a text file and asked for a 2D terminal racing game with no assets or audio, and it blew past my expectations. One unit test failed and it argued the test was wrong, not the code.
This how-to-prompt stuff changes every three months. Remember when you had to tell Claude to keep going or it would just give up? Imagine having to relearn how to drive your car every three months.
Of course it changes every three months, the models are changing in features, scope, intelligence, pricing and communication style all at once. We're in a race and it won't settle down for a while.
I skipped every Opus after 4.8 because they didn't work with my homegrown harness. 5.5 is a pretty happy match, so I'm back.
The frontend design defaults section was worth reading. I always write "don't make it look like generic AI slop" and it kind of works, but now I get why apps still end up with the same styles anyway.
How is anyone supposed to describe what they want when they can't put it into words? That advice is only usable if you already have a design eye or background.
How does the model even know what AI slop is? Telling a kid not to do wrong without saying what's wrong is about as useful as "no mutated hands" in an image prompt.
First thing it did when I tried it was roam through files in directories way outside the project. I asked it to explain why, several times, and never got anything resembling an answer.
The unattended run is the feature that matters most for coding. Just saying continue when it gets stuck makes it repeat the same error, so save the last action and result and force a new approach. If it tries the same thing twice it should stop and ask for help instead of burning money in a loop.
What keeps jumping out at me is that they refuse to give users the thinking tokens and full reasoning in the output. That drives me further away, I'm moving more of my primary workload to Chinese providers.
Yeah, it's annoying. I have to pay for thinking and I still can't see it.
I'm hitting Claude's weekly limits way sooner than before with roughly the same workload. What Chinese provider is actually on par with Claude Code for coding and agentic work?