Cerebras says OpenAI GPT-5.6 Sol reaches 750 tokens/s in limited Ultrafast API tier
- Cerebras and OpenAI are previewing Ultrafast Mode, a limited OpenAI API tier that runs GPT-5.6 Sol at up to 750 output tokens per second; access starts with a select customer group.
- In Cerebras' evaluation, GPT-5.6 Sol Ultrafast completed all 2,500 Humanity's Last Exam questions in 11 hours 11 minutes, versus 78 hours 27 minutes for Claude Fable 5; Cerebras says the results had comparable accuracy.
- Cerebras reports a 5.6x end-to-end speedup for GPT-5.6 Sol Ultrafast on GDP-Val with no quality degradation, based on its July 31 test using medium reasoning in Codex.
- Cerebras attributes the inference speed to its wafer-scale architecture, which places 44 GB of SRAM on each wafer-sized chip so model weights remain on-chip while token processing is pipelined across wafers.
Hacker News opinions
I checked the corresponding OpenAI post and found no pricing. That may mean enterprise-only pricing, or that they are still testing demand before setting a price.
I read that access is expanding to companies that apply and explain their use cases. It seems real, but tightly limited for now.
I suspect OpenAI mainly wants this internally for research workloads with serial bottlenecks. That may matter more than the public partnership.
I am excited about faster inference because speed is underrated. I used Cursor Composer over frontier models for a while largely because it was so fast.
I have been trying DeepSeek Flash this week, and now I want frontier models to feel just as responsive.
What do people actually need this speed for? My constraints are creativity, attention, and budget, and I can batch reviews on slow local models overnight.
Gemini 3.7 Flash may no longer sit on the speed-to-intelligence Pareto frontier.
Price still matters, though.
I would love a thumbnail-sized, replaceable local LLM accessory someday: fully offline and faster than current personal hardware. I do not know the hardware path, but the idea is appealing.
In 5 to 10 years, good phones may run GPT-3 to GPT-4-class models. Memory-heavy laptops can already run GPT-OSS 120B or full Gemma 4 interactively, while phones handle smaller models for tasks such as overnight photo tagging.
The HLE result sounds impressive, but 2,500 independent questions are embarrassingly parallel. I would rather see latency for one hard, complete HLE answer.
I expect Sol goes first while capacity is constrained. I cannot imagine the margins Cerebras will charge for this.
Luna is already adequate for much of my work because speed turns an async task into a real-time interaction. An Ultrafast Luna tier would be great, but I doubt Cerebras capacity will be priced anywhere near current plans.
The comparison chart leaves out Mimo v2.5-Pro Ultraspeed, released in June. It reaches 1,000 tokens/s, scores about 40% lower, and may cost under one-tenth of Sol for many coding jobs.