Cerebras CS-4 claims 30x GPU inference speed with three WSE-3 Turbo wafers
- Cerebras CS-4 uses three WSE-3 Turbo wafers per system, and Cerebras claims up to 30x faster inference than production GPU systems.
- Cerebras says CS-4 delivers over 1,000 tokens per second on models exceeding 10 trillion parameters, with wafer-to-wafer latency reduced to 2 microseconds.
- The system is the first product using the Nexus Platform Architecture, which separates stable power, cooling, and networking infrastructure from modular wafer-scale compute backpacks.
- Each Wafer-Scale Backpack combines the wafer, power conversion, liquid cooling, I/O, and controls in a 3D package with 50% fewer components; Cerebras says this cuts deployment time from days to hours.
- The new programmable wafer I/O subsystem doubles I/O bandwidth and links wafers within and across racks without a switch, while power delivery sits 0.5 mm from the processor instead of roughly 50 mm on conventional GPU boards.
Hacker News opinions
We are only a few years and three or four hardware iterations into LLM-specific systems, so orders-of-magnitude gains in speed or cost over five years seem likely. More than 1,000 tokens per second on models above 10 trillion parameters is wild.
That is why I think the data-center buildout is a bubble. Hardware efficiency and speed will improve exponentially over the next decade, and GPUs are only the current mass-produced option.
Software still has plenty of room too. Qwen 3.8 27B, DeepSeek V4 Flash 0731, and GLM 5.3 suggest intelligence density, kernels, KV caching, MTP, and workload splitting can all improve.
I see the buildout as a scam built on huge debt and the premise that GPUs are the only scaling path. TPUs and ASICs can beat GPU throughput for LLMs, and better software would give developers more hardware choices.
If average ChatGPT users cost about $0.10 per month, that sounds like it would kill OpenAI and Anthropic's economics.
The announcement conspicuously leaves out power consumption.
The figure is 162 kW. Cerebras also claims 10x more throughput per watt than CS-3.
Five years from now, I do not see why Nvidia would dominate inference. Cerebras targets inference rather than training, where Nvidia may still keep a role.
Cerebras is claiming only about 2x Groq's performance, which usually is not enough to make customers switch.
Nvidia is technically well run and will keep improving its inference stack over the next five years too.
Nvidia's supply chain is its major advantage. Nobody else can produce at that scale, so I would not discount its position.
I do not think CS-4's 1,000 tokens per second figure reveals GPT-5.6 Sol's parameter count. Cerebras can raise per-user speed by assigning fewer users to each chip, and its charts omit the numbers needed to infer model weights.
The use of older open-weight models such as GLM 4.7 and Kimi K2.7 is odd when newer versions exist. I would ask whether buyers must fund software work to run recent models.
I think Cerebras runs whichever models enterprise customers pay it to run. It does not look interested in consumer demand.
AMD and Cerebras may compete with Nvidia in the near future. High margins and huge demand attract multiple competitors.
A GPU is only one part of Nvidia's position. Vera Rubin buyers get NVL72 racks alongside Nvidia networking, storage, CUDA, CUDA-X, and data-center administration software.
AMD might do well to buy Cerebras, as Nvidia bought Groq.
I doubt AMD can match Nvidia's large multi-GPU systems such as NVL144 and NVL576. Nvidia is on NVLink generation 9, while UALink is at specification version 1 with incompatible corporate implementations.