All posts
// / Blog

Cerebras' CS-4 is 30x faster than GPUs. It still won't help most production AI I've shipped.

Cerebras shipped the CS-4 this week. Three wafer-scale chips in one rack, 750 PFLOPs of AI compute, and a claim of up to 30x more tokens-per-second-per-user than GPU-based inference on comparable workloads. Memory bandwidth alone is 129.6 petabytes per second across the system.

That's not a typo. That's roughly 2,000x the bandwidth of a next-gen GPU on that one metric, according to Cerebras' own numbers and the early third-party coverage I've read. Wafer-to-wafer latency is down to 2 microseconds. The architecture, on paper, can support models north of 50 quadrillion parameters, which is a number so large it stops meaning anything useful and starts meaning "we haven't hit the ceiling yet."

I build inference infrastructure for a living, so my first reaction was genuine respect. This is real engineering, not a slide-deck chip. Three 4-trillion-transistor wafers wired together with 2-microsecond hops is the kind of thing that used to be a research paper, not a shipping rack.

My second reaction was: this solves a problem I almost never have.

Here's the split I see in production AI, having shipped systems across automotive, defence, and education. There are workloads where the constraint is throughput at scale — you're serving a huge model to a huge number of concurrent users, and shaving milliseconds off time-to-first-token across millions of requests is worth real money. CS-4 is built for exactly that. If you're OpenAI-scale or running a foundation-model API, a 30x throughput gain on GPU-class workloads is a genuine unlock, and the SRAM-heavy, no-HBM design is a legitimately different bet than Nvidia's roadmap.

Then there's the workload I actually spend most of my time on: real-time inference where the datacenter isn't in the loop at all. A drone threat detection system for defence deployment doesn't get to make a network call. Neither does an offline multilingual avatar running in a Mercedes-Benz showroom with no guaranteed connectivity, or a factory floor camera pipeline that has to flag a defect before the part moves three feet down the line. In those systems, the fastest rack in a datacenter a thousand kilometers away is irrelevant. The model has to run where the data is, full stop.

This is the split the industry keeps glossing over when a headline like "30x faster than GPUs" lands. Faster centralized inference and faster edge inference are not the same problem, and they don't share a solution. One is about squeezing more tokens per second out of a cluster you control. The other is about squeezing a capable model onto hardware you can hold in your hand, with no fallback to the cloud if the connection drops — which, in a defence or industrial context, it will.

I'd also flag the "50 quadrillion parameters" figure for what it is: an architectural ceiling, not a trained model anyone is running today. It's the kind of number that's technically true and practically meaningless right now, and I'd rather see Cerebras publish real workload benchmarks against real deployed models than headline figures nobody can currently use. The Register's own coverage makes the same point — theoretical peak and delivered throughput on actual LLM workloads are different conversations.

None of this is a knock on the hardware. If you're running mega-model inference at the scale where a rack like this pays for itself, CS-4 is a serious option and probably the most interesting non-Nvidia architecture I've seen ship this year. Wafer-scale integration solving the memory-bandwidth wall is a real contribution, not marketing.

But if you're building the kind of system I mostly build — something that has to work at the edge, offline, under your control, where the "AI" has to survive without a network — this launch changes nothing about your architecture. Your answer is still a smaller, well-quantized model running locally, not a faster path to a datacenter you can't always reach.

Two different problems, two different chips, two different roadmaps. Know which one you're actually solving before you get excited about the benchmark.

#EdgeAI#AIInfrastructure#OnDeviceAI#ProductionAI#DefenseAI