Cerebras CS-4 rack systems juice chips for every last drop of AI performance

joebuckwilliams1 pts0 comments

Cerebras CS-4 rack systems juice chips for every last drop of AI performance

Jump to main content

Search

REG AD

SYSTEMS

Cerebras CS-4 rack systems juice chips for every last drop of AI performance

Next-gen systems double per-chip performance while cramming 3x as many into a rack

Tobias Mann

Tobias<br>Mann

SYSTEMS EDITOR

Published<br>wed 19 Aug 2026 // 01:00 UTC

If high-speed AI inference is what you’re after, memory bandwidth is the bottleneck to beat. At a mind-numbing 21.6 petabytes per second (PB/s) of memory bandwidth, Cerebras' dinner-plate-sized AI accelerators were already 1,000x faster than Nvidia's or AMD’s best GPUs.<br>The chip newcomer unveiled its next-gen Wafer Scale Engine (WSE) and Nexus rack systems on Tuesday. Cerebras aims to extend that lead by boosting throughput per watt tenfold over the previous generation.<br>Putting the 'T' in Turbo

REG AD

Cerebras accomplishes this in a couple of ways. But, from what we can tell, the primary lever comes from squeezing its chips for every hertz they’ve got. The newly announced WSE-3T — the “T” here stands for “Turbo” — promises twice the compute, memory fabric, and I/O bandwidth of the now two-year-old WSE-3.

REG AD

Yet, if you look at the chart below, you’ll notice it accomplishes this using the same process tech, wafer area size, transistor count, core count, and SRAM capacity. That's because the WSE-3T isn't new silicon. Instead, Cerebras tells us it's just pushing its existing wafer scale engine harder.

WSE-3<br>WSE-3T

Wafer<br>46,225 mm2<br>46,225 mm2

Process Node<br>TSMC 5nm<br>TSMC 5nm

Transistor count<br>4 Trillion<br>4 Trillion

Cores<br>900,000<br>900,000

SRAM<br>44 GB<br>44 GB

Sparse FP16<br>125 PFLOPS<br>250 PFLOPS

Dense FP16<br>12.5 PFLOPS<br>25 PFLOPS

Memory bandwidth<br>21.6 PB/s<br>43.2 PB/s

I/O bandwidth<br>1.2 Tbps<br>2.4 Tbps

TDP (wafer)<br>15 kW<br>33 kW estimated

TDP (System)<br>23 kW<br>46kW estimated

The main innovation this time around seems to be related to power delivery, which is apparently so efficient that they’re able to push twice the power through the chip, which “enables higher operating frequencies and faster token generation.”<br>How much higher does it clock? By our estimate, Cerebras is now running the silicon at 2.8 GHz, up from 1.4 GHz last gen, which would be quite the accomplishment.<br>In any case, each WSE-3T boasts 250 petaFLOPS of AI compute, 44 GB of SRAM (that’s not a typo, there really is that much SRAM on there), good for 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity.<br>On paper that sounds more impressive than it really is. AMD and Nvidia’s latest GPUs offer 4 to 5 petaFLOPS of dense FP16 compute or 35 to 50 petaFLOPS at FP4. Cerebras’ headline performance figure relies heavily on sparsity, which as a general rule doesn't benefit LLM inference.<br>Assuming the same 10x sparsity we saw with the WSE-3, the WSE-3T’s dense FP16 performance should be closer to 25 petaFLOPS, which is still impressive, just not as impressive as the chipmaker would have you believe.<br>We also suspect the WSE-3T’s peak memory bandwidth is purely theoretical. During LLM inference, the WSE-3 lacked the compute necessary to saturate its SRAM on its own, and we have no reason to believe the Turbo variant will be any different.<br>However, this time around Cerebras isn’t trying to run the entire inference stack on its own accelerators. Instead, it has partnered with Amazon Web Services (AWS) and AMD to offload the compute-intensive prompt processing bits of the inference pipeline onto their respective Trainium XPUs and Instinct GPUs.

REG AD

At least for inference, Cerebras’ chips now function primarily as decode accelerators, similar to how Nvidia is using Groq — not to be confused with Elon Musk’s Grok family of models — LPUs in its LPX rack systems.<br>The major benefit for Cerebras is its chips have a whack ton of SRAM on board. So, instead of needing 2,000 LPUs to run a trillion-parameter model, Cerebras can get away with using a few dozen, depending on the precision at which the weights are stored.<br>Curiously, Cerebras opted to double performance this generation rather than boost SRAM capacity, which hasn’t increased meaningfully since the WSE-2 launched five years ago.<br>In a disaggregated inference environment where prefill is handled by GPUs, we’d have expected to see Cerebras prioritize SRAM capacity over compute. However, given that these disaggregated compute architectures are a relatively new phenomenon, it’s possible Cerebras was already too far along in production to pivot.<br>This likely explains the Turbo naming convention. If Cerebras plans to continue down this path, we expect the WSE-4, which is presumably still coming, to offer only modest performance gains at FP16 while roughly doubling SRAM capacity. Our sibling site The Next Platform has drawn up some predictions of what the WSE-4 might look like if you’re interested.<br>Cerebras goes rackscale<br>Cerebras' latest generation of wafer scale accelerators also sees the company get serious about rack-scale compute...

cerebras sram systems performance compute rack

Related Articles