Cerebras CS-4 rack systems juice chips for every last drop of AI performance
Jump to main content
Search
REG AD
SYSTEMS
Cerebras CS-4 rack systems juice chips for every last drop of AI performance
Next-gen systems double per-chip performance while cramming 3x as many into a rack
Tobias Mann
Tobias<br>Mann
SYSTEMS EDITOR
Published<br>wed 19 Aug 2026 // 01:00 UTC
If high-speed AI inference is what you’re after, memory bandwidth is the bottleneck to beat. At a mind-numbing 21.6 petabytes per second (PB/s) of memory bandwidth, Cerebras' dinner-plate-sized AI accelerators were already 1,000x faster than Nvidia's or AMD’s best GPUs.<br>The chip newcomer unveiled its next-gen Wafer Scale Engine (WSE) and Nexus rack systems on Tuesday. Cerebras aims to extend that lead by boosting throughput per watt tenfold over the previous generation.<br>Putting the 'T' in Turbo
REG AD
Cerebras accomplishes this in a couple of ways. But, from what we can tell, the primary lever comes from squeezing its chips for every hertz they’ve got. The newly announced WSE-3T — the “T” here stands for “Turbo” — promises twice the compute, memory fabric, and I/O bandwidth of the now two-year-old WSE-3.
REG AD
Yet, if you look at the chart below, you’ll notice it accomplishes this using the same process tech, wafer area size, transistor count, core count, and SRAM capacity. That's because the WSE-3T isn't new silicon. Instead, Cerebras tells us it's just pushing its existing wafer scale engine harder.
WSE-3<br>WSE-3T
Wafer<br>46,225 mm2<br>46,225 mm2
Process Node<br>TSMC 5nm<br>TSMC 5nm
Transistor count<br>4 Trillion<br>4 Trillion
Cores<br>900,000<br>900,000
SRAM<br>44 GB<br>44 GB
Sparse FP16<br>125 PFLOPS<br>250 PFLOPS
Dense FP16<br>12.5 PFLOPS<br>25 PFLOPS
Memory bandwidth<br>21.6 PB/s<br>43.2 PB/s
I/O bandwidth<br>1.2 Tbps<br>2.4 Tbps
TDP (wafer)<br>15 kW<br>33 kW estimated
TDP (System)<br>23 kW<br>46kW estimated
The main innovation this time around seems to be related to power delivery, which is apparently so efficient that they’re able to push twice the power through the chip, which “enables higher operating frequencies and faster token generation.”<br>How much higher does it clock? By our estimate, Cerebras is now running the silicon at 2.8 GHz, up from 1.4 GHz last gen, which would be quite the accomplishment.<br>In any case, each WSE-3T boasts 250 petaFLOPS of AI compute, 44 GB of SRAM (that’s not a typo, there really is that much SRAM on there), good for 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity.<br>On paper that sounds more impressive than it really is. AMD and Nvidia’s latest GPUs offer 4 to 5 petaFLOPS of dense FP16 compute or 35 to 50 petaFLOPS at FP4. Cerebras’ headline performance figure relies heavily on sparsity, which as a general rule doesn't benefit LLM inference.<br>Assuming the same 10x sparsity we saw with the WSE-3, the WSE-3T’s dense FP16 performance should be closer to 25 petaFLOPS, which is still impressive, just not as impressive as the chipmaker would have you believe.<br>We also suspect the WSE-3T’s peak memory bandwidth is purely theoretical. During LLM inference, the WSE-3 lacked the compute necessary to saturate its SRAM on its own, and we have no reason to believe the Turbo variant will be any different.<br>However, this time around Cerebras isn’t trying to run the entire inference stack on its own accelerators. Instead, it has partnered with Amazon Web Services (AWS) and AMD to offload the compute-intensive prompt processing bits of the inference pipeline onto their respective Trainium XPUs and Instinct GPUs.
REG AD
At least for inference, Cerebras’ chips now function primarily as decode accelerators, similar to how Nvidia is using Groq — not to be confused with Elon Musk’s Grok family of models — LPUs in its LPX rack systems.<br>The major benefit for Cerebras is its chips have a whack ton of SRAM on board. So, instead of needing 2,000 LPUs to run a trillion-parameter model, Cerebras can get away with using a few dozen, depending on the precision at which the weights are stored.<br>Curiously, Cerebras opted to double performance this generation rather than boost SRAM capacity, which hasn’t increased meaningfully since the WSE-2 launched five years ago.<br>In a disaggregated inference environment where prefill is handled by GPUs, we’d have expected to see Cerebras prioritize SRAM capacity over compute. However, given that these disaggregated compute architectures are a relatively new phenomenon, it’s possible Cerebras was already too far along in production to pivot.<br>This likely explains the Turbo naming convention. If Cerebras plans to continue down this path, we expect the WSE-4, which is presumably still coming, to offer only modest performance gains at FP16 while roughly doubling SRAM capacity. Our sibling site The Next Platform has drawn up some predictions of what the WSE-4 might look like if you’re interested.<br>Cerebras goes rackscale<br>Cerebras' latest generation of wafer scale accelerators also sees the company get serious about rack-scale compute...