AMD and Cerebras join forces against Nvidia’s Groq LPUs
Jump to main content
Search
REG AD
SYSTEMS
AMD and Cerebras join forces against Nvidia’s Groq LPUs
The enemy of my enemy is my friend
Tobias Mann
Tobias<br>Mann
SYSTEMS EDITOR
Published<br>thu 23 Jul 2026 // 22:33 UTC
GPUs are great for training, but for inference, you need a heavy dose of speedy memory to churn out the tokens. AMD has tapped Cerebras Systems to develop a disaggregated compute platform combining Instinct GPUs with the chip startup's SRAM-powered AI accelerators. The goal: to deliver ultra-low-latency inference for agentic workloads.<br>The collaboration, announced on stage during AMD CEO Lisa Su's Advancing AI keynote Thursday, closes a gap in AMD's portfolio that cost Nvidia $20 billion to acquihire from Groq back in December.<br>Cerebras CEO and cofounder Andrew Feldman is no fan of Nvidia, having previously denigrated the GPU giant as a mere AI arms dealer. And unlike GPUs, Cerebras' wafer scale engines (WSE) don't rely on HBM4 but instead use on-chip SRAM that's orders of magnitude faster.
REG AD
This has made Cerebras one of the fastest inference providers in the world, with output speeds often exceeding 2,000 tokens a second.
REG AD
By running compute-heavy prompt processing operations on AMD's Instinct GPUs and offloading the memory intensive token generation to Cerebras' WSE accelerator, the duo aims to achieve higher interactivity without compromising on throughput or cost to do it.<br>“What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity. And you marry that with our Wafer Scale Engine, which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched,” Feldman said on stage.<br>Neither company has shared specific figures, but the combination is expected to boost the number of tokens per second generated per watt of electricity consumed by as much as 5x.<br>If any of this sounds familiar, Cerebras' accelerators fill the same role as the Groq 3 LPUs (Language Processing Units) announced alongside Nvidia's Vera Rubin rack systems at GTC in March.<br>But where Nvidia needs two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen.<br>The combined offering will be available in Cerebras Cloud later this year, but may not be AMD's last deal with the upstart.
MORE CONTEXT
AMD attacks the rack with Helios systems that rival Nvidia's
Intel-backed AI chip startup SambaNova breathes new life into aging Nvidia GPUs in latest benchmarks
AMD’s Ryzen AI Halo makes local AI look easy, but at $4K, easy doesn't come cheap
Qualcomm's proposed solution to catch up in AI infra: Bury the compute under the DRAM
“There are lots of ways to get workload-specific acceleration done, and I think Cerebras has a very interesting technology. It works very well with Helios,” Su said during a press conference following the keynote. “The idea of our open ecosystem is frankly that we will work with a number of different companies that may have technology that could be useful.”<br>“You can expect that we're going to do more workload disaggregation going forward,” she added. ®
ai infrastructure month 2026<br>systems<br>ai<br>amd<br>datacenter<br>gpu<br>cerebras
REG AD
off-prem
Microsoft fiber foul-up cut off Azure California for almost five hours
Maintenance mistake caused immediate issues and took out 27 services
Security
OpenAI-Hugging Face attack doesn't mean agents are evil – unless you tell them to be
Attack models gonna attack
Gobi X: Creating more energy for AI, not taking it from society
PARTNER CONTENT: How Envision is reversing the datacenter playbook by making computing chase abundant desert power, not the other way around
Security
Researchers replace downloaded macOS apps with evil twins, Apple shrugs
Gatekeeper has one job and it's not doing it for some software
columnists
Airbus takes flight from AWS. What happens next is critical
Which way to the Land of the Free again?
SYSTEMS
AMD and Cerebras join forces against Nvidia’s Groq LPUs
The enemy of my enemy is my friend
MOST POPULAR
off-prem
Anyone with a shed, an extension cord, a couple of GPUs and an overdraft is building datacenters. Fujitsu just offloaded five
security
Linux kernel team publishes 432 CVEs in two days
AI AND ML
OpenAI admits it was the source of the agent swarm that attacked Hugging Face
off-prem
AWS customer learns the hard way how even the smallest oversight can be mission-critical
columnists
Airbus takes flight from AWS. What happens next is critical
AI
SYSTEMS
AMD and Cerebras join forces against Nvidia’s Groq LPUs
The enemy of my enemy is my friend
AI and ML
Codeberg gives vibe-coded projects the toss, promotes human FLOSS
AI no longer welcome in human-focused community
Systems
AMD attacks the rack with Helios systems that rival Nvidia's
Spec for spec, the House of Zen's first...