AMD and Cerebras Launch AI Inference Solution

rbanffy1 pts0 comments

AMD and Cerebras Launch AI Inference Solution<br>Skip to main content

Products

Ecosystem

Developers

Resources

Pricing<br>Company

Contact usGet Started

*:first-child]:mt-0 [&>*:last-child]:mb-0">Cerebras and AMD Partner on Disaggregated Inference. Learn more >>

Jul 23 2026<br>AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution

*:first-child]:mt-0 [&>*:last-child]:mb-0">News Highlights<br>AMD and Cerebras are collaborating to advance a workload-optimized approach to ultra-low-latency AI inference infrastructure.<br>AMD Helios™ and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow, combining ultra-high-throughput from AMD Instinct™ GPUs, with ultra-fast token generation of Cerebras Wafer-Scale Engine.<br>Cerebras plans to deploy AMD Helios in its data centers, with the joint solution expected to be available first through Cerebras Cloud in the second half of 2026.<br>SAN FRANCISCO and SUNNYVALE, Calif. — July 23, 2026 —  AMD (NASDAQ: AMD) and Cerebras Systems (NASDAQ: CBRS) announced a technical partnership to deliver a new disaggregated AI inference solution that combines AMD Helios™ rackscale solutions with the Cerebras Wafer-Scale Engine. Unveiled at Advancing AI 2026, the solution is designed to deliver the ultra-low latency required for the most advanced AI applications while dramatically increasing the throughput and efficiency.<br>The joint AMD and Cerebras solution will deploy AMD Helios alongside Cerebras Wafer-Scale Engine technology integrated in a single inference workflow for maximum performance and efficiency. AMD Helios will provide a high-performance, scalable throughput engine. Cerebras Wafer-Scale Engine technology will provide ultra-fast, ultra-low latency decode and token generation. Together, the two compute engines are expected to deliver up to 5x higher tokens per second per watt (T/s/W) [i].<br>AI inference workloads increasingly have different requirements across latency, throughput, token capacity, cost and scale. High-volume workloads prioritize maximizing token generation, while coding, real-time copilots, live agents and agentic workflows demand faster response times. These differences are driving demand for heterogeneous infrastructure that matches compute technologies to specific workload requirements.<br>The AMD and Cerebras solution addresses this challenge through disaggregated inference, optimizing the two primary stages of the workflow independently. AMD Helios provides ultra-high throughput, processing prompts and large context windows. The Cerebras Wafer-Scale Engine accelerates the memory-bandwidth-intensive token generation, with ultra-low latency. By connecting these best-in-class engines through one integrated workflow, the companies are creating a differentiated platform for ultra-low-latency inference without sacrificing throughput or scale.<br>“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” said Dr. Lisa Su, chair and CEO, AMD. “AMD Helios delivers leadership performance and scale for the broadest range of inference workloads. Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI.”<br>“The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world’s fastest, ultra-low-latency inference,” said Andrew Feldman, CEO and co-founder, Cerebras. “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.”<br>Fast token generation is becoming increasingly important as AI moves into software development, autonomous agents, robotics, scientific discovery and other applications where response time directly shapes the user experience and the usefulness of the system. The joint solution brings together complementary architectures purpose-built for these demands.<br>AMD Helios provides the high-throughput prompt engine, rack-scale efficiency and deployment scale required to process large numbers of complex requests. Cerebras Wafer-Scale Engine technology provides the ultra-low-latency and decode performance needed to return tokens in real time. The result is a solution designed specifically for the ultra-low-latency segment of the inference market, with AMD Helios as the foundation for high-throughput and balanced inference workloads across the data center.<br>Cerebras plans to deploy AMD Helios systems in its data centers, with the joint solution expected to become available initially through Cerebras Cloud in the second half of 2026.<br>Supporting Resources<br>Follow AMD at Advancing AI 2026 (Press Kit)<br>Learn more about AMD Helios™ rackscale solution<br>Learn more about AMD Instinct™ Accelerators<br>Connect with AMD on LinkedIn<br>Follow AMD on X<br>Learn more about the Cerebras Wafer-Scale Engine<br>Connect with Cerebras on LinkedIn | Follow Cerebras on X<br>About...

cerebras inference ultra scale solution latency

Related Articles