AMD Advancing AI 2026: Talking CDNA5 with AMD's Alan Smith

rbanffy2 pts0 comments

AMD Advancing AI 2026: Talking CDNA5 with AMD’s Alan Smith

Chips and Cheese

SubscribeSign in

AMD Advancing AI 2026: Talking CDNA5 with AMD’s Alan Smith

George Cozma<br>Jul 27, 2026

Share

Hello you fine Internet folks,<br>Today we are over AMD's Advancing AI 2026 event to chit chat on their newly revealed CDNA5 architecture with Alan Smith! AMD's Corporate Fellow and Chief Architect for Datacenter GPUs.<br>Hope y'all enjoy!

The transcript below has been edited for conciseness, readability, and clarifications.<br>George: So, what do you do here at AMD?<br>Alan: Yeah. So, I’m a Corporate Fellow of Graphics Architecture at AMD and responsible for Instinct GPU architecture, which we launched today in MI455 and Helios.<br>George: Starting off, all the way down into the grunt of the engine of MI455. Prior CDNA architectures were based on the GCN architecture, which dates all the way back to Tahiti from I believe 2012, if memory serves, but with CDNA 5, you have now rebased to RDNA. What were some of the considerations for that rebasing of the architecture?<br>Alan: Yeah, that's a great question. I think there were many, many considerations for making that transition, right? First of all, you know, we wanted to move onto a modern architecture. There were some of the things that we were carrying from the previous GCN architecture, as you mentioned, that we wanted to enhance, including the cache system and also just the execution engine itself, right? So in terms of work scheduling, sequencing of the waves, and even instruction issue and things like this. So there were many things that we wanted to improve.<br>And so, we knew we wanted to make a big change in the architecture. We wanted to get significant gains in efficiency, and we felt like the opportunity was there for us to bring together what we had done for RDNA, what we wanted to do with CDNA, in order to bring the AMD GPU roadmap more closely together, which gives us then the opportunity for additional optimization in the future that can benefit both.<br>And one of the reasons we think that’s really important is because AI is everywhere, as you saw in the keynote today. And so, even from gaming and neural rendering and upscaling and all of these things that are now leaning into AI for GPU, even for rendering on the RDNA side, and the things that we’re doing with AI for data center, there’s an opportunity to really bring that together and leverage them on both product lines. So that’s why we wanted to do that.<br>George: Speaking of two different product lines, I believe last week we published an article on what is effectively MI430. What are the differences in terms of design for HPC versus AI, what design considerations do you have to make at sort of the WGP level for those two different product lines?<br>Alan: Yeah, great question. I mean, I think not just the WGP level, but the entire SOC level, how do we think about HPC workloads versus AI workloads? They have a lot of similarities, right? They both thrive on high-bandwidth memory, they both need cluster interconnects, they both really enjoy having a high-performance CPU connected to them with a CPU memory system, cache coherency, and all these things. So all of that is already common, right?<br>And the only thing that's different is the type of numerics that you need for the codes that you're running. And so, the way that we looked at this is we could have included double-precision floating point in the GPU like we've done in the past, or, you know, we could do even better for both if we took advantage of our chiplet architecture and optimized two versions of the compute chiplet: one for traditional simulation with high-precision formats like double-precision floating point, where we really need all of the full IEEE compliance with 64-bit; and then the same thing for the AI workloads, leveraging as highest throughput that we can achieve for both vector operations and the tensor operations for AI.<br>So having said that, even within the workgroup processor, still many things are in common, right? We need similar bandwidth out of the register file, out of the VGPR. So we're able to leverage all of that, and also all the scheduling, etc. So all of the support hardware within the workgroup processor from the LDS, the caching, the front-end sequencers, all those things, and the SIMDs themselves for the vector ALUs are all shared. And all we really need to do is modulate the number of double-precision units that we implement within the SIMD.<br>George: Speaking of the SIMDs, previously, a feature or one of the fundamental building blocks of how GCN worked and prior CDNA architectures was that there were four SIMD16 units that were being issued a Wave64 instruction with one instruction per four cycles. With CDNA 5, you have now moved to four SIMD32 units that are being issued one Wave32 per cycle and you have said that you have deprecated Wave64 support. Why did you choose to deprecate Wave64 support considering that RDNA with a similar SIMD...

architecture things alan wanted george cdna

Related Articles