The market is underpricing memory bandwidth

abhiphull1 pts0 comments

The Market Is Underpricing Memory Bandwidth

Menu

Home

Market Movers

Stock Analysis

Top 50 Stocks

Investment Advisor

Stock Screener

Portfolio Analysis

TradingView Indicators

Options Tools

Options Dashboard

Options Advisor

Options Calculator

Target Playground

Best Covered Calls

Strategy Scanner

Weekly Options

Weekly OI Snapshot

OTM OI Scanner

LEAP Smart Money

Scanners

Signal Strength

Penny Stocks

Options-Ready<br>NEW

Technical Signals

Gamma Exposure<br>NEW

Premium Tools

Market Maker Patterns

Sector Options Flow

AI Stock Picks

Earnings Domino

Options Net Delta

High OI Analysis

Portfolio Builder

Social Sentiment

Reddit Stocks

Twitter Analytics

Tradestie

Home

Market Movers

Stock Analysis

Top 50 Stocks

Investment Advisor

Stock Screener

Portfolio Analysis

TradingView Indicators

Options Tools

Options Dashboard

Options Advisor

Options Calculator

Target Playground

Best Covered Calls

Strategy Scanner

Weekly Options

Weekly OI Snapshot

Scanners

Signal Strength

Penny Stocks

Options-Ready<br>NEW

Technical Signals

Gamma Exposure

Premium Tools

Market Maker Patterns

Sector Options Flow

AI Stock Picks

Earnings Domino

Options Net Delta

High OI Analysis

Portfolio Builder

Social Sentiment

Reddit Stocks

Twitter Analytics

Compute grew 106x since Pascal. Bandwidth grew 11x. The research frontier has already surrendered to this math — the market hasn't finished pricing it.

Every AI accelerator sold today is a memory-bandwidth business wearing a compute costume. The FLOPs are the brochure; the bytes are the product. This piece makes that claim with three datasets we run ourselves: a 3,407-part chip inventory crawled from vendor spec sheets, a daily research radar that tracks new academic publications and community attention, and our fundamentals store. The three tell one story, and the last chart shows the market has only priced half of it.

The arithmetic: 106x vs 10.9x

Start with the chip inventory. We track peak dense tensor FLOPs (no sparsity — vendor "with sparsity" figures are marketing doubling and we strip them) and peak memory bandwidth for every datacenter part in the database. Index NVIDIA's line to the P100, the chip that started the deep learning datacenter era in 2016:

From P100 to B200, dense FP16-class compute grew 106x (21.2 TFLOPS → 2,250 TFLOPS). Memory bandwidth grew 10.9x (732 GB/s → 8 TB/s). Compute compounded roughly ten times faster than the memory system feeding it, for eight consecutive years.

The cleaner way to see it is bytes-per-FLOP: how many bytes of memory bandwidth the chip gives you for every FLOP of compute. If this number falls, the chip is more bandwidth-starved — more of its silicon sits idle waiting for data.

The numbers, at dense FP16, from our inventory:

Chip<br>Year<br>Bandwidth<br>Dense FP16<br>Bytes/FLOP

P100<br>2016<br>732 GB/s<br>21.2 TF<br>0.0345

V100<br>2017<br>900 GB/s<br>125 TF<br>0.0072

A100 80GB<br>2020<br>2,039 GB/s<br>312 TF<br>0.0065

H100 SXM<br>2022<br>3,350 GB/s<br>989.5 TF<br>0.0034

H200 SXM<br>2024<br>4,800 GB/s<br>989.5 TF<br>0.0049

B200<br>2024<br>8,000 GB/s<br>2,250 TF<br>0.0036

A 10x decline in bytes-per-FLOP from P100 to B200. And that flatters the modern chips, because nobody serves models at FP16 anymore. At the precision people actually run inference — INT8 on Ampere, FP8 on Hopper, FP4 on Blackwell — the effective ops-per-byte collapse is far steeper: P100 at 0.0345 bytes per op down to B200-at-FP4 at 0.00089 , a 39x decline. Every FP4 op on a B200 gets one-fortieth the memory bandwidth a P100 op got.

AMD's Instinct line traces the identical slope: MI100 at 0.0067, MI300X at 0.0042, MI355X at 0.0033. Two vendors, two architectures, one wall.

The tell: vendors are shipping bandwidth refreshes

Look at what the H200 actually is. Same silicon as the H100 SXM — identical 989.5 TF dense FP16 in our inventory — with bandwidth raised from 3,350 to 4,800 GB/s. NVIDIA's biggest product refresh of 2024 added zero FLOPs and 43% more bandwidth , purchased entirely with more and faster HBM. AMD did exactly the same thing with the MI325X: identical 1,307 TF compute as the MI300X, bandwidth pushed from 5,427 to 6,144 GB/s.

When both vendors independently decide the way to sell more chips is to bolt on more memory and touch nothing else, they are telling you what the binding constraint is. You don't need our thesis; you can read theirs off the spec sheets.

And the forward roadmap escalates it: AMD's MI455X (July 2026 in our inventory) jumps to 23.9 TB/s of HBM4 bandwidth — a 2.9x single-generation leap, the largest in the dataset. At FP16 that actually restores bytes-per-FLOP to 0.0048, back to 2020 levels. Bandwidth is finally being bought back — and it is being bought with staggering quantities of stacked DRAM. That purchase order lands on the memory supply chain.

The research frontier already pivoted

If bandwidth is the wall, the smartest people in the field should be visibly climbing it. Our research radar — a daily scrape of new academic publication counts, trending-paper...

bandwidth options memory market bytes stock

Related Articles