AI Inference chips will end the same as Crypto Mining ASICs

josefchen1 pts0 comments

Josef Chen on X: "https://t.co/7kneA2Ad11" / X<br>Post

Log inSign up

Post

Josef Chen

@josefchen

The AI Data-Centre Bust Will Look Like a Boom<br>When I was fourteen I bought Bitcoin miners with my own money, and watched newer chips kill them on a spreadsheet while Bitcoin was booming.<br>I am watching the same thing set up again. NVIDIA is arranging half a trillion dollars to finance GPUs. Stocks of the neoclouds, the specialist clouds that borrow to buy GPUs and rent them out, are ripping. Etched just hit a ten billion dollar valuation for an ASIC built to take GPUs' steadiest work, and OLIX raised $312 million at $3.3 billion eleven days behind it.<br>I hold the bullish case. Demand is enormous and still growing: about half of US adults now use AI chatbots and a quarter use them daily. Nobody halves the token budget, and a better model can create its own market. That is why I am bullish in a way I never was on mining, and why custom inference silicon looks like one of the better hardware bets of the decade: the volumes are finally large enough to pay for a chip that does one thing. Even so, that is a quarter of one rich country, and the demand story runs well ahead of the paying base, which is precisely the condition in which build-outs overshoot.<br>None of that saves the machine.<br>On 1 September 2014, Bitmain posted an announcement on the Bitcointalk forum: computing power from its Antminer S2 miners, sold through Hashnest at 0.0016 BTC per GH/s, maintenance fee deducted daily. I was on the other side of that page, buying not the coin but the machinery under it: Antminers, steel boxes built to perform one calculation, and Hashnest claims on somebody else's machines. Later I bought GPU rigs for Ethereum. Demand never killed any of them: a computer can work perfectly and still be a terrible asset.<br>That was my first data-centre asset class. The data centre was smaller.<br>Twelve years later, NVIDIA started talking like a bank. On 10 August 2026 it announced plans with six of the largest names in capital markets, Apollo and BlackRock and Blackstone among them, to mobilise more than $500 billion of outside capital. The agreements were not final. Jensen Huang then said NVIDIA might, project by project, support the value left in the machines for up to 25% of one financing opportunity. The heroic number matters less than the shape of the deal: the firm selling the chips is helping to organise the money, and residual value has become part of the sales pitch before the current build-out is finished.<br>The bust will not begin when people stop using AI. It will begin when the predictable inference jobs that were supposed to pay for a GPU fleet leave before the debt does.<br>Which work sits still<br>Serving a token is mostly a memory problem. The machine reads the model's weights and the running conversation out of memory for every token it emits, and much of the accelerator waits. Everything else here follows from that.<br>A chip built for one fixed model can lay those weights out on the die and stop paying for the trip. That is where the win comes from, and no general-purpose GPU can match it on work that repeats.<br>The catch is the same sentence read backwards. Weights on the die means the chip only serves that model. You cannot buy the win without the welding, so the question is not whether ASICs beat GPUs. It is which work sits still long enough to be welded to.<br>The work that has moved so far is boring. Meta says its MTIA chips serve recommendation models while unsupported models remain on GPUs. Amazon reported in April that it had landed more than 2.1 million AI chips in twelve months, more than half Trainium, alongside more than one million announced NVIDIA GPUs from 2026. It buys both, and the simplest explanation is that the jobs are different.<br>My bet through 2030 is narrower than "custom chips win". The largest platforms move stable, repeated, high-volume inference onto chips they control; GPUs keep the models that change, the software those chips do not support, demand spikes and customers who need broad compatibility. Move enough predictable queues and the GPU survives. The spreadsheet that paid for it does not.<br>The obvious objection is Google, and it is the best one anybody has put to me. Google has served its own steady inference on chips it designed for about a decade, through several generations of them, and NVIDIA's data-centre business compounded through the whole period with Google buying GPUs the entire way. If migration alone were enough, my bust would have arrived years ago. So the call is not that predictable work starts to move, because it has been moving for ten years. The call is that the fleet losing that work is now bought with borrowed money against a six-year book.<br>I am wrong by 2030 if custom chips stay too painful to program, if model changes eat the savings, or if old GPUs keep named workloads and steady cash yields through a major hardware transition. Any one of those kills the call, and more AI demand...

chips gpus work inference nvidia demand

Related Articles