Enjoy the Bubble While It Lasts

romshark2 pts0 comments

: the external Shoelace<br>stylesheets and module scripts below block first paint until fetched, and<br>the body is held hidden until Shoelace upgrades — so without this the<br>WebView2/browser default white flashes for the whole load. A literal colour<br>(not a CSS var, which theme.css defines later) paints the root immediately. -->

Enjoy the Bubble While It Lasts: The Secret Behind Your AI Bill · Kernel Pryanic

so it applies before first paint; a fallback timeout guarantees the<br>body can never stay hidden if an upgrade never resolves. -->

Originally published on dev.to, where it was pulled from search and the newest-posts feed within minutes, with no explanation. Support reinstated it after I reached out, also without saying what triggered it. A re-post was de-listed just as fast. This article is the reason this blog exists.

Somebody else is paying for your AI subscription. Not most of it. Almost all of it.

That's the open secret of this whole era, and it's not on the pricing page. The arithmetic is inconvenient enough that nobody selling you a plan is in a hurry to walk you through it.

SemiAnalysis did the obvious experiment: buy every plan OpenAI and Anthropic sell, then spend a month actually hammering them - long agentic runs, code generation that doesn't stop. A fully maxed-out ChatGPT Pro subscription burns through $14,000 a month of tokens at API prices. Claude Max lands near $8,000 . You're paying $200. Serving costs are lower than list, so the real hole is nearer $3,500 a month - which is still seventeen times what you handed over.

Plan price vs. maximum possible spend at API pricing. Source: SemiAnalysis.

The utilization math is worse than the headline. OpenAI's margin on Plus and Pro goes underwater past ~11.4% of the available limit. On the top plan, it hits zero at around 5.7%. Everything above that is loss. The business model, right now, is that most subscribers barely use what they bought - and the few of us who do are the cost center.

We are the cost center, and that's the good news

If you're a software engineer using AI seriously, you are exactly the user these plans lose money on. Agentic coding tasks consume roughly 1,000x the tokens of a standard chat query, with frontier models burning over 100,000 output tokens on a single task. Every time you point an agent at a repo and let it run, you're spending someone's capital.

That someone is a pension fund, a sovereign wealth fund, a hyperscaler's balance sheet, a chip vendor financing its own customers. Which is to say: ultimately a fairly small number of very rich people and the institutions that move their money. Enormous piles of capital betting that habits formed today convert into pricing power tomorrow. The bet may be correct. It may be not. Either way, the capital has already been spent, and the receipts are landing in your terminal.

This is not a scam and it's not new. It's the same playbook that gave us $5 cross-town rides and two-day free shipping and cloud compute below cost. The difference is the multiple. Uber subsidized maybe 50% of a ride. This is 40-70x.

The part of the bill that is yours

Except we're not getting away completely clean, and this half doesn't make it into the launch keynote.

Go price a RAM kit. A 32GB DDR5 kit that ran $60-80 in early 2025 now asks $180-250. That's not a supply hiccup - Samsung, SK Hynix and Micron make over 95% of the world's DRAM, and they're pointing their wafers at high-margin HBM for AI accelerators. Datacenters are on track to absorb something like 70% of high-end memory output this year, up from 20-30% a few years ago, and the squeeze runs through at least 2027.

So there is a tax, just not on the invoice with our name on it. The giants underwrite our inference, and we underwrite their memory supply chain by paying triple for a component we've bought without thinking for twenty years. A few hundred dollars every few years against thousands a month of frontier compute is still an excellent deal - for us.

Worth noting who else is in that "us", though. The markup is on the memory, not on the AI plan, so it lands on anyone buying a laptop, a phone, a console, a prebuilt. Most of them will never open an agent. The subsidy flows to people who use AI, but the memory tax is spread across everyone who buys a computer.

But it's a tell. When memory - the one input you'd need to run models yourself - is what's being priced out of reach, the cheap path and the independent path are quietly diverging.

Bubble, but which kind

"We're in a bubble" is now the safest opinion in tech, which usually means it's about to stop being useful. The interesting question isn't whether - it's what happens next, and there are roughly three shapes:

It bursts. Capital dries up, the loss-leader plans die, and inference gets priced at what it costs. Uncomfortable, but not the end - the models don't unlearn anything.

It deflates. Prices creep up, limits tighten quietly, the generous tiers get renamed and repriced....

plan memory bubble month capital without

Related Articles