Tokenminning Is the New Tokenmaxxing

kg3631 pts0 comments

Tokenminning is the new tokenmaxxing - Kabalan Gaspard

Kabalan Gaspard

SubscribeSign in

Tokenminning is the new tokenmaxxing

Kabalan Gaspard<br>Jul 21, 2026

Share

Towards the end of 2025, Boris Cherny (creator of Claude Code) shared a phrase that became the default architectural framework for anyone building with AI: “Don’t build for the model of today. Build for the model 6 months from now”.<br>Shortly afterwards, a parallel trend emerged: don’t just use the best available model as soon as possible, use it for everything, as much as possible. This philosophy became known, with a nod to Gen Z slang, as tokenmaxxing.<br>Put together, these two trends - which we’ll group under the “tokenmaxxing” umbrella for convenience - make sense. If we assume that LLMs will continue to improve (which they have and very likely will), and that tokens will continue to get cheaper just like almost any other technology has in the past 100 years, then tokenmaxxing is the self-evident choice. Build your product for where you think Iliad10 or GPT-10.9 will be, and assume that cost isn’t an issue compared to all the value you’ll be getting.<br>Late 2025 - early 2026: the rise of tokenmaxxing

At Tesserae, we bet against that trend in 2025. We assumed that, given LLMs fundamentally just predict the next word, there was a long way to go for them to get good at two crucial tasks involved in creating McKinsey-grade slides, which have little to do with writing or coding: 1) spatial reasoning, and 2) drawing creative business visualisations. We therefore built our own model and logic for those two tasks, and let the LLM do what it was good at in early 2025: determining the user intent, structuring a presentation, and populating the illustrations Tesserae drew with summarised text content. The bet seemed to pay off: the slides this Tesserae + mainstream lab hybrid model could produce were clearly better than what ChatGPT/Claude could produce out of the box.

Enterprise boardroom-style slide comparison late 2025<br>But then, in early 2026, just as tokenmaxxing was becoming a thing, the LLMs won. First, the big labs invested directly in making their models better at creating corporate-style PowerPoint slides (Claude for PowerPoint, released early February 2026, was a turning point). Second, LLMs got better at two fundamental skills crucial in building high-quality slides: drawing HTML on a canvas (which took time to translate into drawing a clear business slide, but the fundamentals were there), and translating grid-based layouts (which most consulting-style slides are based on) into code.<br>On the business side, around the same time (Q1 2026), the priority in most boardrooms was shifting to demonstrating value from the 2025 AI pilots they had launched. Tokenmaxxing provided a turnkey solution to that problem at just the right time - simply instruct the whole company to use the best available models as much as possible, and ride the wave of AI’s exponential pace of improvement as it translates into $$$ for the company. Social media - especially LinkedIn and X - were reflections of this, with Linkedinfluencers and developers alike rushing to brag about how they had burned hundreds of millions of tokens in a week as soon as new models were announced.<br>Things weren’t looking good for our approach. We still had a clear cost advantage - our token-cost-per-slide was (and still is) 10x less than Sonnet 4.6 and 25x less than Opus 4.6 - but really, who cares? Claude is subscription-based anyway, and that cost was never really felt by the users.

We were the midwits in early 2026<br>Q2 2026: the music slows down

Around April/May, several alarm bells started ringing. Uber confirmed it had burned through its entire 2026 AI budget in 4 months through tokenmaxxing. A KPMG survey reported that this wasn’t an isolated incident, and that a significant number of corporate executives are reeling from sticker shock over usage-based AI pricing schemes introduced by newer frontier models. In May, Anthropic announced that paying subscribers would face a separate monthly credit meter for agent tools, billed at full API rates.<br>This reverse trend trickled down from the boardroom. YCombinator’s Summer 2026 batch had significantly more startups building products in AI infrastrucure than the previous batch, many of which are helping AI teams reduce inference costs. While LinkedIn and X still have no shortage of tokenmaxxers, there seems to be a growing trend of posts admitting how expensive/infeasible/slow it is to throw everything at the latest model, and/or giving useful tips on how to get nearly the same level of performance with far fewer tokens or with less expensive models.<br>It turns out that, while the initial assumptions of tokenmaxxers were technically correct - LLMs are continuing to get better and, like any technology, cheaper on average - the price of the newest frontier models continues to increase. Again, this isn’t anything revolutionary, even in the world of...

tokenmaxxing model models llms early slides

Related Articles