The AI Cost Rebellion Is About Unit Economics

voxelperfect1 pts0 comments

The AI Cost Rebellion Is Really About Unit Economics | Kostas Karolemeas

Alex Karp asked the right question in the most provocative way.

If frontier AI companies create so much value, why do they charge for tokens instead of sharing in the outcome?

The question is commercially interested. Palantir sells an alternative vision of enterprise AI, built around controlled deployment, proprietary data, and operational workflows. But the question landed because many enterprise customers were already asking a less theatrical version of it:

What are we actually paying for?

Handelsblatt reports growing customer resistance to the pricing of OpenAI and Anthropic, alongside frustration about control and results. Its follow-up on cost management makes the operational problem explicit: AI spend is rising quickly, but few companies know precisely what they are paying for.

Vanessa Cann, who contributed to the reporting, describes the transition well. During the pilot phase, consumption was too small to dominate the business case. Once companies scaled useful systems—especially agents, where one user request can trigger many model calls—the hidden cost structure became visible. The new question is no longer whether a pilot works. It is what value a specific workflow creates when operated repeatedly.

This is not an AI rejection cycle.

It is the beginning of an AI unit-economics cycle.

The Token Is a Real Meter and the Wrong Management Unit

A token is not imaginary. It measures part of the work a model performs, gives engineers a way to understand consumption, and gives providers a way to invoice variable usage.

But a token has almost no meaning to a business owner.

A claims executive does not want tokens. She wants a correctly processed claim. A support leader wants a resolved case. An engineering manager wants a safe change merged. A compliance officer wants an investigated exception with an auditable decision.

Tokens sit several layers below those outcomes.

They also fail as a simple comparison unit. One token from a small classifier, one from a frontier reasoning model, and one produced inside a long agent loop are not economically equivalent. Provider price lists now distinguish input, output, cached input, long context, service tiers, tools, and batch processing. The public catalogs from OpenAI, Anthropic, and DeepSeek show how wide the price and product range has become.

Those options are useful for optimization. They do not answer whether the workflow earns its place.

The FinOps Foundation's work on token economics makes the distinction clearly. Token cost is only one layer in a wider stack that includes infrastructure, data, networking, engineering, observability, evaluation, governance, and labor. It also notes that a retrieval pipeline with reasoning and multiple tool calls can consume one or two orders of magnitude more tokens than a direct request to a smaller model.

The provider needs a production meter.

The enterprise needs an economic ledger.

Confusing the two is the root of the rebellion.

Scale Exposes What the Pilot Hid

Pilot economics are forgiving.

The user group is small. The workflow is supervised. Failures are treated as learning. Integration and governance labor are often funded centrally. Consumption may sit inside a trial, a generous subscription, or a budget no one has yet challenged.

Production removes those protections.

Every successful deployment increases volume. Agents add planning, retrieval, tool calls, retries, verification, and memory. More users create more edge cases. Higher autonomy requires better evaluation, monitoring, access control, incident response, and human escalation. A system that looked cheap per interaction can become expensive per completed process.

The evidence is still emerging, but the warning signs are strong. In a May 2026 survey with 75 qualified enterprise respondents, McKinsey found that spend rose nearly fourfold as organizations moved from isolated use cases toward enterprise adoption, while 93 percent reported exceeding their AI budgets. The same analysis cites research finding up to 30-fold variation in token use when an agent executes the same task.

That is a small survey, not a universal market estimate. But it describes a structural problem: agentic demand is nonlinear, while most budgets assume a stable relationship between users, requests, and cost.

This is now a mainstream operating concern. The State of FinOps 2026 report says 98 percent of FinOps teams manage AI spend, up from 31 percent two years earlier, and names AI cost management as the leading skill gap.

The market has moved from Can we build it? to Can we operate it economically?

Karp Is Right About Value, but Outcome Pricing Is Not Magic

Karp's implied alternative is attractive: if a vendor claims to create business value, let it charge for the value instead of the tokens.

Parts of the market are already moving that way.

Intercom now prices its Fin agent by outcomes....

cost token from unit economics value

Related Articles