Silicon Valley has a new sport: burning as many AI tokens as possible. They even gave it a name — *tokenmaxxing*. The question is whether the trillion-dollar infrastructure buildout riding on this demand is real, or whether we're measuring activity instead of value.
CNBC's Deirdre Bosa laid out the stakes this week:
"Engineers, they are competing to consume the most AI measured by tokens. It is almost like a sport at this point. Jensen Huang, CEO of Nvidia, he said that he'd be alarmed if a top engineer was not burning 250K a year in AI compute. Shopify told me that they use it as a performance signal. And Meta employees, reportedly they blew through an estimated 900 million tokens in a month." — [Deirdre Bosa, 0:00](https://www.youtube.com/watch?v=2OHMstRVqdE&t=0)
Ramp CEO Eric Glyman, launching a new product to track enterprise AI spend, offered the sharpest data point:
"Across Ramp data, token and AI spend has grown by 13 times over the past year, 50% a quarter. And what is very clear is no one knows how to budget for this." — [Eric Glyman, 4:05](https://www.youtube.com/watch?v=2OHMstRVqdE&t=245)
Thirteen times in a year. That's not organic adoption — that's a firehose.
The core problem is familiar to anyone who's studied measurement: when a metric becomes a target, it ceases to be a good metric. Glyman named it directly:
"There was a lot of discussion even on X a few days ago about incentives at Meta. And I think people are pointing out this idea of Goodhart's law which says a measure of performance that becomes a goal ceases to become a good measure of performance. Once you incentivize use as many tokens, you will see engineers go and count all of the numbers of prime or all the digits of pi and use these tokens and it goes on and on." — [Eric Glyman, 7:11](https://www.youtube.com/watch?v=2OHMstRVqdE&t=431)
Bosa drew the historical parallel herself: Amazon used to grade call center reps on call speed, so reps started hanging up on customers. The metric improved. The service cratered. Same pattern, different domain.
The two biggest AI labs are taking opposite approaches — and the divergence tells you everything about what they think is happening:
"OpenAI is making AI cheaper, easier to use, so more people consume it. It needs the usage numbers to justify spending. Anthropic, meanwhile, putting limits on how much and making people pay for it, maybe because it wants to know the demand it is seeing is real." — [Deirdre Bosa, 1:01](https://www.youtube.com/watch?v=2OHMstRVqdE&t=61)
Investor Dan Niles was blunter:
"OpenAI is cutting prices to try to get customers to them. Meanwhile, Anthropic is actually raising prices and cutting people off because they have so much demand that they are trying to keep people from using too much of it. And to me, that shows a company that is feeling the pressure between Anthropic and Google." — [Dan Niles, 25:38](https://www.youtube.com/watch?v=2OHMstRVqdE&t=1538)
One lab is subsidizing consumption. The other is rationing it. One of these strategies will look very smart in hindsight. The other will look like Amazon's call center metric.
Niles pointed to a structural reason the spend explosion might be partly real — agents burn dramatically more compute than chat:
"If you look at OpenRouter, in the two months prior to the end of January, token growth was up about 20% or so. In the two months after agentic AI got formalized, that token growth has grown about 130%." — [Dan Niles, 28:43](https://www.youtube.com/watch?v=2OHMstRVqdE&t=1723)
Agents don't ask one question and get one answer. They plan, execute, retry, and loop. The token multiplier from agentic workflows is real. But real usage and inflated usage can coexist — the question is the ratio.
Glyman highlighted a signal that gets less attention:
"The frontier models have gone to I think over 20% of the share of tokens used to 4%. I think that is a harbinger of what is going to come. It is very interesting seeing this development coming out of China. It is a culture that is very attuned to efficiency." — [Eric Glyman, 15:22](https://www.youtube.com/watch?v=2OHMstRVqdE&t=922)
Frontier model share of tokens dropping from 20% to 4% while total spend triples — that's the open-source and distillation story playing out in real time. You can track the effect on [benchmark.space/rankings](https://benchmark.space/rankings), where sub-30B models like [Qwen 3.5 9B](https://benchmark.space/model/qwen-3.5-9b) and [Gemma 4 31B](https://benchmark.space/model/gemma-4-31b) increasingly match frontier performance at a fraction of the cost.
The tokenmaxxing narrative isn't wrong — engineers really are burning tokens for sport, and companies really are using AI spend as a vanity metric. But underneath the noise, agentic workflows are creating genuine, structural demand growth. The hard part — and the trillion-dollar question — is separating the two. Ramp's new product is trying to do exactly that. If it works, we'll finally know how much of the AI boom is signal and how much is just counting prime numbers.