The bill for AI inference is getting harder to ignore. Investor Chamath Palihapitiya is warning that soaring token spend will translate into earnings pressure for companies, placing him among a growing chorus of investors and tech executives who argue the tokenmaxxing era is nearing its end.

The case for watching token costs

"Tokenmaxxing" describes the practice of feeding AI systems as many tokens as possible to extract better outputs. For a stretch, the logic was defensible: more context produced better results, and the cost could be absorbed as an early-stage bet. The concern now is that this spending has compounded to a scale where it lands visibly on the operating expense line.

Palihapitiya's warning connects product behavior to margin. AI capabilities have expanded alongside the costs to run them. When enterprise AI spend grows faster than the productivity it generates, it stops being an investment and starts being a drag.

The counterargument

The counterargument is not hard to find. AI adoption is still early, and companies investing heavily in token throughput may be buying efficiency gains that surface over time rather than immediately. Spend now, harvest later has been the operating logic of every major platform cycle, and it has sometimes been correct.

That argument is coherent. It is also the one that gets tested hardest when earnings season arrives and costs are visible while gains remain diffuse.

On balance

What Palihapitiya and the chorus around him are flagging is a question of timing and measurement, not a verdict on AI itself. The tokenmaxxing era assumed volume was the correct optimization target. What's changed is that investors are asking whether each token dollar is earning its keep. The line to watch: whether companies begin breaking out AI inference costs as a named item in financial disclosures. Once spend is visible, it gets managed.

Related reading