Why Chasing More AI Tokens Is Costing Enterprises Real Money

The AI industry is obsessed with scale — bigger models, longer context windows, higher token counts. But evidence suggests that maximizing token consumption (tokenmaxxing) is one of the most expensive habits an enterprise can develop.
What Is Tokenmaxxing?
Tokenmaxxing is the practice of intentionally maximizing AI token consumption to demonstrate higher usage, with the assumption that more tokens equal more productivity. In practice, it does the opposite.
The 2026 BCG AI Radar Survey found that organizations plan to allocate an average of 1.7% of corporate revenue to AI this year. Yet 60% of executives report seeing minimal or no material financial return from those investments. Raw token counts measure computational effort — not accuracy or business value.
Key Takeaways
- Vanity metrics: Large context windows and oversized outputs inflate compute costs and latency without improving outcomes. Fewer than 1% of enterprise executives report a significant AI ROI of 20% or more.
- Cost-per-successful-task is the right metric: Instead of measuring tokens per conversation, measure what it costs to complete a task end to end.
- Right-size your models: Use small, specialized models for high-volume simple tasks (classification, routing, extraction) and reserve large frontier models for genuinely complex work. One misconfigured Amazon AI agent ran up $1.8 million in cloud costs from uncapped repetitive data-matching.
The Fix
Build guardrails from day one: forecast AI costs at full undiscounted prices, set hard token caps and budget alerts, and match model size to task complexity. The most competitive enterprise AI will be defined not by output volume but by the minimum compute required to deliver a result.
Read the full article on Built In
Stay in Rhythm
Subscribe for insights that resonate • from strategic leadership to AI-fueled growth. The kind of content that makes your work thrum.
More from Thrum
Additional pieces exploring adjacent ideas
