An icon of an eye to tell to indicate you can view the content by clicking
Signal
Original article date: Aug 25, 2026

Why Chasing More AI Tokens Is Costing Enterprises Real Money

August 25, 2026
5 min read

The AI industry is obsessed with scale — bigger models, longer context windows, higher token counts. But evidence suggests that maximizing token consumption (tokenmaxxing) is one of the most expensive habits an enterprise can develop.

What Is Tokenmaxxing?

Tokenmaxxing is the practice of intentionally maximizing AI token consumption to demonstrate higher usage, with the assumption that more tokens equal more productivity. In practice, it does the opposite.

The 2026 BCG AI Radar Survey found that organizations plan to allocate an average of 1.7% of corporate revenue to AI this year. Yet 60% of executives report seeing minimal or no material financial return from those investments. Raw token counts measure computational effort — not accuracy or business value.

Key Takeaways

  • Vanity metrics: Large context windows and oversized outputs inflate compute costs and latency without improving outcomes. Fewer than 1% of enterprise executives report a significant AI ROI of 20% or more.
  • Cost-per-successful-task is the right metric: Instead of measuring tokens per conversation, measure what it costs to complete a task end to end.
  • Right-size your models: Use small, specialized models for high-volume simple tasks (classification, routing, extraction) and reserve large frontier models for genuinely complex work. One misconfigured Amazon AI agent ran up $1.8 million in cloud costs from uncapped repetitive data-matching.

The Fix

Build guardrails from day one: forecast AI costs at full undiscounted prices, set hard token caps and budget alerts, and match model size to task complexity. The most competitive enterprise AI will be defined not by output volume but by the minimum compute required to deliver a result.

Read the full article on Built In