Token explosion on OpenRouter masks real AI adoption, not signals a bubble

Weekly token consumption on OpenRouter has surged from 0.5 trillion to 126.2 trillion since January 2025, a 25,000% increase. However, this growth is largely driven by token-hungry reasoning models and inefficient AI agents, not a proportional rise in actual usage. The metric can mislead, as models like OpenAI's GPT 5.6 Luna generate many tokens per prompt without necessarily having more users.
The 25,000% token surge on OpenRouter reflects a shift in model architecture, not just user growth. Reasoning models like OpenAI's GPT 5.6 Luna produce extensive internal "thinking" sequences before answering, so each prompt consumes far more tokens than earlier, simpler models. Meanwhile, inefficient agentic AI systems—which loop through multiple calls and tool use—compound this effect, making raw token counts a poor proxy for real-world adoption or revenue.
Revenue data offers a clearer picture: OpenAI's Astra leads in spending, while Chinese models (Kimi, GLM, DeepSeek) have grown tenfold in monthly spending during 2026, albeit from a small base. This suggests that while token volume is inflated, actual commercial traction is more modest and unevenly distributed across providers.
This metric distortion could mislead investors, developers, and policymakers into overestimating AI's real-world footprint, potentially fueling speculative investment or premature regulatory responses. If token counts are mistaken for genuine usage, resources may flow toward optimizing token efficiency rather than solving actual user problems. Conversely, recognizing the gap could encourage more honest benchmarking—helping businesses make calmer, evidence-based decisions about which models truly deliver value.