For the past two years, the defining question in AI was simple: Whose model is the smartest? That question is becoming less decisive. The model layer is commoditizing faster than expected, with capability gaps that once took a year to close now narrowing within a quarter—often through engineering advances outside the model itself. A new variable is therefore entering the AI decision matrix: How much does a token cost?
The token has evolved far beyond its technical definition as a sub-word unit.
- To an agent, a token is the fundamental unit of work.
- To a CFO, a token is an emerging operating expense.
- To the market, a token increasingly behaves like a tradable commodity, with different suppliers, prices, routing mechanisms and even emerging financial instruments.
Our central thesis is straightforward: the token economy is fundamentally an energy economy cloaked in software. As model capabilities converge, the ability to generate and distribute reliable tokens at lower and more predictable cost will become an increasingly important source of competitive advantage.
Agents have rewritten the economics of AI
Traditional chat interfaces created a natural limit on AI consumption: humans are slow. A user asks a question, waits for an answer, reads it and formulates the next prompt. Human interaction therefore acts as a throttle on token consumption.
Agents remove that throttle. A coding agent may plan a task, inspect files, call tools, execute commands, evaluate results, correct mistakes and repeat the process. Every iteration can expand the context passed back to the model. Industry estimates commonly put agentic workloads at 10–50x the token consumption of traditional single-turn interactions, depending on the task.
This changes the economics of AI. As agent adoption grows, token consumption becomes less a function of human usage and more a function of machine execution. The market is already responding: heavy usage has forced AI products such as Cursor and Claude Code to rethink fixed subscription models, introducing usage limits or additional charges when a small number of users consume disproportionately large amounts of compute.
The logical response is multi-model routing. Not every step in an agent workflow requires a frontier model. Complex planning can be assigned to a premium model, while retrieval, summarization, classification or routine tool execution can be handled by smaller, cheaper models. Research frameworks such as RouteLLM and FrugalGPT have demonstrated that substantial cost reductions are possible without proportionally sacrificing performance.
This is where the agent harness becomes strategically important. The harness—the layer responsible for planning, tool orchestration, context management, error recovery and integrations—can compensate for weaknesses in the underlying model.
Better engineering can therefore narrow the practical gap between expensive frontier models and cheaper alternatives.
The implication is significant: the intelligence premium of the model is shrinking, while the economics of inference are becoming more important.
Caching reveals the physical cost of a token
Routing is only one side of token optimization. Prompt caching attacks the other: redundant computation.
Agent workflows repeatedly send the same system prompts, tool definitions and persistent context. Instead of recomputing these tokens every time, providers can cache the underlying KV states and reuse them.
Major providers now offer substantial discounts for cached input, in some cases approaching 90%. The exact pricing differs by model and cache duration, but the economic principle is consistent: if computation has already been performed, subsequent reuse costs dramatically less.
This distinction matters because a 90% cache discount is not the same as a 90% cache hit rate. If 90% of tokens are cached, and cached tokens cost only 10% of the list price, the blended cost is roughly 19% of the original price: an effective saving of about 80%.
For enterprise financial planning, these distinctions matter. Token economics is not simply about list price; it is about effective unit cost under real workloads.
Caching also exposes something deeper. Why is a cached token cheaper? Because less computation is required. Less computation means fewer GPU cycles. GPU cycles consume electricity. In other words, the price of a token ultimately reflects physical infrastructure.
A token is crystallized electricity
At the infrastructure level, the price of a token incorporates GPUs and their depreciation, datacenter capacity, networking, cooling, and, above all, electricity.
The precise cost breakdown varies by architecture and accounting methodology, but the direction is clear: as AI inference scales, energy is becoming one of the most important constraints on the economics of intelligence.
The numbers are already material. A typical AI interaction may consume only a fraction of a watt-hour, but multiplied across hundreds of billions of interactions and amplified by agentic workflows involving repeated reasoning and tool calls, the aggregate requirement becomes enormous.
The International Energy Agency expects global datacenter electricity consumption to approach 950 TWh by 2030, nearly doubling from 2025 levels.
Geography consequently matters. Industrial electricity prices vary significantly across major AI markets, while China has expanded renewable generation at extraordinary speed and invested heavily in energy-abundant western regions through its “East Data, West Computing” strategy.
Software intelligence can scale globally within months, but cheap, reliable electricity cannot. Building generation capacity, transmission infrastructure, and datacenters requires years of capital investment. As model capabilities become increasingly substitutable, access to abundant and low-cost compute becomes a more durable advantage.
Tokens turn energy into a globally tradable service
Electricity itself is difficult to export across borders. Power grids remain fundamentally regional. Tokens solve this problem.
Inside a datacenter, electricity becomes compute. Compute becomes tokens. Tokens can then be transmitted instantly through APIs and fiber networks to users anywhere in the world. The token is therefore a form of digitized energy: physical energy transformed into a globally distributable service.
This helps explain the growing competitiveness of Chinese AI models. Industry estimates suggest that inference costs for leading Chinese models can be substantially below Western equivalents, although the exact differential varies by model, provider and workload.
The market response is increasingly visible. OpenRouter and other routing platforms allow developers to compare models on price, latency, and performance and to dynamically redirect traffic. Chinese models have gained significant share of routed traffic, while open-weight models such as Qwen and DeepSeek have become increasingly prominent in global developer ecosystems.
This does not mean Western frontier models are becoming irrelevant. Premium models continue to command pricing power for complex, high-stakes reasoning. Instead, the market is bifurcating: expensive frontier models serve the highest-value cognitive workloads, while cheaper models capture the much larger volume of routine inference.
The strategic question is therefore no longer simply who has the most capable model. It is increasingly ”Who can deliver the required level of intelligence at the lowest reliable unit cost?”
Open weights accelerate commoditization
Chinese model providers have also pursued a different commercial strategy: open weights.
DeepSeek, Qwen, GLM, Kimi and other Chinese model families have released downloadable weights to varying degrees, often under permissive licenses. By contrast, leading Western providers remain predominantly centered on closed API ecosystems.
The two approaches reflect different economic models. Closed models monetize scarce cognitive capabilities through proprietary access. Open-weight models use the model itself as an adoption mechanism, shifting competition downstream toward inference infrastructure, cloud services, developer tools, fine-tuning and enterprise integration.
DeepSeek-R1’s release in January 2025 marked an important inflection point. As high-performing model weights became increasingly accessible, the foundation model itself became less of a moat.
Once model weights approach zero marginal distribution cost, the competitive battleground moves to inference economics and the enterprise layer. This is the essence of commoditization: when intelligence becomes abundant, the scarce resources move elsewhere.
Token economics increasingly resembles a commodity market
The structure emerging around tokens already resembles a mature commodity market:
- Price differentiation: The same model can have dramatically different inference prices depending on the provider, region and infrastructure.
- Routing and arbitrage: Platforms such as OpenRouter allow users to dynamically select among models and providers based on price, latency and throughput.
- Tiered pricing: Providers increasingly differentiate prices according to priority, speed, batch processing and guaranteed capacity.
- Quotas and allocation: Usage limits and enterprise tiers effectively allocate scarce compute according to customer value and demand.
- Financialization: Emerging indices and compute futures point toward a future in which enterprises may manage compute exposure much like other variable input costs.
The result is a rapid compression of the commodity cycle. Routing, caching, model substitution and low-cost electricity are all solving the same underlying problem: How to minimize the delivered cost of useful intelligence. And in commodity markets, unit economics ultimately matter.
The Artefact perspective: Enterprises need to manage their token economy
For enterprises, token economics should therefore become a management discipline, not simply an infrastructure concern.
1. Treat tokens as a FinOps issue: AI cost needs to be tied to business outcomes. How many tokens does it take to resolve a support ticket? What is the token cost of an automated pull request? What is the blended AI cost per customer interaction? Without this attribution, enterprises cannot distinguish productive AI consumption from waste.
2. Make multi-model routing an architectural principle: Frontier models should not be the default for every task. Enterprises should dynamically match model capability to task complexity, supported by intelligent routing, caching and observability. The enduring value will increasingly sit outside the model: proprietary workflows, enterprise context, orchestration logic, domain expertise and trusted data.
3. Evaluate Chinese tokens pragmatically: The growing cost differential between Chinese and Western models deserves serious evaluation, particularly for high-volume, lower-sensitivity workloads.
For sensitive data, open-weight models deployed in private or sovereign environments can provide greater control over data residency and infrastructure. For standardized, high-frequency tasks, cost-efficient API endpoints may offer significant improvements in unit economics.
4. Build falling model costs into business models: Model prices will continue to decline. Sustainable margins will therefore depend less on access to raw intelligence and more on workflow integration, proprietary data, orchestration, domain expertise and customer trust.
5. Put energy on the AI strategy agenda: For technology companies, cloud providers, manufacturers and policymakers, AI infrastructure can no longer be separated from energy strategy. The next competitive advantage may come not from owning the smartest model, but from securing abundant, resilient and affordable compute.
Conclusion: The shift from smart to high-quality/low cost
The primary competitive axis of AI has shifted from “Which model is smartest?” to “Who can generate and deliver high-quality tokens at the lowest cost?”
Because a token is ultimately crystallized compute and electricity, the future map of AI dominance will not be drawn solely by elite research labs, but by the distribution of abundant, low-cost, and resilient energy infrastructure.

BLOG





