AI Customers Learn to Spend Less, and Model Makers Feel It

12_openai_anthropic_efficiency

The finance chief of a mid-sized software company began receiving a new kind of report this spring: a weekly breakdown of how much the company’s engineers were spending on artificial-intelligence tokens, broken down by team, by model and by prompt. The reports came from an internal tool the company built after its AI bill tripled in four months. Within two quarters, the company had cut its spending by a third without reducing the number of tasks its AI systems handled.

That story, told by executives and confirmed by industry analysts, is playing out across the technology economy. OpenAI and Anthropic, the two most prominent makers of frontier AI models, are confronting a problem the industry had not taken seriously: their customers are getting efficient. The shift from token maximization to token value maximization is rewriting the economics of AI software, and it has created what Cloud Wars, an industry analysis firm, called a chief-executive-level dilemma for OpenAI.

The mechanics are straightforward. Most AI companies charge by the token, the unit of text a model processes. Revenue grows when customers use more tokens. For the past two years, usage grew automatically, as companies rushed to deploy AI and models consumed more tokens per task. That growth is slowing as customers learn to control costs: they compress prompts, cache repeated requests, route easy tasks to cheaper models and reserve the most expensive models for the hardest problems.

The behavioral change is showing up in the data. Enterprise procurement teams now treat AI spending like any other line item, with usage dashboards, budgets and quarterly reviews. Startups that once used frontier models for everything now use them selectively. And the tools of efficiency, prompt optimizers, model routers and caching layers, have become a small industry of their own, all of them designed to reduce the number of tokens that reach OpenAI and Anthropic.

The consequences for the model makers are uncomfortable. If customers use fewer tokens per task, revenue per task falls, and growth must come from a rising number of tasks instead. That is not an impossible model; the number of AI tasks in the economy is still growing quickly. But it changes the unit economics, and it puts pressure on pricing, because a customer that has learned to economize is a customer that will compare prices across providers.

Cloud Wars argues the situation is the most underrated development in the AI industry this year. The analysis firm’s researchers note that public discussion has focused on compute costs, model quality and regulation, while the quiet efficiency movement among customers has received far less attention. Yet it is the efficiency movement that will determine whether the companies selling AI can turn their enormous infrastructure spending into durable profits.

The dilemma for OpenAI, and for Anthropic, is balancing two goals that pull in opposite directions. Revenue growth depends on usage, and usage growth depends on customers finding the models valuable enough to keep paying. But the more valuable the models become, the more customers invest in using them efficiently, and the less revenue each unit of value generates. The companies must decide whether to defend price per token or volume of tokens, and the choice shapes everything from product design to sales strategy.

There are offsets. Reasoning models, which think step by step, consume far more tokens than their predecessors, and their adoption has pushed usage back up even as customers economize. Agentic applications, in which software performs multi-step tasks autonomously, can multiply token consumption many times over. The question is whether these new sources of usage outpace the efficiency gains, and the answer so far is unclear.

The efficiency movement also changes the competitive structure. If customers are price-sensitive and usage-conscious, then cheaper models, smaller specialized models and open-source alternatives become more attractive. The economics that once favored the biggest, most capable models now favor the models that deliver the most value per token, a standard that is easier for smaller competitors to meet.

The efficiency wave is visible in the numbers that matter. Enterprise customers report that the average tokens consumed per task have fallen sharply over the past year, even as the volume of tasks has grown, and the model makers’ own disclosures show revenue per token declining. The trend is partly a function of better models, which need fewer tokens to do the same work, and partly of customers who have learned to stop wasting tokens.

The model makers have responded with pricing experiments of their own. OpenAI and Anthropic have introduced subscription tiers, volume discounts and agent-specific pricing, and both are pushing customers toward applications that consume tokens in bursts that are harder to optimize. The goal is the same on both sides of the market: find a structure in which usage, efficiency and revenue can grow at the same time.

For the finance chief who built the token dashboard, the calculation is simple: his company’s AI bill is now a managed cost, not a runaway one. For the industry selling AI, the dashboard is a symbol of something harder, a customer base that has stopped being impressed by capability alone and started asking what capability costs. The companies that answer that question best, analysts say, will be the ones that win the next phase of the AI market.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…