03_altman_microsoft_tokens.md

The question came from the audience on July 9, and Sam Altman paused before answering: with compute and memory costs climbing, how does OpenAI protect its margins? The chief executive’s reply was direct. Rising costs for computation and memory are “an unfavorable factor, no question,” he said, before adding what has become his standard reassurance: Microsoft will remain one of OpenAI’s largest customers.

The remark, made at a technology conference in San Francisco, was a rare concession from a company that has spent two years framing AI economics as a story of falling prices. Altman did not quantify the cost pressure, but the context was clear. Memory prices have climbed through 2026 as hyperscalers compete for high-bandwidth memory, the chips packed into AI accelerators, and compute demand continues to outrun supply.

Altman paired the warning with a counterweight: efficiency. He said OpenAI’s newest model has improved token efficiency by 54 percent on agentic coding tasks, meaning the model can complete roughly the same coding work while consuming about a third less compute than its predecessor. The number matters because agentic coding is among the most token-hungry workloads in the industry; a small improvement in tokens per task translates into large changes in the cost of serving enterprise customers.

The Microsoft relationship anchors OpenAI’s commercial model. Microsoft has invested more than $13 billion in OpenAI since 2019, holds a seat on its board and runs OpenAI models through its Azure cloud, where many enterprise customers access them. Altman’s insistence that Microsoft “will remain one of OpenAI’s largest customers” was aimed at investors who have speculated about a cooling of the partnership as OpenAI builds its own data centers and signs compute deals with other clouds.

People familiar with the matter said the two companies continue to negotiate how the relationship evolves as OpenAI grows its direct enterprise sales. Analysts said the direction of travel is clear: OpenAI will keep selling through Azure, but it will also sell directly to large corporations, and the balance between the two channels will shift as its own infrastructure comes online.

The efficiency figure also speaks to a broader industry trend. Token efficiency, the amount of useful work a model performs per unit of compute, has become the metric that chief financial officers track when they evaluate AI spending. Frontier labs are competing on it, and OpenAI’s 54 percent improvement on agentic coding is among the larger single-generation gains publicly claimed this year.

Altman’s appearance was itself a signal. OpenAI’s chief executive rarely takes the stage to discuss cost structure, and the decision to engage the question directly, rather than deflect it, suggests the company believes the numbers can carry the argument. His message to the room was layered: costs are rising, efficiency is rising faster, and the relationship with Microsoft remains the anchor of the model’s distribution. Each clause answered a question investors have been asking for months.

The cost picture is not entirely under OpenAI’s control. Memory is the most visible pressure point: high-bandwidth memory prices have roughly doubled over the past year, according to industry estimates, driven by the AI accelerator buildout, and the shortage has become a topic in earnings calls from Nvidia and the memory makers. Compute costs are stickier still, tied to data center construction, power contracts and the supply of advanced chips.

Altman’s framing, analysts said, is designed for a skeptical audience. By conceding the cost headwind publicly, he pre-empts questions about margins while pointing to efficiency as the offset. The 54 percent figure is the payload of the message: even if input costs rise, the cost per completed task falls, and that is the number enterprise buyers ultimately care about.

The efficiency gains also change the conversation with enterprise customers. OpenAI has been pushing businesses to move from experimental use to production workloads, and the cost per task is the number procurement teams actually sign against. A model that completes agentic coding work with roughly a third less compute gives sales teams a concrete answer to the question of how much the system will cost at scale. People who work with OpenAI’s enterprise unit said the 54 percent figure is already being used in customer conversations as the headline justification for expanding deployments.

Investors have reason to watch both numbers. OpenAI is the largest private consumer of AI compute in the industry, and its cost structure is a proxy for the entire model market. If token efficiency gains keep pace with hardware and memory price increases, the industry’s gross margin story holds. If not, the pressure will show up in pricing, which is why OpenAI’s largest customer relationship matters more than ever.

Microsoft’s own incentives align with that logic. Azure sells the models, and Microsoft’s profit on OpenAI workloads depends on usage growing faster than serving costs. A 54 percent efficiency gain helps both sides of that equation. Altman said as much in his closing remark: the way to outrun rising costs, he said, is to make the models do more with less. The market will get a fuller picture in the coming months, when the efficiency gains show up in enterprise usage data and, eventually, in the financial statements of the companies that sell the compute.

Related Posts

  • September 6, 2026
  • 10 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 11 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…