Microsoft Pushes Its Own AI Models to Cut Costly Dependence

Microsoft is quietly shifting more of its artificial-intelligence workloads onto models it built itself, a move designed to cut its dependence on outside providers and rein in a cost line that has been eating into profits, according to a report from TechCrunch. The company is increasing the share of traffic handled by its in-house models across products like Copilot and its Azure AI services.

The reasoning is arithmetic. Microsoft’s AI business is growing quickly, but the bill for running it is growing faster. Each query that travels through a third-party model carries a per-token price, and the company has found that inference costs, the computing required to run models once they are trained, are substantial enough to erode margins even as revenue climbs. Building its own models lets Microsoft swap that external price tag for internal cost.

The company is not starting from scratch. Microsoft has been developing small language models for years, and it has fielded larger in-house systems designed for specific product surfaces. Those models have been gradually taking over routine tasks, with third-party frontier models reserved for the most demanding workloads, according to people familiar with the company’s plans.

The shift reflects a broader industry pattern. Hyperscalers that once treated frontier models as a commodity they would buy have concluded that owning the model stack is a strategic necessity. Google has built Gemini from the ground up, Amazon has fielded its Nova family, Meta has released its Llama models, and xAI has become a major player in months. Microsoft was the outlier, leaning on its partnership with OpenAI while rivals built in-house.

That partnership remains central. Microsoft has invested tens of billions in OpenAI, uses its models across Windows, Office, and GitHub, and has said the relationship is a pillar of its AI strategy. Executives have described the arrangement as complementary rather than competitive, but the in-house push suggests the company wants an option beyond its biggest supplier, a hedge that analysts said is becoming standard practice among the largest buyers of AI compute.

The economics of inference explain the urgency. Model providers have cut prices repeatedly over the past two years, yet the volume of usage has grown so quickly that total spend keeps rising. For a company running AI at Microsoft’s scale, shaving a fraction of a cent per query from hundreds of millions of daily queries produces savings that compound into billions over a year.

The company has also been building the other half of the cost equation: its own compute. Microsoft has committed to spending tens of billions of dollars a year on data centers and chips, and it has designed custom silicon to run AI workloads more efficiently. Owning the models and the hardware that runs them gives it the same vertical control that Google has long enjoyed.

Customers may notice little difference. Microsoft has said its in-house models are competitive for many enterprise tasks, and that it will route workloads to whichever model performs best for the job. For most users, the change is invisible; for Microsoft’s finance team, it shows up as margin.

There are limits to the strategy. Building frontier-class models is expensive and uncertain, and Microsoft’s internal efforts have not matched the most advanced external systems on every benchmark. Analysts said the company is likely to keep buying top-tier models for the hardest problems while using its own for the bulk of traffic, a portfolio approach that balances cost and capability.

The strategy has deep roots at Microsoft. The company has fielded its own language models for years, from the Turing models that powered early Bing features to the Phi family of small models designed to run efficiently on modest hardware. What has changed is the scale of deployment: in-house models are now handling a meaningful share of production traffic rather than serving as research projects.

The economics of the OpenAI relationship add another layer. Microsoft’s investment entitles it to a share of OpenAI’s profits and to exclusive cloud hosting for its models, but it also pays OpenAI for the compute its products consume. As OpenAI’s pricing has evolved, Microsoft’s cost per query has become a line item that its own models can shrink, a tension that analysts said will shape the relationship for years.

The shift does not mean Microsoft is walking away from OpenAI. The companies’ partnership agreement runs for years, Microsoft’s products depend on OpenAI’s frontier models for their most demanding features, and executives have said the relationship will continue. What it means is that Microsoft intends to be a customer with options, the same position every other large AI buyer is trying to reach.

The timing is notable. Microsoft reports earnings this month, and investors have been pressing management on how much AI spending will compress profits. The in-house model push is part of the answer: more output per dollar, less dependence on external pricing, and a path toward an AI business that earns its keep. Whether it is enough to satisfy the market will be clear when the numbers come out.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…