
Meta’s artificial-intelligence teams have been told to cut back their use of Google’s Gemini models, according to people familiar with the matter, after Google concluded it could no longer supply the computing capacity its biggest cloud rival was requesting. The restriction, first reported by the Financial Times and CNBC, is the most visible sign yet that compute allocation has become a strategic weapon among the handful of companies that both sell and consume AI infrastructure.
Google told Meta that demand for AI computing had outstripped what its data centers could deliver, the people said. Rather than simply expand capacity for a customer that is also a direct competitor, Google began steering Gemini access through an internal quota system that left Meta engineers with less headroom than they had planned for. Meta declined to comment. Google did not respond to requests for comment.
The move inverts the usual logic of the cloud business, where a vendor’s incentive is to sell every available chip to anyone willing to pay. Google still sells computing capacity, but it now also decides who may run Gemini, the model family that has become a default for a large slice of the industry’s generative AI work. Quota allocation, once an internal plumbing detail, is now a commercial and strategic instrument in its own right.
For Meta the constraint is doubly awkward. The company is one of the largest buyers of cloud capacity in the world, spending billions of dollars a year across Google Cloud, Microsoft Azure and Amazon Web Services to train and serve its Llama models. It also competes with Google directly in consumer AI, messaging and digital advertising. A supplier that is also a rival can raise the drawbridge at any moment, and this episode gives that abstract risk a concrete example.
Analysts said the episode points to a broader change in how the AI supply chain operates. The bottlenecks have moved from raw chips to the systems around them: memory, power, data-center space and, increasingly, the model layer itself. “If you are a cloud customer that competes with your provider, you are renting capacity from someone who can decide your access,” one technology analyst said. “That changes the negotiating table.”
The scarcity is real, according to people familiar with Google’s capacity planning. The company has been racing to bring data centers online, but demand from its own products, including Search, Gemini, Android and Workspace, has consumed capacity faster than it can be built. In that environment, allocation decisions become a form of rationing, and the criteria are no longer purely commercial. Internal teams have been told to treat every major Gemini deployment as a capacity commitment that must be planned months in advance.
Google executives have framed the policy in neutral terms, describing it as standard capacity management under unusual pressure. But the practical effect is that the world’s largest model developers are learning to treat compute as a diplomatic resource. Microsoft, Amazon and Google each control different layers of the stack, and each has shown where its loyalties lie when its own models are at stake. Enterprises that rent capacity from a hyperscaler while competing with that hyperscaler’s products now have a documented case study of the risk.
The episode also complicates the assumption that openness will keep the market competitive. Gemini is licensed to outside companies, and Meta used it for specific tasks, the people said, but the terms of that access now appear to include a judgment about the customer. Regulators have been watching the AI infrastructure market closely; whether compute allocation amounts to exclusionary conduct could be a question for antitrust agencies next, especially in Europe, where digital-market rules give regulators a direct line into such practices.
For now, Meta is absorbing the change quietly. The company has its own chips in development and has been working to reduce its dependence on external model providers, but its needs are so large that no single supplier can be abandoned overnight. The episode has already changed how Meta plans capacity, according to a person close to the company: contracts now include language about access guarantees, and internal teams have been told not to assume Google Cloud is a permanent home for Gemini workloads.
There is an irony in the timing. Meta has spent the past year arguing that open models will prevent any single company from controlling AI, and it has released its Llama models in open form as part of that argument. Google’s restriction shows that access to models is only half the question; the machines underneath matter just as much, and those are far from open. Meta has spent the past year arguing that open models will prevent any single company from controlling AI, releasing its Llama models openly as part of that argument. Google’s restriction shows that access to models is only half the question; the machines underneath matter just as much, and those are far from open.”
Google’s message to the industry, delivered in allocation tables rather than press releases, is direct: whoever controls the models and the machines that run them sets the terms. Compute has become a moat, and for the first time, the drawbridge has been raised in full view of the market.


