
On 28 June 2026, Google announced it would size down its support for Meta’s Gemini AI, putting a hard cap on the allocated cloud resources as the demand for LLM services explodes worldwide.
Event itself
Google caps Meta’s Gemini AI usage after users flocked to the new LLM, putting pressure on Google’s infrastructure. The move came after the AI cluster’s utilisation rate hit levels that threatened to choke the rest of Google’s own service portfolio.
Google’s letter to Meta, published on its press page, detailed that Gemini would be limited to a fixed number of requests per minute and a hard ceiling on GPU hours. The decision was justified as a measure to protect “critical services” such as Search, Ads and Cloud for other customers.
Key players
At the centre is Sundar Pichai, Google’s CEO, who led the decision and signed off the communique sent to Meta’s chief executive, Mark Zuckerberg. Meta’s own tech head, Elon Musk, has been publicly demanding more capacity for the company’s advertising arm, but the letter clarified that Google’s commitment to Meta had been “progressing responsibly”.
Other stakeholders include Google Cloud’s infrastructure team, Meta’s AI engineering squad, and third‑party developers who rely on Gemini’s APIs for AI‑driven applications. The headline-breaking clash has already spurred speculation that Meta might look to alternative cloud providers to host Gemini if the relationship dilutes.
What was said
In the announcement, Google clarified that the caps were a “technical necessity” to maintain latency targets for core services. The company also noted that it had calculated the impact on Meta’s AI workloads and deemed the reduction minimal relative to overall system capacity.
Meta, in turn, rebutted that the decision would “disrupt the AI ecosystem” and that Gemini “is drawing on a fraction of Google’s total capacity”. They pointed out that few other vendors match Gemini’s capabilities, calling the move a potentially strategic squeeze on competition.
Background and context
Gemini debuted in late 2024 as Meta’s flagship language model, promising advanced reasoning and reduced hallucinations compared to earlier iterations. Its launch coincided with a boom in generative AI services, forcing major cloud providers to re‑balance their server allocations.
Google’s partnership with Meta began in 2025 under a shared‑infrastructure agreement. However, the surge in overall AI workloads from summer 2025 onward outstripped predictions, leading to performance slowdowns that forced Google to re‑evaluate its commitments.
Consequences
For developers, the cap means slower response times and higher cost per inference if they rely on Google’s infrastructure for Gemini. This may push the market toward alternative LLMs, potentially reshuffling the competitive landscape.
For Meta, the restriction could delay feature rollouts dependent on Gemini’s insights, weakening its advertisement targeting engine. If the partnership frays, Meta may face higher operational costs deploying Gemini on third‑party clouds.
Regulators might now scrutinise the new limits, questioning whether Google is wielding its infrastructure as a competitive weapon against Meta, a move that could invite antitrust investigations.
Personal take
All this is a textbook case of a tech giant playing hardball to safeguard its own interests. Google’s decision feels a bit like a corporate toddler refusing to share a toy because the other kid keeps taking it away. Meta, meanwhile, is left wondering whether it should start brewing its own cloud instead of playing the long game of “let Google keep the speed while we enjoy the ease.”
It will be fascinating to watch whether this showdown forces a shift toward a more fragmented AI ecosystem, or whether Google will backtrack after backlash from developers and regulators. Either way, it reminds us that even in the age of AI, resource allocation still feels like a game of Monopoly.