Key Takeaways
CAUTION: Google's Gemini 2 is ACTUALLY useful
- Google's Frozen v2 chip will embed Gemini model elements directly in hardware, potentially achieving 6-10x efficiency over current TPUs
- The chip targets deployment by 2028 and addresses an internal compute crunch that has forced Google Cloud to turn down customer deals
- Frozen v2 is a separate chip line from TPUs, not a replacement, signaling Google's multi-pronged approach to AI infrastructure
Google is developing a new server chip that hardwires elements of its Gemini AI model directly into silicon, aiming to serve AI inference 6 to 10 times more efficiently than its current custom chips. The project, internally called Frozen v2, targets deployment by 2028 and represents Google's response to a compute capacity crisis that has forced its Cloud division to decline customer deals.

The Information reported Monday that the Alphabet-owned company expects Frozen v2 to ease internal tensions over scarce AI computing resources. Engineers are still finalizing the design, particularly how much model information will be permanently embedded in the hardware. Alphabet shares rose 3% on the news.
Why bake the model into the chip?
Traditional AI chips load model weights from memory for each inference request. This constant data shuffling burns power and adds latency. By embedding Gemini's weights directly into the chip's circuitry, Google can skip that bottleneck entirely.
The reported efficiency gain of 6 to 10 times is measured in AI tokens served per unit of power. For a company running billions of daily Gemini queries across Search, Gmail, and Cloud APIs, even modest per-query savings compound into massive reductions in electricity bills and data center footprint.
The tradeoff is flexibility. A chip with hardwired model weights cannot run a different model without new silicon. Google appears willing to accept that constraint for its flagship Gemini workloads.
Frozen v2 vs. TPUs: separate tracks, not a replacement
Google's tensor processing units have been its primary custom AI accelerators since 2016. The company has deployed over a million TPU v5p chips across its data centers. Frozen v2 is not meant to replace them.
Instead, the Frozen project creates a parallel chip line optimized for a specific, high-volume task: serving Gemini inference at scale. TPUs remain the general-purpose training and inference workhorses. Think of Frozen chips as specialists and TPUs as generalists.
This dual-track strategy mirrors what hyperscalers do in other domains. AWS runs custom Graviton chips for general compute alongside Inferentia chips for ML inference. Google is now applying the same logic to its AI stack.
The compute crunch forcing Google's hand
The Information's report highlights internal friction at Google over AI computing capacity. Google Cloud has reportedly turned away outside customers because it lacks the hardware to serve them. That's a painful position for a company competing with AWS and Microsoft Azure for enterprise AI workloads.
Google plans to spend over $80 billion in capital expenditure in 2025, much of it on AI infrastructure. But building data centers and buying NVIDIA GPUs takes time and money. A chip that squeezes 6 to 10 times more work from the same power envelope offers a faster path to capacity than simply adding more machines.
The timing is notable. Bloomberg reported last week that Google delayed the launch of its latest Gemini model after it fell short of internal benchmarks, particularly in coding. The company is simultaneously trying to improve model quality and the infrastructure to serve it.
What 2028 deployment means for competition
Three years is a long runway in AI hardware. NVIDIA, AMD, and a wave of startups will ship multiple chip generations before Frozen v2 reaches production. Amazon's Trainium 3 and Microsoft's Maia chips will also mature.
Google's bet is that embedding model weights in silicon delivers efficiency gains that general-purpose accelerators cannot match. If that thesis holds, Frozen v2 could give Google a structural cost advantage in serving Gemini queries. Lower inference costs translate to either higher margins or lower API prices for Cloud customers.
The risk is that Gemini itself evolves. If Google releases a fundamentally different model architecture before 2028, Frozen v2's hardwired weights become a liability. The company appears confident enough in Gemini's direction to commit silicon to it.
Logicity's Take
Google's willingness to lock Gemini weights into hardware signals confidence in the model's longevity, but it also reveals desperation. Turning away Cloud customers due to compute shortages is not a strategy; it's a symptom. The 2028 timeline means Google needs bridge solutions. Expect more aggressive TPU v6 rollouts and possibly increased reliance on NVIDIA hardware in the interim. For enterprise buyers evaluating Google Cloud vs. AWS or Azure, the question is whether Google can stabilize capacity in the next 12 to 18 months, not 2028.
Frequently Asked Questions
What is Google's Frozen v2 chip?
Frozen v2 is a new server chip Google is developing that embeds elements of its Gemini AI model directly into hardware, targeting 6 to 10 times greater efficiency than current TPUs for AI inference.
When will Google's Frozen v2 chip be available?
Google plans to deploy Frozen v2 as soon as 2028, though engineers are still finalizing the chip's design and how much model information will be hardwired.
Will Frozen v2 replace Google's TPU chips?
No. Frozen v2 is a separate chip line meant to complement TPUs, not replace them. TPUs remain Google's general-purpose AI training and inference accelerators.
Why is Google building model-specific AI chips?
Google faces an internal AI compute shortage that has forced its Cloud division to decline customer deals. Model-specific chips could multiply efficiency and ease capacity constraints.
Context on the broader AI infrastructure investment wave
Need Help Implementing This?
If your team is evaluating cloud AI infrastructure or custom chip strategies, Logicity's consulting network can connect you with specialists in AI deployment architecture. Reach out via our contact page.
Source: Tech-Economic Times / ET
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in Trending Tech
AI Revolution: How Tech is Transforming the World, One Industry at a Time
From desalination plants in Iran to AI-powered manufacturing, the tech world is abuzz with innovation. Discover how AI is changing the game for small entrepreneurs and what it means for the future of industry. Explore the latest developments in cybersecurity, robotics, and more.

Revolutionizing AI: The Game-Changing Tech That's Making Agents Smarter
A new technology is set to revolutionize the way AI agents learn and adapt, enabling them to accumulate wisdom and apply it to new situations. This innovation has the potential to significantly boost the reliability of AI agents, especially in complex tasks. By converting raw agent trajectories into reusable guidelines, this tech is poised to transform the AI landscape.

The Dark Side of AI: How Bots Are Fueling a Monetized Abuse Ecosystem
A recent analysis of 2.8 million Telegram messages reveals a shocking truth: AI-powered bots are being used to create and sell non-consensual intimate images. These bots can turn ordinary photos into synthetic nude images, and the abuse is being monetized through affiliate programs and subscription-based archives. The researchers behind the study are calling for stricter regulations to combat this growing problem.

AI's Secret Sauce: How Journalism Became the Unlikely Ingredient
A recent study reveals that AI chatbots rely heavily on journalistic sources for their quotes, with one in four coming from news outlets. This shocking discovery has significant implications for the media industry and our understanding of AI's information gathering processes. As AI technology continues to evolve, it's essential to consider the role of journalism in shaping its responses.


