Key Takeaways
How Google Makes Custom Cloud Chips That Power Apple AI And Gemini
- Frozen v2 embeds Gemini's model architecture into silicon, potentially delivering 6-10x efficiency gains over current TPUs
- Unlike the original Frozen design that baked in weights, v2 embeds architecture only, allowing new weights to be loaded
- Google plans deployment in 2028, targeting internal AI compute capacity rather than external customers
Google is developing a custom server chip called Frozen v2 that hardcodes Gemini's model architecture directly into silicon. According to The Information, the chip could serve AI responses 6 to 10 times more efficiently than Google's current TPU chips, with deployment planned for 2028.
The approach marks a significant departure from general-purpose AI accelerators. Where TPUs can run many different models, Frozen v2 trades flexibility for raw efficiency by permanently embedding parts of Gemini's structural blueprint into the hardware itself.

Why freeze the architecture instead of the weights?
The name follows a familiar concept in machine learning. When training a model, engineers sometimes "freeze" certain parameters so they stop updating. Frozen v2 applies that logic to hardware: a portion of Gemini's structure gets locked into the chip, cutting compute steps and speeding up responses.
Jeff Dean, Google DeepMind's chief scientist, reportedly conceived the original Frozen design. His first version would have embedded the model weights directly into the chip. Weights are the specific numerical values that determine how a model responds to queries. Google scrapped that approach because the chip would only work with a single Gemini version and become obsolete too quickly.
Frozen v2 takes a more practical path. By embedding the architecture (the underlying blueprint) rather than the weights (the tuned parameters), the chip can accept new weight updates. Think of it like building a car factory optimized for a specific chassis design: you can still swap out the engine and interior, but the frame stays the same.
The Information reports that Google hasn't yet decided how much of the architecture will actually be hardcoded. That decision will likely depend on the tradeoff between efficiency gains and the risk that Gemini's architecture evolves in ways the chip can't accommodate.
Internal tool, not a product for customers
Google already commercializes its TPUs aggressively. The company leases them to Meta, offers them to external cloud customers, and pitches its "TPU@Premises" program as a direct Nvidia alternative. Internally, Google has reportedly set a goal of capturing 10% of Nvidia's annual revenue through these efforts.
Frozen v2 won't follow that path. Because the chip only works as long as Google sticks with the same Gemini architecture, it's impractical as a product for outside customers. Instead, Google views it as a way to ease internal AI compute constraints. The company expects smaller production volumes than its TPU line.
That doesn't mean Frozen v2 lacks commercial impact. If the chip delivers on its efficiency promises, Google could run Gemini at dramatically lower costs than competitors running equivalent models on general-purpose hardware.
The inference cost war
For AI companies, inference costs increasingly determine margins. Training a model is a one-time expense. Serving it to millions of users is ongoing. Every percentage point of efficiency gain compounds across billions of API calls.
A 6-10x efficiency improvement would be substantial. Google could use that headroom to undercut OpenAI and Anthropic on pricing, or pocket the savings as margin, or fund more expensive model capabilities at the same price point. Probably some combination of all three.
The strategy carries risk. Model architectures evolve. Google has already moved through multiple Gemini generations, and techniques like mixture-of-experts, speculative decoding, and sparse attention continue to reshape what "optimal" looks like. A chip designed around 2024's architecture might look suboptimal by 2028's standards.
Google's bet seems to be that Gemini's core architecture is stable enough to justify the investment. The smaller production volume suggests they're treating Frozen v2 as an experiment rather than a wholesale TPU replacement.
What this means for the broader AI chip market
Nvidia dominates AI training and inference because its GPUs are general-purpose. You can run any model architecture on an H100. That flexibility comes at an efficiency cost compared to hardware optimized for a specific workload.
The history of computing is full of examples where general-purpose hardware eventually loses to specialized silicon once a workload matures. ASICs replaced GPUs for cryptocurrency mining. Custom video encoding chips replaced software encoders in data centers. If Frozen v2 works, it could signal that AI inference is mature enough for similar specialization.
The catch: cryptocurrency mining algorithms are fixed, while AI architectures are not. Whether model-specific chips make sense depends entirely on how quickly architectures evolve relative to chip development cycles. A two-year chip development timeline works fine if your model architecture is stable for five years. It's a disaster if architectural breakthroughs arrive every 18 months.
Logicity's Take
For AI teams building on Gemini, this is worth tracking but not worth planning around yet. A 2028 deployment means any efficiency benefits are still years out, and Google hasn't confirmed whether those gains would trickle down to API pricing. The more interesting signal is strategic: Google is betting that Gemini's architecture is stable enough to justify custom silicon, which suggests confidence in their current technical direction. For teams choosing between model providers, the question is whether Google will translate hardware advantages into pricing advantages, or simply use them to fund more capable models at current prices.
Frequently Asked Questions
What is Google's Frozen v2 chip?
Frozen v2 is a custom server chip that embeds Gemini's model architecture directly into silicon. Unlike general-purpose TPUs, it's optimized specifically for running Gemini models, potentially delivering 6-10x efficiency gains.
When will Google deploy Frozen v2?
Google plans to deploy Frozen v2 starting in 2028. The company views it as a test run for specialized chips, with smaller production volumes than its TPU line.
Why didn't Google embed the model weights instead of the architecture?
The original Frozen design would have embedded weights directly, but that chip would only work with a single Gemini version and become obsolete when the model updated. Embedding architecture instead allows new weights to be loaded while still gaining efficiency.
Will Frozen v2 be available to external customers?
No. Because the chip only works with Gemini's specific architecture, Google plans to use it internally to reduce AI compute costs rather than selling or leasing it to outside customers.
How does Frozen v2 compare to Google's TPUs?
TPUs are general-purpose AI accelerators that can run many different models. Frozen v2 sacrifices that flexibility to achieve higher efficiency for Gemini specifically, trading versatility for performance.
Cost-efficiency tradeoffs in AI inference are central to both stories
Need Help Implementing This?
If you're evaluating AI infrastructure decisions or comparing model providers, Logicity's consulting team can help you map the cost and capability tradeoffs for your specific use case. Reach out at consulting@logicity.in.
Source: The Decoder / Matthias Bastian
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




