by Denkstrom
All storiesGoogle Frozen v2: Gemini baked directly into silicon

Google Frozen v2: Gemini baked directly into silicon

Google is developing Frozen v2, a server chip that permanently encodes parts of the Gemini AI model directly into silicon. Energy efficiency is expected to be six to ten times higher than current TPU chips. Planned deployment: 2028.

Moving data between processor and memory consumes most of the energy when an AI responds. Google wants to solve this problem radically: A new server chip called Frozen v2 will permanently embed parts of the Gemini model into hardware rather than reloading them from memory with each request. The Information reported this on July 21, citing engineers working on the project.

How Frozen v2 differs from a TPU

Google's current chips are Tensor Processing Units, or TPUs. They are generalists: flexible enough to run many different models, but precisely for this reason forced to compromise on efficiency. Each time the model answers a request, the chip loads the model parameters from RAM. This costs energy and time.

Frozen v2 takes a different approach. Instead of keeping Gemini model parameters in memory, parts of them are permanently embedded into the chip's circuits. The name comes from AI jargon: "frozen" parameters are parameters that no longer change during operation. With Frozen v2, this freezing is literal. Google would blur the line between software and hardware.

The result, according to insiders: six to ten times more tokens per watt than current TPUs. A token is the unit of measurement for AI outputs; one word equals approximately 1.3 tokens. More tokens per watt means: more AI responses for the same energy expenditure. Given the volume Google processes daily with Gemini, this efficiency gain translates directly to billions in operating costs.

Why Google must solve this problem now

Google faces a bottleneck. Demand for AI compute power is growing faster than available infrastructure. TPUs, which Google introduced in a new generation in April 2026, are its current answer. They are fast, but expensive and energy-intensive. Frozen v2 addresses a different bottleneck: not compute power, but energy efficiency during inference.

Inference is the operating mode in which AI responds to requests, as opposed to training, where the model learns. Training happens once. Inference happens billions of times daily. Google therefore sees Frozen v2 initially as a pilot for specialized silicon architectures while Gemini matures into a stable model. According to reports, Frozen v2 will be produced in significantly smaller quantities than TPUs.

What this means for Gemini users

Frozen v2 is not directly visible to users. No new user interface, no new model version. The chip works in data centers. What theoretically improves: faster response times and lower costs per AI query, both prerequisites for cheaper or faster services.

There is also a risk. A chip with an embedded Gemini model loses flexibility. Once Google introduces a new version of Gemini, older Frozen v2 chips won't match the new architecture. TPUs solve this problem simply through new software. How and whether Google handles this tradeoff in future chip generations remains unclear.

Apple, Meta, and Microsoft building in parallel

Frozen v2 is not a standalone phenomenon. Major tech companies have been building their own AI chips for years to reduce dependence on Nvidia. Apple is working under the codename ACDC on data center inference chips expected to enter mass production in the second half of 2026. Meta has announced four new generations of its MTIA chips to be deployed from 2026 to 2027. Meta's MTIA family increases HBM bandwidth by 4.5 times and compute capacity by 25 times from generation 300 to 500. Microsoft operates Maia 200 internally.

The difference with Frozen v2 lies in its radicality. Meta, Apple, and Microsoft are building more powerful general-purpose chips for their AI models. Google goes a step further and tests whether one can permanently bake a specific model into hardware. This is not a gradual efficiency gain, but a paradigm shift in chip design that only works if the model itself has stabilized.

Frozen v2 as a pilot: What must be proven by 2028

Google does not plan to deploy Frozen v2 before 2028. By then, the concept must prove more than laboratory efficiency. The central question is longevity: How often does Google update Gemini, and how well can a chip custom-built for today's architecture swap against a new model version? If Gemini evolves quickly, a chip tailored to today's architecture could become outdated within two years.

Google itself explicitly calls Frozen v2 a pilot, not the planned replacement for the TPU fleet. If the test shows that stable core areas of Gemini can be embedded efficiently, future generations could pursue this approach at larger scale. That would represent a paradigm shift in AI infrastructure: from flexible general-purpose chips to highly efficient, model-specific silicon components.