Artificial Intelligence

Google’s Frozen v2 Chip Targets Better Gemini Efficiency

by Vivek Gupta - 2 days ago - 3 min read

Google is reportedly developing a new server chip designed specifically to run its Gemini AI models more efficiently.

The processor, internally known as Frozen v2, could deliver six to ten times more Gemini tokens per unit of electricity than Google’s latest custom AI chips, according to The Information. The chip is still in early development and may not be deployed before 2028.

Google has not officially confirmed the project, its specifications or its release schedule.

Frozen v2 could be built around Gemini

Unlike Google’s broader Tensor Processing Units, Frozen v2 is reportedly being designed around Gemini’s model architecture.

This could allow the chip to handle common Gemini operations with less data movement and lower power consumption. In AI infrastructure, moving data between memory and processing units is a major source of energy use.

The approach could improve efficiency, but it also carries risk. AI models evolve quickly, while chip design and manufacturing can take several years. Google would need to ensure that the chip remains useful even if Gemini’s architecture changes before 2028.

Reported details

AreaReported information
Internal nameFrozen v2
Main workloadGemini inference
Efficiency targetSix to ten times more tokens per watt
Possible deployment2028
StatusEarly development
Official confirmationNone

The efficiency figure is a development target, not a verified result from production hardware.

Actual performance would depend on the Gemini model, memory configuration, numerical precision, batch size and software optimisation.

Google is under pressure to expand AI capacity

The project comes as Google faces growing demand for AI computing power.

Reuters reported that limited infrastructure capacity has affected Google Cloud and restricted how much computing access the company can provide to some customers.

This makes efficiency increasingly important. Producing more tokens from the same amount of electricity could allow Google to serve more Gemini requests without increasing data-centre capacity at the same rate.

How it fits with Google’s existing chips

Google already uses several generations of Tensor Processing Units for AI training and inference.

ChipPrimary role
TPU 8tTraining large AI models
TPU 8iLow-latency inference
IronwoodLarge-scale training and inference
Frozen v2Reported Gemini-specific inference

Frozen v2 is expected to complement Google’s TPU family rather than replace it.

The company could continue using TPUs for broader AI workloads while relying on Frozen v2 for high-volume Gemini requests.

The wider impact

The chip would strengthen Google’s control over the full Gemini stack, from model development to cloud infrastructure and hardware.

It could also reduce Google’s dependence on third-party processors for some inference workloads. However, the project remains unconfirmed, and its reported efficiency gains may change before any commercial deployment.

For now, Frozen v2 is best viewed as an early indication of Google’s effort to design more specialised hardware for Gemini.