by Vivek Gupta - 2 days ago - 3 min read
Google is reportedly developing a new server chip designed specifically to run its Gemini AI models more efficiently.
The processor, internally known as Frozen v2, could deliver six to ten times more Gemini tokens per unit of electricity than Google’s latest custom AI chips, according to The Information. The chip is still in early development and may not be deployed before 2028.
Google has not officially confirmed the project, its specifications or its release schedule.
Unlike Google’s broader Tensor Processing Units, Frozen v2 is reportedly being designed around Gemini’s model architecture.
This could allow the chip to handle common Gemini operations with less data movement and lower power consumption. In AI infrastructure, moving data between memory and processing units is a major source of energy use.
The approach could improve efficiency, but it also carries risk. AI models evolve quickly, while chip design and manufacturing can take several years. Google would need to ensure that the chip remains useful even if Gemini’s architecture changes before 2028.
| Area | Reported information |
|---|---|
| Internal name | Frozen v2 |
| Main workload | Gemini inference |
| Efficiency target | Six to ten times more tokens per watt |
| Possible deployment | 2028 |
| Status | Early development |
| Official confirmation | None |
The efficiency figure is a development target, not a verified result from production hardware.
Actual performance would depend on the Gemini model, memory configuration, numerical precision, batch size and software optimisation.
The project comes as Google faces growing demand for AI computing power.
Reuters reported that limited infrastructure capacity has affected Google Cloud and restricted how much computing access the company can provide to some customers.
This makes efficiency increasingly important. Producing more tokens from the same amount of electricity could allow Google to serve more Gemini requests without increasing data-centre capacity at the same rate.
Google already uses several generations of Tensor Processing Units for AI training and inference.
| Chip | Primary role |
|---|---|
| TPU 8t | Training large AI models |
| TPU 8i | Low-latency inference |
| Ironwood | Large-scale training and inference |
| Frozen v2 | Reported Gemini-specific inference |
Frozen v2 is expected to complement Google’s TPU family rather than replace it.
The company could continue using TPUs for broader AI workloads while relying on Frozen v2 for high-volume Gemini requests.
The chip would strengthen Google’s control over the full Gemini stack, from model development to cloud infrastructure and hardware.
It could also reduce Google’s dependence on third-party processors for some inference workloads. However, the project remains unconfirmed, and its reported efficiency gains may change before any commercial deployment.
For now, Frozen v2 is best viewed as an early indication of Google’s effort to design more specialised hardware for Gemini.