Google is developing a new server chip that embeds parts of the Gemini model directly into the hardware. The chip, informally called “Frozen v2,” is intended to serve Gemini more efficiently and help Google address a pressing shortage of AI computing power. Rollout is scheduled for 2028.
This is according to The Information, citing sources. According to the report, the chip could be six to ten times more efficient than Google’s latest custom AI chips, measured by the number of AI tokens served per unit of power. Engineers are still working on the design and determining exactly how much model information will be embedded. The figure is therefore a projection, not a measured performance metric.
A complement to the TPUs, not a replacement
Frozen v2 is part of the broader “Frozen” project, which centers on a new line of proprietary chips alongside Google’s Tensor Processing Units (TPUs). It is explicitly not a replacement for them. In recent years, Google has invested heavily in its TPU family. For example, the company previously introduced the Ironwood TPU, an inference-focused chip that scales to nearly 10,000 units per pod and runs Gemini 2.5, among other models.
More recently, at Cloud Next, Google unveiled two new TPUs, in which training (8t) and inference (8i) are separated. Frozen v2 would thus be added on top of that as a highly specialized option for Gemini-specific inference.
Google is grappling with a shortage of AI computing power, which is fueling internal tensions and leading Google Cloud to turn down deals with external customers.