Google Locks Gemini Into Silicon with Frozen v2 – Targeting 10x Tokens per Watt

Release date:2026-07-21 Number of clicks:186

Google is developing a server chip codenamed Frozen v2 that hardwires parts of its Gemini model architecture directly into silicon. The project aims to slash inference latency and power consumption, with internal projections showing 6 to 10 times more tokens processed per unit of power compared to Google's latest TPUs.

The chip is driven by a severe internal compute shortage that has reportedly forced Google Cloud to turn away external customers. Deployment is targeted for 2028.

Frozen v2 marks a departure from general-purpose AI accelerators like GPUs and TPUs. By etching Gemini's computation logic into hardware, it reduces real-time decision-making overhead and data movement. The trade-off is flexibility: future Gemini models will only run on the chip if they keep the same underlying architecture.

1784623884172524.png

This is not Google's first attempt. An earlier "Frozen" version led by Jeff Dean burned full model weights into silicon, but was shelved due to a prohibitively short hardware lifecycle. The v2 iteration hardcodes architecture only, while allowing weight updates.

Frozen v2 will complement rather than replace TPUs, and Google has no plans for mass production at TPU scale. The project is partly treated as a trial run, building engineering expertise for more specialized chips as AI model architectures stabilize.


From ICgoodFind: Hardcoding the model is the ultimate efficiency move – but only if the model doesn't change. Google is betting Gemini's architecture is stable enough to freeze in silicon. If they're right, it's a 10x power play.

Home
TELEPHONE CONSULTATION
Whatsapp
Semiconductor Technology