Google Frozen chip: Gemini integrated into the silicon.
Most AI chips are designed for general use, where you load a model onto them for execution. Google, however, is exploring a more unconventional concept: a chip that effectively embodies the model, with the framework of Gemini integrated into the hardware itself.
This initiative, informally dubbed “Frozen v2,” has been reported by The Information and picked up by Reuters and Bloomberg Law. Following this news, Alphabet's shares increased by as much as 3.7%. Though Google hasn't officially confirmed this project and the chip will take years to develop, it represents a significant wager on the future direction of AI infrastructure.
Understanding ‘frozen’
Current chips operate by keeping the model in memory, transferring its data back and forth, which consumes both power and time.
In contrast, Frozen v2 aims to embed the architecture of Gemini’s neural network directly into the circuitry. This hardware would lock in the design of Google’s current AI model, allowing engineers to update the model by loading new weights while maintaining a fixed, or “frozen,” underlying structure. The extent of the model that will be permanently integrated is still reportedly under consideration.
The advantage lies in efficiency. The Information suggests that the chip could achieve 6 to 10 times the efficiency of Google's latest custom AI chips, based on tokens served per unit of power. Instead of replacing Google’s TPUs, it would represent a new category of silicon. The expected rollout could be as soon as 2028.
The significance
The timing for this is intentional. The Information indicates that Frozen v2 is partly a reaction to a severe AI capacity crunch at Google, which has led Google Cloud to decline some external customer requests and has created internal conflicts.
Efficiency is currently paramount. Operating AI models is exceptionally costly, and saving any wattage at the data center level equates to significant savings. A chip optimized for a single model can eliminate much of the inefficiency found in general-purpose chips.
Additionally, the chip would be fast. With a fixed design, it can respond with minimal delay, making it ideal for real-time applications like voice assistants, where lag is problematic.
There's also a strategic aspect. Google already develops its own TPUs to lessen its dependence on Nvidia. A chip specifically for Gemini would further this independence, contributing to Google's strategy of diversifying its chip suppliers.
Not unique to Google
This strategy isn’t exclusive to Google. A startup named Taalas is already marketing a similar approach, embedding a model’s weights and architecture into a chip they call Hardcore.
Taalas shares impressive figures, claiming its chip can handle up to 17,000 tokens per second, compared to around 150 per user on a top Nvidia GPU. Additionally, it claims to operate without needing costly high-bandwidth memory, which could alleviate some of the memory pressures affecting the industry.
This reflects a broader theory: if a model can be integrated into silicon, the trade-off between flexibility, speed, cost, and power becomes favorable. For a company catering to billions with a single model, this balance can be advantageous. It parallels initiatives aimed at miniaturizing models for mobile devices.
The potential drawbacks
However, the evident risk is inflexibility. AI is evolving rapidly, and a chip tailored to the current Gemini could become outdated by 2028. Although Google’s design allows for weight updates, the overall architecture remains predetermined.
Moreover, there’s the issue of confirmation; Google has not yet acknowledged this initiative. A spokesperson merely stated that the teams experiment with high-efficiency concepts, with the caveat that not all lab projects transition into production.
Thus, Frozen v2 should be seen as an indicator rather than a finalized product. The message is clear: the competition for custom AI silicon is shifting from accommodating various models to integrating a singular model with the hardware. If Google succeeds, its competitors will need to respond.
Other articles
Google Frozen chip: Gemini integrated into the silicon.
A report indicates that the Google Frozen chip will integrate Gemini directly into silicon, achieving up to 10 times the efficiency by 2028. This has implications for AI.
