Google Frozen chip: Gemini integrated into the silicon.

Google Frozen chip: Gemini integrated into the silicon.

      Most AI chips are designed for general use, where you load a model onto them for execution. Google, however, is exploring a more unconventional concept: a chip that effectively embodies the model, with the framework of Gemini integrated into the hardware itself.

      This initiative, informally dubbed “Frozen v2,” has been reported by The Information and picked up by Reuters and Bloomberg Law. Following this news, Alphabet's shares increased by as much as 3.7%. Though Google hasn't officially confirmed this project and the chip will take years to develop, it represents a significant wager on the future direction of AI infrastructure.

      Understanding ‘frozen’

      Current chips operate by keeping the model in memory, transferring its data back and forth, which consumes both power and time.

      In contrast, Frozen v2 aims to embed the architecture of Gemini’s neural network directly into the circuitry. This hardware would lock in the design of Google’s current AI model, allowing engineers to update the model by loading new weights while maintaining a fixed, or “frozen,” underlying structure. The extent of the model that will be permanently integrated is still reportedly under consideration.

      The advantage lies in efficiency. The Information suggests that the chip could achieve 6 to 10 times the efficiency of Google's latest custom AI chips, based on tokens served per unit of power. Instead of replacing Google’s TPUs, it would represent a new category of silicon. The expected rollout could be as soon as 2028.

      The significance

      The timing for this is intentional. The Information indicates that Frozen v2 is partly a reaction to a severe AI capacity crunch at Google, which has led Google Cloud to decline some external customer requests and has created internal conflicts.

      Efficiency is currently paramount. Operating AI models is exceptionally costly, and saving any wattage at the data center level equates to significant savings. A chip optimized for a single model can eliminate much of the inefficiency found in general-purpose chips.

      Additionally, the chip would be fast. With a fixed design, it can respond with minimal delay, making it ideal for real-time applications like voice assistants, where lag is problematic.

      There's also a strategic aspect. Google already develops its own TPUs to lessen its dependence on Nvidia. A chip specifically for Gemini would further this independence, contributing to Google's strategy of diversifying its chip suppliers.

      Not unique to Google

      This strategy isn’t exclusive to Google. A startup named Taalas is already marketing a similar approach, embedding a model’s weights and architecture into a chip they call Hardcore.

      Taalas shares impressive figures, claiming its chip can handle up to 17,000 tokens per second, compared to around 150 per user on a top Nvidia GPU. Additionally, it claims to operate without needing costly high-bandwidth memory, which could alleviate some of the memory pressures affecting the industry.

      This reflects a broader theory: if a model can be integrated into silicon, the trade-off between flexibility, speed, cost, and power becomes favorable. For a company catering to billions with a single model, this balance can be advantageous. It parallels initiatives aimed at miniaturizing models for mobile devices.

      The potential drawbacks

      However, the evident risk is inflexibility. AI is evolving rapidly, and a chip tailored to the current Gemini could become outdated by 2028. Although Google’s design allows for weight updates, the overall architecture remains predetermined.

      Moreover, there’s the issue of confirmation; Google has not yet acknowledged this initiative. A spokesperson merely stated that the teams experiment with high-efficiency concepts, with the caveat that not all lab projects transition into production.

      Thus, Frozen v2 should be seen as an indicator rather than a finalized product. The message is clear: the competition for custom AI silicon is shifting from accommodating various models to integrating a singular model with the hardware. If Google succeeds, its competitors will need to respond.

Other articles

AI agent security: four attacks in July, one common vulnerability AI agent security: four attacks in July, one common vulnerability In a span of ten days, four research teams discovered vulnerabilities in AI agents through four different methods, ranging from Claude for Chrome to compromised memory. This highlights the security risks associated with AI agents. Z.AI constructed a large AI data center utilizing chips manufactured in China. Z.AI constructed a large AI data center utilizing chips manufactured in China. The Chinese laboratory Z.AI has finished constructing a 1GW data center that exclusively utilizes chips produced in China to train its GLM models, entirely without Nvidia components. This development indicates certain implications for the AI sector. This Mac application can transform any window into a CRT-inspired surreal experience or evoke the nostalgia of a Game Boy. This Mac application can transform any window into a CRT-inspired surreal experience or evoke the nostalgia of a Game Boy. Glaze 1.9 employs real-time GPU shaders to enhance every Mac window, video, and game, offering more than 50 different appearances for a one-time cost of $9.99. A new AI model aims to make self-driving cars think carefully before making a sudden turn. A new AI model aims to make self-driving cars think carefully before making a sudden turn. A team from Seoul National University developed an AI that evaluates the safety decisions of self-driving cars in real-time, and it recently received a noteworthy highlight paper recognition at CVPR. Samsung's latest OLED laptop screens have become significantly brighter and will also have an extended lifespan. Samsung's latest OLED laptop screens have become significantly brighter and will also have an extended lifespan. Samsung Display's latest tandem OLED panel achieves a brightness of 1,600 nits and has a lifespan that is double that of standard OLED. It is already featured in a Lenovo laptop. Samsung's latest OLED laptop displays have become significantly brighter and also have an extended lifespan. Samsung's latest OLED laptop displays have become significantly brighter and also have an extended lifespan. Samsung Display's latest tandem OLED panel reaches a brightness of 1,600 nits and has a lifespan that is double that of standard OLED, and it has already been incorporated into a Lenovo laptop.

Google Frozen chip: Gemini integrated into the silicon.

A report indicates that the Google Frozen chip will integrate Gemini directly into silicon, achieving up to 10 times the efficiency by 2028. This has implications for AI.