Google Frozen chip: Gemini embedded in the silicon.

Google Frozen chip: Gemini embedded in the silicon.

      Most AI chips are designed for general purposes. You upload a model onto them, and they execute it. Google is reportedly exploring a more unconventional approach: a chip that integrates the model directly into the hardware itself, based on Gemini’s design.

      This initiative, informally dubbed “Frozen v2,” was reported by The Information and subsequently covered by Reuters and Bloomberg Law. Following the news, Alphabet's stock rose by as much as 3.7%. Google has not confirmed this project, and the chip is still several years from completion. However, the concept represents a significant gamble on the future direction of AI infrastructure.

      Understanding ‘frozen’

      Current chips maintain the model in memory and transfer its data back and forth, which incurs costs in terms of power and time. In contrast, Frozen v2 would incorporate Gemini’s neural network architecture directly into the circuitry. This hardware would conform to Google’s existing AI design. While engineers could still modify the model by loading new weights, the fundamental structure would remain unchanged, or “frozen.” Reportedly, discussions are ongoing regarding how much of the model will be integrated into the hardware.

      The advantage here is efficiency. According to The Information, this chip could achieve 6 to 10 times the efficiency of Google’s latest custom AI chips, measured by tokens processed per unit of power. This would represent a new category of silicon, distinct from Google’s TPUs, rather than a substitute. Deployment is expected as soon as 2028.

      Importance of the development

      The timing of this venture is significant. The Information notes that Frozen v2 is partly a reaction to a capacity crunch within Google’s AI capabilities. The situation is critical enough that Google Cloud has had to decline some external clients, leading to internal conflicts.

      Efficiency is currently paramount. Operating AI models is extremely costly, and every watt conserved at the data center level equates to savings. A chip specifically designed for one model can eliminate much of the overhead associated with a general-purpose chip.

      Speed is another benefit. With a fixed design, the chip can respond with minimal delay. As one expert pointed out, this is advantageous for real-time applications like voice assistants, where latency is a major concern.

      There’s also a strategic dimension. Google already creates its own TPUs to reduce dependence on Nvidia. A chip tailored to Gemini would further enhance that self-sufficiency, reinforcing an initiative that has seen Google diversify its chip suppliers.

      Google is not the only player

      This approach is not exclusive to Google. A startup named Taalas is marketing a similar concept, embedding a model’s weights and architecture directly onto a chip it calls Hardcore.

      The claimed performance is impressive. Taalas asserts that its chip delivers up to 17,000 tokens per second, compared to about 150 tokens per user on a leading Nvidia GPU. It also claims that it does not require expensive high-bandwidth memory, which could alleviate the memory limitations affecting the industry.

      This represents a broader gamble. If a model can be embedded in silicon, one sacrifices flexibility for gains in speed, cost, and power. For a company serving a single model to billions, this trade-off can be advantageous. It mirrors efforts to compress models for mobile devices.

      The risks involved

      The main risk is inflexibility. AI evolves rapidly, and a chip built around the current Gemini could become outdated by 2028. Google’s design aims to mitigate this by allowing updates to weights, but the architecture itself remains fixed.

      Additionally, there’s the matter of confirmation. Google has yet to acknowledge the initiative. A spokesperson stated only that the teams are exploring high-efficiency methods and that not all laboratory projects advance to production.

      Therefore, Frozen v2 should be viewed as an indication rather than an impending product. The message is clear: the competition for custom AI silicon is shifting from the capacity to run any model to integrating a specific model into the hardware. If Google succeeds, its competitors will need to react.

Other articles

AMD's Helios incorporates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72. AMD's Helios incorporates 72 GPUs and 31 terabytes of HBM4 within a single rack. This is AMD's response to Nvidia's NVL72. The AMD Helios contains 72 MI455X GPUs, 31TB of HBM4, and delivers 2.9 exaflops of inference computing capabilities within a single rack. Engineering samples are expected to ship in the second half of 2026, with mass production set to begin in the second quarter of 2027. A mini lab the size of AirPods could assist in detecting disease markers within minutes. A mini lab the size of AirPods could assist in detecting disease markers within minutes. The VPodDuo, which is roughly the size of an AirPods case, links to a smartphone and can analyze molecular tests for viruses, bacteria, and cancer-related biomarkers in approximately 10 minutes. Singles are utilizing ChatGPT and Claude to craft their dating messages, and the apps have no intention of putting a stop to it. Singles are utilizing ChatGPT and Claude to craft their dating messages, and the apps have no intention of putting a stop to it. 26% of adults in the US have sought assistance from AI for dating purposes. Match, Hinge, and Bumble have stated that they do not filter for messages generated by AI. This trend is referred to as chatfishing. Samsung's latest OLED laptop screens have become significantly brighter and will also have an extended lifespan. Samsung's latest OLED laptop screens have become significantly brighter and will also have an extended lifespan. Samsung Display's latest tandem OLED panel achieves a brightness of 1,600 nits and has a lifespan that is double that of standard OLED. It is already featured in a Lenovo laptop. OneDrive will prevent screenshots of sensitive documents, but this restriction applies only when using Microsoft's browser. Beginning in August, OneDrive and SharePoint will prevent screen captures of PDFs that have sensitivity labels when using Edge. This feature will not be available for Chrome, Firefox, or Safari. Microsoft aims to prevent screenshot leaks, focusing on one Edge tab at a time. Microsoft aims to prevent screenshot leaks, focusing on one Edge tab at a time. Microsoft is addressing a security vulnerability that permitted sensitive PDFs to be intercepted within web browsers. However, this protection currently necessitates the use of Edge and an enterprise account.

Google Frozen chip: Gemini embedded in the silicon.

A report indicates that the Google Frozen chip will integrate Gemini directly into silicon, potentially achieving up to 10 times efficiency by 2028. This development could have significant implications for AI.