Google is working on a new AI chip designed to make Gemini more efficient

The "Frozen v2" chip is projected to deliver a significant leap in efficiency, potentially performing between six and ten times better than Google’s current AI chips. This efficiency is measured by the number of tokens generated per unit of power, a critical metric in the energy-intensive world of large language models (LLMs) and other generative AI applications. Such a substantial improvement could dramatically lower the operational costs associated with running Google’s sophisticated AI models, particularly for inference workloads, which represent the vast majority of AI computing once models are trained and deployed for real-world use.
The Intensifying AI Arms Race and the Drive for Efficiency
The past few years have witnessed an unprecedented acceleration in artificial intelligence development, particularly with the advent of generative AI. Large language models like Google’s Gemini, OpenAI’s GPT series, and Anthropic’s Claude require colossal computational resources not just for their initial training, but also for their ongoing operation (inference) as users interact with them. This computational demand translates directly into massive energy consumption and significant financial outlays.
Google, through its parent company Alphabet, has publicly committed to staggering investments in its AI strategy, with plans to spend between $180 billion and $190 billion. This capital expenditure is directed towards building out the necessary infrastructure, including data centers, specialized hardware, and research and development, to support its ambitious AI initiatives. However, such monumental spending has, at times, caused apprehension among investors, who naturally seek clear pathways to profitability and a robust return on investment. The development of a highly efficient custom chip like "Frozen v2" directly addresses these concerns by promising to optimize the cost-performance ratio of Google’s AI operations.
The economic imperative extends beyond mere cost reduction. The sheer scale of AI inference required for billions of user queries across Google’s diverse product ecosystem—from Search and Assistant to Google Cloud and Workspace—demands hardware that can execute these tasks with minimal latency and maximum throughput. Energy efficiency is not just a financial consideration but also an environmental one, as the carbon footprint of large-scale AI operations becomes an increasingly scrutinized aspect of technological advancement. A chip that can generate more tokens per watt directly contributes to a more sustainable AI future.
Google’s Legacy in Custom Silicon: A Chronology of Innovation
Google’s foray into custom silicon is not new; it dates back nearly a decade, establishing a deep-rooted commitment to vertical integration in its hardware and software stack. This strategy began with the introduction of the Tensor Processing Unit (TPU) in 2016, specifically designed to accelerate machine learning workloads.
- 2016: First-Generation TPU (Cloud TPU v1): Google unveiled its first-generation TPU, initially used internally for applications like AlphaGo and Street View. It was later made available to Google Cloud customers, marking a pivotal moment in custom AI hardware.
- 2017: Cloud TPU v2: This iteration brought significantly more power, supporting both training and inference tasks, and was crucial for the development of early Transformer models.
- 2018: Cloud TPU v3: Offering liquid cooling and even greater computational density, v3 further scaled Google’s AI capabilities.
- 2020: Cloud TPU v4: Introduced with improved efficiency and scalability, v4 became a cornerstone for training some of the largest AI models at the time.
- 2023: Cloud TPU v5e and v5p: The most recent public iterations, v5e focused on cost-efficiency for inference and smaller training tasks, while v5p pushed the boundaries for large-scale, high-performance training.
The "Frozen v2" project builds directly on this extensive lineage. While details on its specific architecture remain under wraps, its internal codename suggests a distinct evolution from prior "Frozen" iterations, presumably optimized for the unique demands of Gemini models. This continuous innovation in custom hardware underscores Google’s belief that tightly integrating hardware and software design yields superior performance and efficiency compared to relying solely on off-the-shelf solutions. As Google itself stated in response to TechCrunch, "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers… By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads." This statement, while not directly confirming "Frozen v2," strongly aligns with the company’s established philosophy and ongoing efforts.
"Frozen v2": Technical Aspirations and Operational Impact
The projected 6-10x efficiency gain for "Frozen v2" is a staggering figure that could redefine the economics of AI inference for Google. In practical terms, this means that for the same amount of power consumption, Google could process six to ten times more AI queries, generate more tokens of text, or run more complex models. This efficiency is particularly critical for inference, where models are deployed at scale to serve billions of users. Training a model is a one-time, albeit intensive, event, but inference is continuous and constitutes the vast majority of ongoing AI compute needs.
The unit of measurement – "tokens generated per unit of power" – highlights the focus on the practical output of AI models. A "token" can be a word, a part of a word, or even a character, and it is the fundamental unit of data processed by LLMs. Maximizing tokens per watt directly addresses the high operational costs and energy footprint associated with running these models. For instance, a single query to a large language model can consume several kilojoules of energy, equivalent to charging a smartphone several times over. Scaling this across billions of queries daily necessitates extreme efficiency at the chip level.
While specific architectural details of "Frozen v2" are proprietary, it is logical to infer that Google is leveraging its deep expertise in custom accelerators, specialized memory architectures, and efficient data flow within its data centers to achieve these gains. This could involve optimizations for sparse computations common in neural networks, dedicated hardware for specific AI operations, or novel approaches to memory access and bandwidth. The 2028 target release date provides Google ample time for iterative design, rigorous testing, and seamless integration into its existing infrastructure, ensuring the chip is finely tuned for the evolving demands of its Gemini AI models.
The Industry-Wide Push for Custom AI Silicon
Google is by no means alone in its pursuit of custom AI chips. The entire technology industry is undergoing a significant paradigm shift, moving away from a near-total reliance on general-purpose GPUs, primarily from Nvidia, towards bespoke silicon solutions. This trend is driven by several factors:
- Nvidia’s Dominance and Supply Chain Concerns: Nvidia has historically held an overwhelming market share (estimated at 80-90%) in the high-performance AI chip market. While its GPUs are incredibly powerful, this dominance has created a single point of failure in the supply chain and allowed Nvidia significant pricing power. Major tech companies, unwilling to be beholden to a single vendor for such critical infrastructure, are actively seeking alternatives.
- Tailored Performance: General-purpose GPUs are designed for a wide range of computational tasks. Custom chips, however, can be meticulously designed and optimized for specific AI workloads (e.g., inference for LLMs, specialized training for vision models), often leading to superior performance-per-watt and cost-efficiency for those particular tasks.
- Cost Management: As AI models grow in complexity and scale, the cost of acquiring and operating GPUs becomes astronomical. Custom chips offer the potential for significant long-term cost savings, especially for companies like Google, Meta, Amazon, and Microsoft, which operate AI at a hyper-scale.
- Competitive Advantage: Owning the full stack – from software algorithms to the underlying hardware – provides a significant competitive advantage, allowing for tighter integration, faster innovation cycles, and optimized performance that competitors relying on generic hardware cannot easily match.
Recent announcements from other major AI players highlight this accelerating trend:
- June 2026: OpenAI’s "Jalapeño": OpenAI, the developer of ChatGPT, announced its first custom inference processor, internally dubbed "Jalapeño," built in collaboration with Broadcom. This move signals OpenAI’s commitment to optimizing its own large-scale inference operations.
- July 2026: Anthropic and Samsung: It was reported that Anthropic, another leading AI research company behind the Claude LLM, was in discussions with Samsung for a new chipmaking partnership.
- Meta Platforms: Meta has been developing its own custom silicon, including the MTIA (Meta Training and Inference Accelerator) chips, to power its AI models and reduce reliance on external vendors.
- Amazon Web Services (AWS): AWS has successfully deployed its custom Inferentia and Trainium chips for its cloud customers, offering specialized AI acceleration.
- Microsoft: Microsoft has also unveiled its own custom AI chips, Maia for AI inference and Cobalt for general-purpose computing, for use in its Azure cloud infrastructure.
This concerted effort across the industry signifies a strategic shift, where silicon design is becoming as critical as software development for AI leadership.
Market Reaction and Investor Confidence
News of the more efficient "Frozen v2" chip appears to have resonated positively with investors, providing a timely boost to Alphabet’s stock ahead of its earnings report. Following The Information’s publication, Alphabet’s stock climbed approximately 3% on Monday morning. This immediate positive reaction underscores investor relief and confidence in Google’s long-term strategy.
Investors have previously expressed concerns about the massive capital expenditures required to fuel Google’s AI ambitions. While the promise of AI is immense, the initial outlay is equally substantial. Announcements like "Frozen v2" serve as tangible evidence that Google is not only investing heavily but also strategically to ensure these investments yield a strong return. Analysts view such vertical integration as a prudent financial move, enabling Google to control its total cost of ownership (TCO) for AI infrastructure, improve profit margins on AI-powered services, and maintain a competitive edge in a rapidly evolving market.
The market’s enthusiasm reflects a belief that increased efficiency translates directly into reduced operational expenses over time, effectively de-risking Google’s considerable AI investments. This strategic independence from external chip suppliers is also seen as a positive, mitigating potential supply chain vulnerabilities and pricing pressures.
Strategic Implications and the Future Landscape of AI Hardware
The development of "Frozen v2" carries profound strategic implications for Alphabet and the broader AI ecosystem.
- For Google:
- Cost Savings and Profitability: The most immediate impact will be substantial cost savings on running its vast AI inference workloads, directly boosting profitability for Google’s AI-powered products and services.
- Competitive Advantage: Superior hardware tailored for its Gemini models will give Google a distinct performance advantage, potentially enabling more complex, faster, and more capable AI applications.
- Strategic Independence: Further reduces reliance on third-party chipmakers, providing greater control over its supply chain, intellectual property, and product roadmap.
- Google Cloud Offerings: While initially for in-house use, Google’s history with TPUs suggests that future iterations of custom silicon could eventually be offered to Google Cloud customers, providing a unique and highly optimized platform for AI workloads.
- For the AI Industry:
- Accelerated Innovation: The race for custom silicon will drive even faster innovation in chip design, pushing the boundaries of what’s possible in AI computation.
- Diversification of Supply Chain: As more companies develop their own chips, the AI hardware market will become more diversified, potentially leading to more competitive pricing and robust supply chains.
- Sustainability: The emphasis on energy efficiency will contribute to a more environmentally conscious AI industry, reducing the overall carbon footprint of advanced computing.
- For Nvidia: While Nvidia remains the dominant force, the proliferation of custom chips from major tech giants poses a long-term challenge to its market share. Nvidia will need to continue innovating at a rapid pace and potentially pivot its strategy to offer more specialized, configurable, or services-based solutions to maintain its leadership.
In conclusion, Alphabet’s "Frozen v2" chip represents a significant strategic investment in its AI future. By designing highly efficient custom silicon, Google aims to not only optimize the performance and cost-effectiveness of its Gemini models but also to solidify its competitive position, enhance its operational independence, and demonstrate a clear path to profitability for its substantial AI expenditures. This development is emblematic of a broader industry shift, where the battle for AI supremacy is increasingly being fought at the fundamental level of hardware design.







