- calendar_today August 17, 2025
Google unveiled Ironwood, its 7th-generation Tensor Processing Unit, which is a custom chip designed to transform its AI capabilities. This new architecture represents a strategic leap that directly addresses the evolving needs of Google’s most advanced Gemini models. Ironwood has been engineered to master simulated reasoning tasks, which Google refers to as “thinking.”
The company stresses the mutually beneficial connection between its advanced AI models and specialized infrastructure. Ironwood demonstrates this philosophy by delivering rapid inference speeds and extending the context window capabilities of Google’s powerful AI models. Google believes that Ironwood represents its most advanced and scalable TPU to date, which sets the stage for future AI systems capable of autonomous data collection and output generation on behalf of users. Google defines its “agentic AI” vision through this proactive user-focused methodology, while Ironwood serves as the driving force behind this new inference era.
Performance Unleashed: Ironwood’s Impressive Specs
Ironwood exhibits substantial throughput improvements beyond earlier Google TPUs’ capabilities. The organization plans to deploy large clusters consisting of up to 9,216 liquid-cooled Ironwood chips operating together. The newly upgraded Inter-Chip Interconnect (ICI) facilitates uninterrupted communication between these substantial arrays, enabling high-bandwidth and low-latency data transfer throughout the whole system.
Google’s internal teams as well as cloud developers will be able to access this powerful processing capability. Ironwood will be available in two configurations: Ironwood will be offered in two configurations including a 256-chip server for basic needs and a comprehensive 9,216-chip cluster intended for extreme AI processing demands.
The sheer computational power of a full Ironwood pod is staggering: 42.5 Exaflops of inference computing. Google reports that every single Ironwood chip achieves 4,614 TFLOPs peak throughput, which marks significant progress from earlier iterations. Each Ironwood chip contains 192GB of memory, which represents a sixfold improvement compared to the Trillium TPU. Memory bandwidth received a 4.5x increase to achieve 7.2 Tbps.
Contextualizing the Power: Ironwood’s Place in the AI Landscape
Different measurement techniques make it difficult to compare AI chip performance directly. Google establishes FP8 precision as the performance standard for Ironwood. The company states Ironwood “pods” perform at 24 times the speed of the world’s leading supercomputers although caution is needed because many supercomputers lack native FP8 hardware support.
Google’s direct performance comparisons do not include their TPU v6, known as Trillium. The company claims Ironwood delivers double the performance per watt compared to v6. A Google spokesperson explained that Ironwood replaces the TPU v5p while Trillium succeeded the less capable TPU v5e. At its highest capability, Trillium reached about 918 TFLOPS when working at FP8 precision.
The Road Ahead: Ironwood and the Future of AI
Despite the complexities of benchmarking, the message is clear: Ironwood marks a major advancement in Google’s artificial intelligence infrastructure. Ironwood’s accelerated performance and improved efficiency expand upon the solid base that has supported swift development in models such as Gemini 2.5, which runs on earlier TPU versions.
Google expects Ironwood’s advanced inference performance and efficiency to lead to major artificial intelligence breakthroughs next year. Ironwood’s ability to supply the computational power required for complex models and autonomous capabilities makes it a central player in Google’s “age of inference” concept where AI becomes an essential component of our digital existence.





