The Evolution of Google's TPU: 7th Generation Ironwood Ushers in the Era of Inference AI

Exploring the development history and evolution of Google's Tensor Processing Unit (TPU). From its inception in 2013 to the 7th generation Ironwood announced in 2025, the journey of specialized AI chips designed for inference.

The Evolution of Google's TPU: 7th Generation Ironwood Ushers in the Era of Inference AI

Google’s Tensor Processing Unit (TPU), which plays a crucial role in the company’s AI strategy, has undergone rapid evolution since development began in 2013. The 7th generation “Ironwood” announced in April 2025, with its inference-specific design, heralds the arrival of the “age of inference” where AI proactively generates insights.

TPU development began with challenges Google faced in its own data centers. Traditional CPUs and GPUs struggled to efficiently process neural network computations, and power efficiency was also a concern. Google developed TPU v1 in just 15 months—an exceptionally short timeframe—starting in late 2013, and began internal deployment in 20151.

History and Evolution of TPU Development

Early Breakthrough: TPU v1 (2015)

TPU v1 was designed primarily for inference tasks, manufactured on a 28nm process, operating at 700MHz with 40W power consumption. This chip achieved 15-30x performance improvement and 30-80x better power efficiency compared to contemporary CPUs and GPUs2. Google packaged this TPU as an external accelerator card that could be inserted into SATA hard disk slots, facilitating deployment in existing data centers.

Addressing Training: TPU v2/v3 (2017-2018)

The second generation TPU addressed memory bandwidth limitations by incorporating 16GB of High Bandwidth Memory (HBM), increasing bandwidth to 600GB/s. By supporting the bfloat16 format developed by Google Brain, it became capable of both training and inference for machine learning models. Performance reached 45 teraFLOPS3.

The third generation TPU was announced on May 8, 2018, with further performance improvements and enhanced scalability.

Current Generation TPU Specification Comparison

GenerationRelease YearPrimary UsePeak PerformanceMemory CapacityFeatures
TPU v12016Inference--INT8 operations, 40W power
TPU v22017Training/Inference45 TFLOPS16GB HBMbfloat16 support
TPU v32018Training/Inference90 TFLOPS32GB HBMLiquid cooling
TPU v42021Training/Inference275 TFLOPS32GB HBM2x+ performance vs v3
TPU v5e2023Inference-focused197 TFLOPS16GB HBM2Cost-efficiency focused
TPU v5p2023Training-focused459 TFLOPS95GB HBMLarge model support
TPU v6 (Trillium)2024Training/Inference925.9 TFLOPS32GB HBM4.7x performance vs v5e
TPU v7 (Ironwood)2025Inference-focusedFP8 support192GB HBM2x power efficiency vs v6

Latest Generation: Ironwood (TPU v7) Innovation

Ironwood, announced at Google Cloud Next 25 in April 2025, was designed as the first TPU specifically for inference. Key features include 2x power efficiency compared to Trillium and 192GB of HBM per chip (6x that of Trillium)4.

Ironwood’s most significant feature is being the first TPU to support FP8 calculations. In a full pod system, 9,216 chips are interconnected, achieving 21.26 exaflops at FP16 and 42.52 exaflops at FP8. This performance exceeds the world’s largest existing supercomputers5.

Edge TPU Deployment

In January 2019, Google announced the Edge TPU for edge computing. This chip can perform 4 trillion operations per second at 2W power consumption and is offered to developers under the Coral brand6. Edge TPU enables local AI inference execution without cloud connectivity.

Transition to the Age of Inference

Ironwood’s release marks a significant turning point in Google’s AI strategy. It aims to transition from traditional “responsive AI” to the “age of inference” where AI proactively retrieves and generates data, providing insights and interpretations. This change enables AI agents to not just provide data but collaboratively generate insights and answers.

Impact of TPU

Google’s TPU development has significantly impacted the industry in several ways:

  1. Cost Efficiency Improvement: Specialized design achieves substantial cost reduction compared to general-purpose processors
  2. Energy Efficiency: Contributing to reduced power consumption in data centers
  3. Service Enhancement: Enabling advanced AI features in services like Google Search, Google Photos, and Google Translate
  4. Cloud Service Strengthening: Providing TPUs to Google Cloud customers, promoting AI development democratization

Google’s TPU development is expected to continue playing a crucial role in the evolution and proliferation of AI technology.

Sources

  1. An in-depth look at Google’s first Tensor Processing Unit (TPU) - Google Cloud Official Blog (TPU v1 detailed specifications)
  2. Google’s First Tensor Processing Unit : Origins - The Chip Letter (TPU development history)
  3. Tensor Processing Unit - Wikipedia - Wikipedia (Overview of all TPU generations)
  4. Ironwood: The first Google TPU for the age of inference - Google Official Blog (Ironwood announcement)
  5. Stacking Up Google’s “Ironwood” TPU Pod To Other AI Supercomputers - The Next Platform (Ironwood performance analysis)
  6. What is a tensor processing unit (TPU)? - TechTarget (TPU and Edge TPU explanation)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →