Google’s Tensor Processing Unit (TPU), which plays a crucial role in the company’s AI strategy, has undergone rapid evolution since development began in 2013. The 7th generation “Ironwood” announced in April 2025, with its inference-specific design, heralds the arrival of the “age of inference” where AI proactively generates insights.
TPU development began with challenges Google faced in its own data centers. Traditional CPUs and GPUs struggled to efficiently process neural network computations, and power efficiency was also a concern. Google developed TPU v1 in just 15 months—an exceptionally short timeframe—starting in late 2013, and began internal deployment in 20151.
History and Evolution of TPU Development
Early Breakthrough: TPU v1 (2015)
TPU v1 was designed primarily for inference tasks, manufactured on a 28nm process, operating at 700MHz with 40W power consumption. This chip achieved 15-30x performance improvement and 30-80x better power efficiency compared to contemporary CPUs and GPUs2. Google packaged this TPU as an external accelerator card that could be inserted into SATA hard disk slots, facilitating deployment in existing data centers.
Addressing Training: TPU v2/v3 (2017-2018)
The second generation TPU addressed memory bandwidth limitations by incorporating 16GB of High Bandwidth Memory (HBM), increasing bandwidth to 600GB/s. By supporting the bfloat16 format developed by Google Brain, it became capable of both training and inference for machine learning models. Performance reached 45 teraFLOPS3.
The third generation TPU was announced on May 8, 2018, with further performance improvements and enhanced scalability.
Current Generation TPU Specification Comparison
| Generation | Release Year | Primary Use | Peak Performance | Memory Capacity | Features |
|---|---|---|---|---|---|
| TPU v1 | 2016 | Inference | - | - | INT8 operations, 40W power |
| TPU v2 | 2017 | Training/Inference | 45 TFLOPS | 16GB HBM | bfloat16 support |
| TPU v3 | 2018 | Training/Inference | 90 TFLOPS | 32GB HBM | Liquid cooling |
| TPU v4 | 2021 | Training/Inference | 275 TFLOPS | 32GB HBM | 2x+ performance vs v3 |
| TPU v5e | 2023 | Inference-focused | 197 TFLOPS | 16GB HBM2 | Cost-efficiency focused |
| TPU v5p | 2023 | Training-focused | 459 TFLOPS | 95GB HBM | Large model support |
| TPU v6 (Trillium) | 2024 | Training/Inference | 925.9 TFLOPS | 32GB HBM | 4.7x performance vs v5e |
| TPU v7 (Ironwood) | 2025 | Inference-focused | FP8 support | 192GB HBM | 2x power efficiency vs v6 |
Latest Generation: Ironwood (TPU v7) Innovation
Ironwood, announced at Google Cloud Next 25 in April 2025, was designed as the first TPU specifically for inference. Key features include 2x power efficiency compared to Trillium and 192GB of HBM per chip (6x that of Trillium)4.
Ironwood’s most significant feature is being the first TPU to support FP8 calculations. In a full pod system, 9,216 chips are interconnected, achieving 21.26 exaflops at FP16 and 42.52 exaflops at FP8. This performance exceeds the world’s largest existing supercomputers5.
Edge TPU Deployment
In January 2019, Google announced the Edge TPU for edge computing. This chip can perform 4 trillion operations per second at 2W power consumption and is offered to developers under the Coral brand6. Edge TPU enables local AI inference execution without cloud connectivity.
Transition to the Age of Inference
Ironwood’s release marks a significant turning point in Google’s AI strategy. It aims to transition from traditional “responsive AI” to the “age of inference” where AI proactively retrieves and generates data, providing insights and interpretations. This change enables AI agents to not just provide data but collaboratively generate insights and answers.
Impact of TPU
Google’s TPU development has significantly impacted the industry in several ways:
- Cost Efficiency Improvement: Specialized design achieves substantial cost reduction compared to general-purpose processors
- Energy Efficiency: Contributing to reduced power consumption in data centers
- Service Enhancement: Enabling advanced AI features in services like Google Search, Google Photos, and Google Translate
- Cloud Service Strengthening: Providing TPUs to Google Cloud customers, promoting AI development democratization
Google’s TPU development is expected to continue playing a crucial role in the evolution and proliferation of AI technology.
Sources
- An in-depth look at Google’s first Tensor Processing Unit (TPU) - Google Cloud Official Blog (TPU v1 detailed specifications)
- Google’s First Tensor Processing Unit : Origins - The Chip Letter (TPU development history)
- Tensor Processing Unit - Wikipedia - Wikipedia (Overview of all TPU generations)
- Ironwood: The first Google TPU for the age of inference - Google Official Blog (Ironwood announcement)
- Stacking Up Google’s “Ironwood” TPU Pod To Other AI Supercomputers - The Next Platform (Ironwood performance analysis)
- What is a tensor processing unit (TPU)? - TechTarget (TPU and Edge TPU explanation)