Cerebras WSE-3: The Wafer-Scale Chip Built to Speed Up AI Inference
Cerebras Systems' WSE-3 chip features 4 trillion transistors and 44GB of on-chip SRAM, achieving over 20x faster inference than traditional GPUs. An exploration of its technological innovation and founder's vision.
In an era where AI inference speed can determine business success or failure, Cerebras Systems is challenging the limits of AI processing with a fundamentally different approach from traditional GPU architectures. The company’s Wafer Scale Engine (WSE) uses an entire silicon wafer as a single chip—a design that defies semiconductor industry conventions and achieves unprecedented processing speeds.
Announced in March 2024, the WSE-3 features 4 trillion transistors and 900,000 AI-optimized cores, achieving processing speeds of 450 tokens per second for the Llama 3.1 70B model and 1,800 tokens per second for the 8B model. This represents over 20x the speed of comparable NVIDIA GPU solutions12. TIME Magazine selected this technology as one of the Best Inventions of 20243.
What is Cerebras: An AI Infrastructure Company Built by Serial Entrepreneurs
Cerebras Systems is an AI infrastructure company founded in 2015 by five co-founders, including Andrew Feldman. Feldman is known as the founder of SeaMicro, a serial entrepreneur who sold that company to AMD in 2012 for $334 million ($281 million in cash plus $53 million in stock)4.
Before SeaMicro, Feldman was involved in launching networking companies including Force10 Networks and Riverstone Networks. With an MBA from Stanford University, he is recognized as a pioneer in energy-efficient computing who forever changed the server industry trajectory by creating the microserver category5.
The Co-founding Team
Cerebras’ founding team consists of members who worked together at SeaMicro:
- Andrew Feldman (CEO)
- Gary Lauterbach
- Michael James
- Sean Lie
- Jean-Philippe Fricker
The team had raised over $720 million in total funding as of November 20216.
WSE-3’s Technical Innovation: Why It’s So Fast
Fundamental Solution to the Memory Bottleneck
In traditional GPU architectures, memory is located in external HBM (High Bandwidth Memory) outside the compute cores. During LLM inference, parameters for each layer must be loaded from memory into compute cores, matrix multiplication performed, and this process repeated across all transformer blocks. During this time, most of the GPU’s compute capacity sits idle waiting for parameters to be fetched from HBM7.
Cerebras solved this problem fundamentally. The WSE-3 features 44GB of SRAM directly on silicon, approximately 1,000 times the capacity of an NVIDIA H100. This SRAM is distributed near compute cores, eliminating the need for external memory access during inference8.
Overwhelming Memory Bandwidth
WSE-3 technical specifications9:
- Memory Bandwidth: 21 petabytes/second (7,000x the H100)
- Processor Interconnect Bandwidth: 214 petabits/second (3,715x graphics processors)
- Transistor Count: 4 trillion
- AI Core Count: 900,000
- Peak Performance: 125 petaflops
- Chip Size: 46,225mm² (over 56x the maximum size of traditional chips at ~815mm²)
On-Wafer Interconnect Innovation
The WSE-3’s on-wafer interconnect eliminates communication delays and inefficiencies from connecting hundreds of small devices via wires and cables. With all communication and memory on a single silicon slice, data moves unimpeded, achieving core-to-core bandwidth of 1,000 petabits per second and SRAM-to-core bandwidth of 9 petabytes per second10.
“It’s not just a little more,” says Feldman. “It’s four orders of magnitude greater bandwidth, because we stay on silicon”11.
Real-World Performance: The Numbers Tell the Story
Inference Speed Achievements
Performance at August 2024 announcement12:
- Llama 3.1 8B: 1,800 tokens per second (2.4x faster than Groq)
- Llama 3.1 70B: 450 tokens per second (the only platform enabling instant responses)
By November 2024, with further optimization13:
- Llama 3.1 70B: Achieved 2,100 tokens per second
Training Speed Improvements
WSE-3 serves as the foundation for the CS-3 computer system, compared to NVIDIA DGX H10015:
- Training Speed: 8x faster
- Maximum Model Size: Supports up to 24 trillion parameters
- Llama 70B Training Time: 30 days on GPUs completed in 1 day on CS-3 cluster
- Power Efficiency: One-third the power consumption of DGX solutions
Why Such a Dramatic Difference in LLM Inference?
The relationship between LLM characteristics and architecture creates Cerebras’ advantage. LLM inference is inherently sequential—generating each word requires passing through the entire model. One word takes one pass, 100 words require 100 passes, and since each word depends on the previous one, this process cannot be parallelized16.
Traditional GPU architectures require repeatedly loading model weights from external memory, creating a bottleneck. Cerebras’ wafer-scale approach fundamentally solves this memory bandwidth bottleneck by keeping the entire model on-chip.
From a developer’s perspective, the hardware appears as a gigantic GPU with all model weights on-chip. As a result, inference runs “unreasonably fast” at over 2,500 tokens per second17.
Key Partnerships and Deployments
Strategic Partnership with Meta
In April 2025, Meta announced a partnership with Cerebras to power the new Llama API with Cerebras technology, enabling developers to access inference speeds up to 18x faster than traditional GPU-based solutions.
Other Customers
- G42: AI application deployment in the Middle East
Cerebras’ Vision: The Transformation Brought by Instant AI Inference
The Cerebras Scaling Law
Cerebras advocates its own “scaling law”—the concept that improved inference speed doesn’t just reduce response times but enhances AI “intelligence” itself21.
What high-speed inference enables:
- Complex Workflows: Real-time execution of multi-model processing
- Interactive Experiences: Continuous user dialogue without waiting
- Large-Scale Reasoning Chains: Deep expansion of Chain of Thought
Future Developments
The company’s technology was selected for Forbes AI 50 (April 2024) and TIME’s 100 Most Influential Companies (May 2024)23.
For those interested in deeper understanding of Cerebras’ technological innovation, several resources are available.
The Cerebras official site provides the latest technical specifications and case studies. The architecture deep dive explains technical details of hardware-software co-design. IEEE Spectrum’s technical article offers third-party technical evaluation.
Sources
- Cerebras Launches the World’s Fastest AI Inference - Cerebras official announcement (August 2024)
- Introducing Cerebras Inference: AI at Instant Speed - Cerebras official blog
- TIME Best Inventions 2024 - TIME Magazine selection
- Cerebras - Wikipedia - Company overview and history
- Andrew Feldman Interview - Founder interview
- Cerebras Systems - Crunchbase - Funding information
- Beyond GPUs: Cerebras’ Wafer-Scale Engine - Technical explanation
- Cerebras Architecture Deep Dive - Architecture details
- Cerebras WSE-3 Announcement - WSE-3 official announcement
- IEEE Spectrum: Cerebras Chip - Technical specification details
- IEEE Spectrum: Giant Chip Analysis - Technical analysis
- Cerebras Inference Launch - Inference service launch
- How Cerebras Made Inference 3X Faster - Performance improvement details
- Data Center Knowledge Report - CS-3 system details
- The Cerebras Scaling Law - Scaling law
- Product - Chip - Product specifications
- Cerebras Scaling Law Blog - Scaling law explanation
- Forbes AI 50 - Forbes AI 50 selection
Was this article helpful?
Thank you!
Received. Thank you!