Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Cerebras WSE-3: The Wafer-Scale Chip Built to Speed Up AI Inference

Cerebras WSE-3: The Wafer-Scale Chip Built to Speed Up AI Inference

Cerebras Systems' WSE-3 chip features 4 trillion transistors and 44GB of on-chip SRAM, achieving over 20x faster inference than traditional GPUs. An exploration of its technological innovation and founder's vision.

In an era where AI inference speed can determine business success or failure, Cerebras Systems is challenging the limits of AI processing with a fundamentally different approach from traditional GPU architectures. The company’s Wafer Scale Engine (WSE) uses an entire silicon wafer as a single chip—a design that defies semiconductor industry conventions and achieves unprecedented processing speeds.

Announced in March 2024, the WSE-3 features 4 trillion transistors and 900,000 AI-optimized cores, achieving processing speeds of 450 tokens per second for the Llama 3.1 70B model and 1,800 tokens per second for the 8B model. This represents over 20x the speed of comparable NVIDIA GPU solutions12. TIME Magazine selected this technology as one of the Best Inventions of 20243.

What is Cerebras: An AI Infrastructure Company Built by Serial Entrepreneurs

Cerebras Systems is an AI infrastructure company founded in 2015 by five co-founders, including Andrew Feldman. Feldman is known as the founder of SeaMicro, a serial entrepreneur who sold that company to AMD in 2012 for $334 million ($281 million in cash plus $53 million in stock)4.

Before SeaMicro, Feldman was involved in launching networking companies including Force10 Networks and Riverstone Networks. With an MBA from Stanford University, he is recognized as a pioneer in energy-efficient computing who forever changed the server industry trajectory by creating the microserver category5.

The Co-founding Team

Cerebras’ founding team consists of members who worked together at SeaMicro:

  • Andrew Feldman (CEO)
  • Gary Lauterbach
  • Michael James
  • Sean Lie
  • Jean-Philippe Fricker

The team had raised over $720 million in total funding as of November 20216.

WSE-3’s Technical Innovation: Why It’s So Fast

Fundamental Solution to the Memory Bottleneck

In traditional GPU architectures, memory is located in external HBM (High Bandwidth Memory) outside the compute cores. During LLM inference, parameters for each layer must be loaded from memory into compute cores, matrix multiplication performed, and this process repeated across all transformer blocks. During this time, most of the GPU’s compute capacity sits idle waiting for parameters to be fetched from HBM7.

Cerebras solved this problem fundamentally. The WSE-3 features 44GB of SRAM directly on silicon, approximately 1,000 times the capacity of an NVIDIA H100. This SRAM is distributed near compute cores, eliminating the need for external memory access during inference8.

Overwhelming Memory Bandwidth

WSE-3 technical specifications9:

  • Memory Bandwidth: 21 petabytes/second (7,000x the H100)
  • Processor Interconnect Bandwidth: 214 petabits/second (3,715x graphics processors)
  • Transistor Count: 4 trillion
  • AI Core Count: 900,000
  • Peak Performance: 125 petaflops
  • Chip Size: 46,225mm² (over 56x the maximum size of traditional chips at ~815mm²)

On-Wafer Interconnect Innovation

The WSE-3’s on-wafer interconnect eliminates communication delays and inefficiencies from connecting hundreds of small devices via wires and cables. With all communication and memory on a single silicon slice, data moves unimpeded, achieving core-to-core bandwidth of 1,000 petabits per second and SRAM-to-core bandwidth of 9 petabytes per second10.

“It’s not just a little more,” says Feldman. “It’s four orders of magnitude greater bandwidth, because we stay on silicon”11.

Real-World Performance: The Numbers Tell the Story

Inference Speed Achievements

Performance at August 2024 announcement12:

  • Llama 3.1 8B: 1,800 tokens per second (2.4x faster than Groq)
  • Llama 3.1 70B: 450 tokens per second (the only platform enabling instant responses)

By November 2024, with further optimization13:

  • Llama 3.1 70B: Achieved 2,100 tokens per second

Training Speed Improvements

WSE-3 serves as the foundation for the CS-3 computer system, compared to NVIDIA DGX H10015:

  • Training Speed: 8x faster
  • Maximum Model Size: Supports up to 24 trillion parameters
  • Llama 70B Training Time: 30 days on GPUs completed in 1 day on CS-3 cluster
  • Power Efficiency: One-third the power consumption of DGX solutions

Why Such a Dramatic Difference in LLM Inference?

The relationship between LLM characteristics and architecture creates Cerebras’ advantage. LLM inference is inherently sequential—generating each word requires passing through the entire model. One word takes one pass, 100 words require 100 passes, and since each word depends on the previous one, this process cannot be parallelized16.

Traditional GPU architectures require repeatedly loading model weights from external memory, creating a bottleneck. Cerebras’ wafer-scale approach fundamentally solves this memory bandwidth bottleneck by keeping the entire model on-chip.

From a developer’s perspective, the hardware appears as a gigantic GPU with all model weights on-chip. As a result, inference runs “unreasonably fast” at over 2,500 tokens per second17.

Key Partnerships and Deployments

Strategic Partnership with Meta

In April 2025, Meta announced a partnership with Cerebras to power the new Llama API with Cerebras technology, enabling developers to access inference speeds up to 18x faster than traditional GPU-based solutions.

Other Customers

  • G42: AI application deployment in the Middle East

Cerebras’ Vision: The Transformation Brought by Instant AI Inference

The Cerebras Scaling Law

Cerebras advocates its own “scaling law”—the concept that improved inference speed doesn’t just reduce response times but enhances AI “intelligence” itself21.

What high-speed inference enables:

  • Complex Workflows: Real-time execution of multi-model processing
  • Interactive Experiences: Continuous user dialogue without waiting
  • Large-Scale Reasoning Chains: Deep expansion of Chain of Thought

Future Developments

The company’s technology was selected for Forbes AI 50 (April 2024) and TIME’s 100 Most Influential Companies (May 2024)23.

For those interested in deeper understanding of Cerebras’ technological innovation, several resources are available.

The Cerebras official site provides the latest technical specifications and case studies. The architecture deep dive explains technical details of hardware-software co-design. IEEE Spectrum’s technical article offers third-party technical evaluation.

Sources

  1. Cerebras Launches the World’s Fastest AI Inference - Cerebras official announcement (August 2024)
  2. Introducing Cerebras Inference: AI at Instant Speed - Cerebras official blog
  3. TIME Best Inventions 2024 - TIME Magazine selection
  4. Cerebras - Wikipedia - Company overview and history
  5. Andrew Feldman Interview - Founder interview
  6. Cerebras Systems - Crunchbase - Funding information
  7. Beyond GPUs: Cerebras’ Wafer-Scale Engine - Technical explanation
  8. Cerebras Architecture Deep Dive - Architecture details
  9. Cerebras WSE-3 Announcement - WSE-3 official announcement
  10. IEEE Spectrum: Cerebras Chip - Technical specification details
  11. IEEE Spectrum: Giant Chip Analysis - Technical analysis
  12. Cerebras Inference Launch - Inference service launch
  13. How Cerebras Made Inference 3X Faster - Performance improvement details
  14. Data Center Knowledge Report - CS-3 system details
  15. The Cerebras Scaling Law - Scaling law
  16. Product - Chip - Product specifications
  17. Cerebras Scaling Law Blog - Scaling law explanation
  18. Forbes AI 50 - Forbes AI 50 selection

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →