AMD Helios Enters Full Production: 72 MI455X GPUs per Rack, Up to 2GW Anthropic Deal and $5B Investment

AMD's Helios rack-scale AI system is now in full production, pairing 72 Instinct MI455X GPUs with EPYC Venice CPUs to challenge NVIDIA. OpenAI targets Q4 2026 deployment, while Anthropic commits to up to 2 gigawatts backed by an AMD equity investment of up to $5 billion.

AMD Helios Enters Full Production: 72 MI455X GPUs per Rack, Up to 2GW Anthropic Deal and $5B Investment

AMD announced at its Advancing AI 2026 event in San Francisco on July 23, 2026 that Helios, the company’s first rack-scale AI system, is now in full production1. CEO Lisa Su told the audience that “Helios is in full production,” and named OpenAI, Meta, Microsoft, and Oracle among the major AI companies and cloud providers lining up to deploy it13.

The day before, on July 22, AMD and Anthropic separately announced a strategic partnership under which Anthropic will deploy up to 2 gigawatts of AMD Instinct MI450 Series GPUs, with AMD committing to a strategic equity investment of up to $5 billion in Anthropic2. Together, the announcements mark the arrival of a production-ready alternative in a training and inference infrastructure market that NVIDIA’s rack systems have largely dominated.

Inside Helios: 72 MI455X GPUs and EPYC “Venice”

Helios integrates GPUs, CPUs, networking, and software at the rack level, combining 72 Instinct MI455X GPUs with 18 6th Gen EPYC “Venice” CPUs, connected through AMD’s Pensando networking1. AMD claims the configuration delivers “up to 30% more tokens per dollar than the leading competitive solution” — a design squarely aimed at NVIDIA’s 72-GPU racks such as Vera Rubin NVL721.

The 6th Gen EPYC “Venice” CPUs announced alongside Helios top out at 256 cores and 512 threads, with support for up to 16 channels of 12.8 GT/s MRDIMM memory and PCIe Gen 6 connectivity1. On the GPU side, AMD says the MI455X delivers 34x higher token throughput than the MI355X, while the MI350P — targeted at existing infrastructure — claims “up to 4.2x more tokens per second per dollar than the competition”1. According to ITPro, AMD also positions Helios as offering “15% more AI compute, 50% more memory, and 50% more scale-out bandwidth” than NVIDIA’s Vera Rubin NVL724.

These are all vendor claims, and the benchmark conditions have not been disclosed. Still, leading with cost per token rather than raw performance speaks directly to the industry’s current preoccupation with driving down inference costs.

The Customer List: OpenAI in Q4, Anthropic Scaling Up from MI355X

The roster of named adopters is broad. OpenAI expects to bring Helios online beginning in the fourth quarter of 2026, with deployments accelerating throughout 20271. Meta has begun testing and validating workloads on Helios racks1, and Microsoft, Oracle, HPE, Lenovo, Supermicro, AT&T, and Cisco also appear on the deployment side14. AMD is additionally collaborating with Cerebras on ultra-low-latency inference serving1.

The Anthropic partnership is the most concrete. Under the July 22 agreement, Anthropic will deploy up to 2 gigawatts of MI450 Series GPUs (MI455X in the Helios configuration), with deployment of the first gigawatt beginning in the first half of 20272. Anthropic already uses AMD’s MI355X GPUs, making this a substantial expansion2. Alongside the equity commitment of up to $5 billion, the two companies launched a multiyear engineering collaboration that will use Claude to optimize workloads for AMD Instinct GPUs and accelerate ROCm software development — and AMD will broadly adopt Claude across its engineering and product development teams2.

“By partnering with AMD across the stack, we are securing the capacity we need,” said Tom Brown, Anthropic co-founder and Chief Compute Officer2.

The Backstory: Alternatives Filling the Gaps in an NVIDIA-First Market

NVIDIA’s GPUs and NVLink-based rack systems have been the default AI infrastructure for years, but alternative procurement has visibly accelerated over the past year. Inference-focused challengers have raised heavily — see SambaNova’s $1 billion round anchored by an on-premises deployment at JPMorgan Chase, and General Compute’s credit facility collateralized by SambaNova inference chips — while on the supply side, TSMC just posted record earnings on AI demand. With demand still outrunning supply, simply having a credible non-NVIDIA option carries negotiating leverage and supply-security value for buyers.

AMD’s entry is notable for being a full rack-scale system rather than a standalone GPU. On the software front, the company introduced ROCm.ai, an AI-driven development platform that supports GPU software development with coding agents such as Claude, Codex, and Cursor1 — a plausible attempt to close the gap with NVIDIA’s CUDA ecosystem by leaning on AI-assisted development.

AMD projects its total addressable market to reach roughly $2 trillion by 2030 on AI demand1, and Su said the AI accelerator market alone will reach about $1.4 trillion by 2030, with GPUs making up the vast majority3. The roadmap extends to the MI500 Series in 2027 and MI600 in 20281.

API Costs and Claude’s Compute Runway

For companies consuming AI through APIs, this announcement won’t change prices overnight. But model providers’ inference costs set the floor for API pricing, so if rack-level competition genuinely pushes down cost per token, better pricing and higher usage limits become plausible over the medium term. For Claude users specifically, Anthropic locking in up to 2 gigawatts of capacity is meaningful reassurance about the model’s supply and stability.

That said, AMD’s performance and cost claims currently rest on its own disclosures, and real-workload evaluations are still to come. Readers selecting hardware themselves should — as with choosing hardware for local LLMs — verify that benchmark conditions match their actual use case before deciding. Verifiable milestones arrive over the next year: OpenAI’s Q4 2026 deployment and the start of Anthropic’s first-gigawatt deployment in the first half of 2027.

Sources

  1. AAI 2026: AMD Delivers Full-Stack Compute for the Agentic AI Era - AMD official press release (July 23, 2026)
  2. AMD and Anthropic Announce Strategic Partnership to Deploy Up to 2 Gigawatts of AMD Instinct MI450 Series GPUs - AMD official press release (July 22, 2026)
  3. AMD takes on Nvidia with its Helios AI rack scale system - TechCrunch (July 23, 2026)
  4. AMD Advancing AI 2026: Helios on the rise, with launch of new Instinct GPUs - ITPro (July 23, 2026)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →