NVIDIA Puts the Groq 3 LPX Inference Accelerator Into Full Production - and Names the Benchmark Conditions

NVIDIA Puts the Groq 3 LPX Inference Accelerator Into Full Production - and Names the Benchmark Conditions

NVIDIA announced on August 24, 2026 at Hot Chips that its Groq 3 LPX AI inference accelerator has entered full production. In Artificial Analysis benchmarking it recorded 3,400 output tokens per second running Gemma 4 31B at a 100,000-token context. AI cloud Nebius is the first to adopt it.

At the Hot Chips conference on August 24, 2026, NVIDIA announced that its AI inference accelerator NVIDIA Groq 3 LPX has entered full production1. The company calls it an “interactive AI inference accelerator” and positions it as an extension of the Vera Rubin platform1. CNBC framed the announcement as the moment technology from NVIDIA’s largest-ever acquisition became commercial3.

The number NVIDIA put forward is 3,400 output tokens per second, recorded in Artificial Analysis benchmarking. But that figure comes from running Gemma 4 31B, an open source agentic model, at a 100,000-token context, and NVIDIA describes it as “the fastest performance ever recorded for the model”1. It is a conditional record, not a general-purpose maximum — worth keeping in view.

Splitting “Speed” Into Two Problems

What stands out in the release is that it separates agentic AI’s computing demands into two distinct challenges: efficiently processing enormous amounts of context, and generating tokens with extremely low latency. NVIDIA treats these as different problems1.

Groq 3 LPX handles the second one. NVIDIA calls this interactivity — “the rate at which tokens are generated for an individual user, determining how quickly an agent can complete each step of its work”1. Faster generation, the argument goes, gives agents more time to inspect files, write and test code, call tools, verify results, and iterate, all while keeping the experience responsive1.

Agentic systems can generate massive volumes of tokens across hundreds or thousands of inference steps, which is why token generation speed determines whether complex tasks can be completed in real time1. The 100,000-token measurement condition was chosen with that in mind. Our explainer on context windows covers how long contexts are handled, which makes it easier to read what this benchmark setup is actually testing.

NVIDIA further claims that agentic tasks such as coding move from hours to minutes, and that the platform delivers 4x faster responsiveness than the nearest alternative platform1. The release does not say what that “nearest alternative platform” is.

Jensen Huang, NVIDIA’s founder and CEO, said inference is the growth engine of AI, that Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency, and that Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI1.

Who Gets It, and When

Two companies have been named so far, and both are still at the plan stage.

AI cloud Nebius is the first adopter, with plans to bring Groq 3 LPX to Nebius Token Factory, its production inference platform1. Nebius CTO Danila Shtan said generation is the phase of inference that determines how responsive an AI system actually is, and that as the first AI cloud bringing it to production via Nebius Token Factory, the company is making sure every step of an agent’s loop feels instant — through the same API developers already use, with no migration to a new stack1.

Following Nebius, purpose-built AI inference cloud Groq plans to be among the platform’s earliest adopters1.

Timing is absent from the release, but CNBC’s reporting fills it in. The Nebius deployment is expected to come online before the end of the year, sitting there beside Vera CPUs and Rubin GPUs; the timing came from NVIDIA senior director Dion Harris in a briefing with reporters3.

The same report sketches some of the hardware. A single LPX rack holds 256 Groq 3 chips. Samsung fabricates them — not TSMC, which builds NVIDIA’s GPUs. And each die carries 500 megabytes of SRAM, intended to keep memory from becoming the constraint3.

Neither the release nor the reporting gives concrete pricing or supply figures. If you are running agents in production and inference latency is your bottleneck, what you have so far is “this option is now in volume production” — not enough to put into a cost model.

Where the Name “Groq” Comes From

The trademark note at the end of the release states that Groq and LPU are used under license from Groq, Inc.1. Behind that sentence sits a substantial deal from December 2025.

The sum involved, as reported by CNBC, was about $20 billion in cash — paid not for a company but for the assets of Groq, a chip startup building high-performance AI accelerators. Nothing NVIDIA had bought before came close; the prior record was Mellanox in 2019, at roughly $7 billion4.

That distinction between assets and company is the whole point. Per CNBC, Groq described the arrangement as a non-exclusive licence to NVIDIA covering its inference technology. Its founder and CEO Jonathan Ross, its president Sunny Madra, and other senior leaders were to join NVIDIA, while the company itself was to continue as an independent business, with its cloud arm left outside the deal entirely4. Jensen Huang reportedly told staff in an internal note that NVIDIA was adding talented people and licensing Groq’s IP, but was not acquiring Groq as a company4.

With that structure in view, the $350 million Series A that Groq announced — covered here in August, where a company known for its own silicon named NVIDIA hardware as the use of proceeds — reads more naturally. That said, neither this release nor the earlier announcements say Groq has stepped back from designing its own chips.

Designing at the Rack Level

A recurring theme in the release is that the unit of design is not a chip but the AI factory as a whole. NVIDIA describes Vera Rubin as a platform spanning seven chips and five purpose-built racks1.

Rack platforms including Groq 3 LPX feature BlueField-4 DPUs and work in combination with Vera CPU racks, Vera BlueField-4 STX storage, and Spectrum-6 SPX Ethernet to optimize multi-agent systems for the highest throughput per watt and the lowest-latency inference1.

One of those components, the Vera CPU, got its own press release the same day: SpaceXAI will adopt it, and Vera is described as featuring 88 Olympus cores and up to 1.2TB/s of memory bandwidth2. It reinforces NVIDIA’s argument, from the hardware side, that GPUs are not the only thing that determines inference speed.

This Month’s Race to Sell Speed

Explicitly packaging inference speed as a product has been a running theme for weeks. On August 13, Cerebras announced that it would supply its inference technology to “Ultrafast mode,” a new service tier of the OpenAI API, running GPT-5.6 Sol at up to 14x the speed of standard processing. CNBC’s reporting places this in a crowded field: the figure Ultrafast mode advertises is 750 tokens per second, with Cerebras behind it, and AMD has said this year that Cerebras chips would go into its own rack-scale systems3.

On NVIDIA’s own side, securing sites and power is moving in parallel. On August 17 the company announced it had exclusively secured land in Ohio through a partnership with SB Energy. Chips, racks, power, land — the range the company gathers under the phrase “AI factory” keeps widening.

That said, this class of chip does not replace the GPU. As reported by CNBC, Harris framed it as a matter of matching the right processor at the right price to each part of the workload rather than displacing GPUs. The territory these chips cover sits mostly in the stage of model serving known as “decode”3.

The reach of this announcement also has limits. The 3,400 tokens per second figure is measured on one model at one context length; what your own models would do is a separate question. So is a “4x” with no named comparison. Unless you have first worked out whether your agents are slow because of token generation, or because of context processing and tool-call round trips, numbers like these have nothing to attach to.

Sources

  1. NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI - NVIDIA official press release (August 24, 2026)
  2. SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scale - NVIDIA official press release (August 24, 2026)
  3. Nvidia says Groq racks will be online this year following $20 billion purchase - CNBC (August 24, 2026)
  4. Nvidia buying AI chip startup Groq’s assets for about $20 billion in its largest deal on record - CNBC (December 24, 2025)

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →