Nvidia expands Vera Rubin with Groq 3 LPX for agents

Nvidia expands Vera Rubin with Groq 3 LPX for agents

Nvidia confirms that Groq 3 LPX is in full production. The inference accelerator complements Vera Rubin NVL72 with fast token generation for agentic AI. Nebius, CoreWeave, and SpaceXAI are the first customers.

Groq 3 LPX is a rack-scale system designed in conjunction with Vera Rubin NVL72. In a benchmark by Artificial Analysis using the open-source model Gemma 4 31B, the system achieved 3,400 output tokens per second with a long context size of 100,000 tokens. According to Nvidia, that’s four times faster than the closest alternative.

GPUs and LPUs work together

Rubin GPUs handle the heavy-duty context processing, while the LPUs focus on the latency-sensitive decoding phase. That is precisely where the bottleneck lies: an agent generates tokens one by one, and every delay accumulates in long chains of tasks.

A rack can contain 256 LP30 accelerators, connected via direct chip-to-chip links. According to reports on the specifications, this amounts to approximately 315 PFLOPS of FP8 computing power and 128 GB of on-chip SRAM per rack.

Nebius, CoreWeave, and SpaceXAI join the fray

Nebius is the first AI cloud to purchase Groq 3 LPX and is deploying the system within its Token Factory. CoreWeave is now running Spectrum-X Multiplane in production to interconnect Vera Rubin racks. SpaceXAI is opting for Vera CPUs for orchestration, tooling, code execution, and simulation, from data centers on Earth to satellites in orbit.

Spectrum-X Multiplane splits each server connection into multiple independent paths. This lets a flat, two-layer network scale to 512,000 GPUs without a third network layer. If one plane fails in an eight-plane topology, approximately 90 percent of the bandwidth remains intact.

Tip: SpaceXAI is sending Nvidia Vera Rubin into space