NVIDIA CFO expects Groq 3 LPX volume shipments this quarter
- NVIDIA CFO expects volume shipments of Groq 3 LPX to early adopters later this quarter
- Nebius confirmed as the first customer for the new inference accelerator
- System delivers record 3,400 output tokens per second in benchmark tests
- Platform offers 4x faster responsiveness compared to nearest alternatives

*this image is generated using AI for illustrative purposes only.
NVIDIA Corporation (NASDAQ: NVDA) CFO stated on a conference call that the company expects to ship the NVIDIA Groq 3 LPX in volume to early adopters later this quarter.
Nebius, a leading AI cloud provider, is confirmed as the first customer for the platform. The accelerator extends the NVIDIA Vera Rubin platform, designed to boost token generation rates for latency-sensitive agentic systems.
Performance Benchmarks
In Artificial Analysis benchmarking running Gemma 4 31B, an open source agentic model, the system delivered a record 3,400 output tokens per second. This test utilized a 100,000-token context, which is critical for agentic systems requiring massive volumes of tokens across hundreds or thousands of inference steps.
The platform enables agentic tasks such as coding in minutes rather than hours. It provides 4x faster responsiveness for agents compared to the nearest alternative platform.
| Metric | Value |
|---|---|
| Output Speed | 3,400 tokens per second |
| Model Tested | Gemma 4 31B |
| Context Window | 100,000 tokens |
| Responsiveness Gain | 4x vs nearest alternative |
Cloud Adoption
Nebius plans to bring NVIDIA Groq 3 LPX to Nebius Token Factory, its production inference platform. Danila Shtan, chief technology officer of Nebius, noted that generation determines how responsive an AI system actually is. As the first AI cloud bringing it to production, Nebius aims to make every step of an agent’s loop feel instant through existing APIs.
Purpose-built AI inference cloud Groq also plans to be among the earliest adopters of the platform.
How might the 4x responsiveness gain impact the competitive landscape between NVIDIA, Groq, and other AI inference providers in the enterprise sector?
What are the projected cost implications for Nebius customers when migrating to the Groq 3 LPX platform compared to current GPU-based inference solutions?
Will the success of the Vera Rubin platform accelerate the shift from training-focused to inference-focused hardware investments among major cloud providers?

































