CoreWeave trains DeepSeek-V3 in 2 minutes in MLPerf benchmark
CoreWeave, Inc. set a new record in the MLPerf Training v6.0 benchmark by training the DeepSeek-V3 671B model in 2.02 minutes using 8,192 NVIDIA GB300 NVL72 GPUs. The company demonstrated near-linear scaling efficiency across three different cluster sizes and achieved top results for Llama-3.1-405B and GPT-OSS-20B models. These benchmarks were conducted on the same production infrastructure available to customers, utilizing full-stack optimizations including CoreWeave Mission Control and a topology-aware scheduler.

*this image is generated using AI for illustrative purposes only.
CoreWeave, Inc. (NASDAQ: CRWV) announced record-breaking results in the MLPerf Training v6.0 benchmark suite, training the DeepSeek-V3 671B model in 2.02 minutes on 8,192 NVIDIA GB300 NVL72 GPUs. The performance marks the fastest DeepSeek-V3 training result in the benchmark and was achieved on the largest GB300 cluster submitted in this round. This speed addresses the critical constraint of training performance as frontier models scale to trillion-parameter sizes and agentic workloads become standard.
The benchmark results reflect CoreWeave's full-stack infrastructure optimizations across networking, orchestration, scheduling, storage, and software. The company submitted three GB300 NVL72 configurations on the DeepSeek-V3 671B workload, achieving the fastest results across all Closed/Available-cloud submissions. CoreWeave was the only submitter in the v6.0 round to scale a GB300 platform beyond 2,048 GPUs on this specific workload.
DeepSeek-V3 671B Performance Metrics
CoreWeave demonstrated consistent, near-linear scaling efficiency as the cluster size doubled. The training time improved predictably across the different node configurations.
| GPUs | Nodes | Training Time (Minutes) |
|---|---|---|
| 8,192 | 2,048 | 2.02 |
| 4,096 | 1,024 | 3.09 |
| 2,048 | 512 | 5.54 |
Additional Benchmark Results
Beyond the DeepSeek-V3 results, CoreWeave reported performance metrics for other models on different hardware configurations. On a 4,096-GPU NVIDIA GB300 NVL72 deployment, the company reached the Llama-3.1-405B reference quality target in 9.77 minutes. This run utilized the NVIDIA NeMo Framework Release 26.04, CUDA graphs, and NVIDIA Spectrum-X Ethernet running RoCE.
On a smaller 8-node, 64-GPU NVIDIA HGX B200 cluster connected via InfiniBand, CoreWeave trained GPT-OSS-20B in 26.98 minutes and Llama-3.1-8B in 16.54 minutes. The company attributed these results to optimizations in orchestration, communication libraries, and distributed training configuration.
Infrastructure and Optimization
CoreWeave attributed its performance to several key infrastructure layers. CoreWeave Mission Control performs continuous health checks across rack-scale systems to validate hardware, firmware, network, and thermal health. The CoreWeave SUNK scheduler is topology-aware, placing workloads to maximize locality and minimize inter-rack communication for Mixture of Experts (MoE) workloads. Additionally, a rail-aware networking strategy balances traffic to prevent hotspots within the fabric at multi-thousand-GPU scale.
Chen Goldberg, Executive Vice President of Product and Engineering at CoreWeave, stated that the results came from the same infrastructure customers run in production today. Brendan Burke, Research Director at Futurum Research, noted that the results demonstrate full-stack AI expertise compounds real-world performance gains as new hardware arrives.
How will CoreWeave's record-breaking training speeds influence the pricing models and competitive positioning of its cloud services against hyperscalers?
Can CoreWeave maintain this near-linear scaling efficiency as it expands to clusters exceeding 16,000 GPUs for trillion-parameter models?
What impact will these benchmarks have on enterprise adoption of CoreWeave for latency-sensitive agentic AI workloads?



























