QumulusAI signs $71.9 million, three-year AI inference capacity deal
QumulusAI announces a $71.9 million, three-year deal for NVIDIA Blackwell GPU capacity, bringing its total announced agreements since June to over $246 million. The contract supports an AI inference platform serving LLM and generative AI workloads, with capacity expected in Q3 2026.

*this image is generated using AI for illustrative purposes only.
QumulusAI (NASDAQ: QMLS) has signed a three-year agreement valued at more than $71 million to supply NVIDIA Blackwell B300 and B200 GPU capacity to an unnamed AI inference platform provider. The deal, which includes renewal options, ranks among the company’s largest customer commitments announced to date and underscores growing demand for dedicated high-performance compute in production AI workloads. For investors, this agreement signals continued momentum in QumulusAI’s demand-led deployment model, adding significant contracted revenue visibility to its book of business.
The customer operates a platform that helps companies deploy large language models, vision models, speech models, and other AI applications with low latency and high reliability. Its clients specialize in LLMs, image generation, and video generation. The dedicated Blackwell capacity will provide the platform with the high-performance compute necessary to serve production inference workloads as demand from its customers scales. The agreement adds more than $71 million in contracted, multiyear commitments to QumulusAI’s total book of business.
Capacity under the agreement will be served from QumulusAI’s U.S. data center footprint and is expected to be ready for customer use in the third quarter of 2026. The company utilizes a demand-led deployment model that places capacity into available pockets of power across a distributed network of colocation and owned facilities. This approach enables QumulusAI to bring GPU capacity online in months rather than years, addressing the urgent infrastructure needs of AI enterprises.
"Inference is where AI meets the real world, and the platforms serving it can't afford to wait on capacity," said Michael Maniscalco, CEO of QumulusAI. "Our customer runs production workloads for companies that need the right combination of flexibility, access, cost, trust and speed in their infrastructure. A three-year commitment of this size reflects what our model is built to do: procure and deploy Blackwell capacity where demand already exists, and do it fast."
Recent Contract Momentum
This agreement extends a series of recent demand announcements made by QumulusAI since early June. The company has now announced more than $246 million in customer agreements during this period. Key recent deals include:
| Date | Value | Term | Customer Type | Details |
|---|---|---|---|---|
| July 23 | $32 million | Two years | AI inference platform | Focused on generative AI applications; NVIDIA Blackwell B300 |
| July 22 | >$18 million | Two years | GPU cloud marketplace | Take-or-pay agreement; serves AI teams in 100+ regions; NVIDIA Blackwell B300 |
| June 11 | $124.4 million | Three years | Two customers | Inference agreements |
| May 28 | Not specified | Not specified | Shadeform | Two NVIDIA H200 cluster deployments |
The rapid succession of these contracts highlights the scalability of QumulusAI’s distributed AI cloud platform. By combining rapid deployment with flexible private cloud infrastructure, the company aims to offer customers a faster, more adaptable path beyond the capacity constraints of traditional centralized and hyperscale cloud models.
What the Numbers Show
The concentration of large-value contracts over a short timeframe suggests strong market validation for QumulusAI’s niche focus on inference workloads. With more than $246 million in announced agreements since early June, the company is demonstrating its ability to secure multi-year commitments from diverse customer segments, including specialized inference platforms and global cloud marketplaces. This pipeline provides near-term revenue visibility while the company executes on its deployment timeline, with the latest $71 million deal coming online in Q3 2026.
How will QumulusAI secure the necessary power and cooling infrastructure to meet the Q3 2026 deployment deadline for the Blackwell capacity given current grid constraints?
What is the potential impact on QumulusAI's gross margins if NVIDIA adjusts Blackwell GPU pricing or availability in the interim before the 2026 deployment?
How does the concentration of revenue from a few large inference platform providers affect QumulusAI's risk profile compared to diversified hyperscale competitors?




























