Cerebras unveils CS-4 AI system claiming 30x faster inference than Nvidia GPUs
Cerebras Systems launched the CS-4, a rack-scale AI accelerator claiming 30x faster inference than Nvidia GPUs. The system uses three Wafer Scale Engine 3 Turbo processors, delivering 750 PFLOPs of compute. OpenAI is testing the technology with its Ultrafast tier. Despite the launch, Cerebras shares fell 12.7% due to rising bond yields.

*this image is generated using AI for illustrative purposes only.
Cerebras Systems (NASDAQ: CBRS) has officially unveiled its fourth-generation CS-4 system, marking a significant expansion in its hardware capabilities. The new rack-scale solution is built from three newly released Wafer Scale Engine 3 Turbo (WSE-3T) processors. According to the company, the CS-4 delivers up to 30 times faster inference speeds than competing GPU solutions and up to 10 times more throughput per watt than the previous CS-3 generation.
The CS-4 represents the first iteration of the Cerebras Nexus platform architecture. It provides 750 PFLOPs of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O bandwidth. By integrating these components into a modular design, Cerebras aims to reduce deployment time from days to hours while supporting models with over 50 trillion parameters.
Technical Specifications
The CS-4 is powered by the WSE-3T, which contains four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon. Each wafer includes 44GB of SRAM integrated directly on-chip. The system’s modular "backpack" design decouples compute from power supplies, reducing component count by 50% compared to the prior generation.
Key performance improvements include:
- Compute: Doubles AI compute to 250 PFLOPS per wafer (750 PFLOPs total for CS-4).
- Memory Bandwidth: Doubles to 43.2 petabytes per second per wafer.
- Latency: Wafer-to-wafer latency drops to as low as two microseconds, enabling massive cluster creation.
- Power Efficiency: Power conversion is moved 100x closer to processors, nearly eliminating board-level power loss.
| Metric | CS-3 (one wafer) | CS-4 (3 wafers) |
|---|---|---|
| AI compute | 125 PFLOPS | 750 PFLOPS |
| Memory bandwidth | 21.6 PByte/s | 129.6 PByte/s |
| On-chip fabric bandwidth | 26.7 PByte/s | 160.5 PByte/s |
| System I/O bandwidth | 1.2 Tbit/s | 7.2 Tbit/s |
| I/O latency | 5 microseconds | 2 microseconds |
Inference Performance
In head-to-head comparisons on the GPT-OSS-120B model, the CS-4 delivered more than 4,400 tokens per second per user (TPS/user). This performance is up to 30 times faster than GPU solutions under identical prompt conditions. The company states that this speed allows agentic systems to perform an order of magnitude more reasoning and verification within the same wall-clock time.
SemiAnalysis estimates CS-4 could reach about 4,000 tokens per second per user on frontier models, compared with roughly 100 to 200 for Nvidia’s Blackwell chips. The performance jump comes partly from running the existing silicon much harder. SemiAnalysis estimates a three-wafer CS-4 rack at roughly 125 to 135 kilowatts and says performance per watt improves only modestly over the previous generation. Even so, The Register notes that is well below the 240 to 250 kilowatt racks Nvidia and AMD are preparing to ship later this year.
Commercial Context
This product launch follows a period of significant commercial traction for Cerebras. In the second quarter, the company signed six individual contracts, each valued at north of $30 million. This deal activity signals growing enterprise adoption of its silicon architecture.
The CEO previously highlighted that the company’s total revenue pipeline (RPO) stands at $25.4 billion. He clarified that this figure does not currently reflect any backlog from Amazon Web Services (AWS) or other hyperscalers, distinguishing the current pipeline composition from broader market expectations regarding hyperscaler engagement.
OpenAI is already putting that speed to work. The company last week previewed an Ultrafast tier powered by Cerebras that runs its flagship GPT-5.6 Sol model at up to 750 output tokens per second, up to 14 times faster than standard processing.
Market Reaction
The launch comes during a rough stretch for AI stocks. Cerebras shares fell 12.7% Tuesday, erasing Monday’s 15% rally, as surging bond yields pressured high-growth technology stocks. Prediction markets suggest demand for Nvidia compute could remain firm. Kalshi traders put a 64% chance on H200 rental prices ending the year above $6.69 an hour and a 61% chance on H100 prices staying above $3.38.
What the Numbers Show
The combination of doubled compute density and halved latency positions the CS-4 as a specialized inference engine rather than a general-purpose training accelerator. With memory bandwidth jumping from 21.6 PByte/s to 129.6 PByte/s, the bottleneck shifts from data movement to raw compute capacity. This architectural shift supports the company’s claim of vastly improved data center economics, as higher throughput per watt directly reduces operational costs for large-scale token generation. However, the modest improvement in performance per watt noted by SemiAnalysis suggests gains are driven primarily by increased power delivery and cooling improvements rather than fundamental efficiency leaps in silicon design.
Availability
First shipments of the CS-4 begin this quarter. Full system specifications are available in the official datasheet.
How might Cerebras' $25.4 billion revenue pipeline evolve if major hyperscalers like AWS or Microsoft begin integrating the CS-4 into their cloud infrastructure?
Will the significant power consumption of the CS-4 rack (125-135 kW) limit its adoption in data centers with strict thermal constraints compared to more efficient GPU alternatives?
Could the CS-4's specialized inference architecture disrupt Nvidia's dominance in the AI inference market, or will it remain a niche solution for specific high-throughput workloads?

































