Cerebras unveils CS-4 AI system claiming 30x faster inference than Nvidia GPUs

scanx
Reviewed by
Jubin VScanX News Team
Key Highlights

Cerebras Systems launched the CS-4, a rack-scale AI accelerator claiming 30x faster inference than Nvidia GPUs. The system uses three Wafer Scale Engine 3 Turbo processors, delivering 750 PFLOPs of compute. OpenAI is testing the technology with its Ultrafast tier. Despite the launch, Cerebras shares fell 12.7% due to rising bond yields.

powered bylight_fuzz_icon
48118857

*this image is generated using AI for illustrative purposes only.

Cerebras Systems (NASDAQ: CBRS) has officially unveiled its fourth-generation CS-4 system, marking a significant expansion in its hardware capabilities. The new rack-scale solution is built from three newly released Wafer Scale Engine 3 Turbo (WSE-3T) processors. According to the company, the CS-4 delivers up to 30 times faster inference speeds than competing GPU solutions and up to 10 times more throughput per watt than the previous CS-3 generation.

The CS-4 represents the first iteration of the Cerebras Nexus platform architecture. It provides 750 PFLOPs of AI compute, 129.6 petabytes per second of memory bandwidth, and 7.2 terabits per second of I/O bandwidth. By integrating these components into a modular design, Cerebras aims to reduce deployment time from days to hours while supporting models with over 50 trillion parameters.

Technical Specifications

The CS-4 is powered by the WSE-3T, which contains four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon. Each wafer includes 44GB of SRAM integrated directly on-chip. The system’s modular "backpack" design decouples compute from power supplies, reducing component count by 50% compared to the prior generation.

Key performance improvements include:

  • Compute: Doubles AI compute to 250 PFLOPS per wafer (750 PFLOPs total for CS-4).
  • Memory Bandwidth: Doubles to 43.2 petabytes per second per wafer.
  • Latency: Wafer-to-wafer latency drops to as low as two microseconds, enabling massive cluster creation.
  • Power Efficiency: Power conversion is moved 100x closer to processors, nearly eliminating board-level power loss.
Metric CS-3 (one wafer) CS-4 (3 wafers)
AI compute 125 PFLOPS 750 PFLOPS
Memory bandwidth 21.6 PByte/s 129.6 PByte/s
On-chip fabric bandwidth 26.7 PByte/s 160.5 PByte/s
System I/O bandwidth 1.2 Tbit/s 7.2 Tbit/s
I/O latency 5 microseconds 2 microseconds

Inference Performance

In head-to-head comparisons on the GPT-OSS-120B model, the CS-4 delivered more than 4,400 tokens per second per user (TPS/user). This performance is up to 30 times faster than GPU solutions under identical prompt conditions. The company states that this speed allows agentic systems to perform an order of magnitude more reasoning and verification within the same wall-clock time.

SemiAnalysis estimates CS-4 could reach about 4,000 tokens per second per user on frontier models, compared with roughly 100 to 200 for Nvidia’s Blackwell chips. The performance jump comes partly from running the existing silicon much harder. SemiAnalysis estimates a three-wafer CS-4 rack at roughly 125 to 135 kilowatts and says performance per watt improves only modestly over the previous generation. Even so, The Register notes that is well below the 240 to 250 kilowatt racks Nvidia and AMD are preparing to ship later this year.

Commercial Context

This product launch follows a period of significant commercial traction for Cerebras. In the second quarter, the company signed six individual contracts, each valued at north of $30 million. This deal activity signals growing enterprise adoption of its silicon architecture.

The CEO previously highlighted that the company’s total revenue pipeline (RPO) stands at $25.4 billion. He clarified that this figure does not currently reflect any backlog from Amazon Web Services (AWS) or other hyperscalers, distinguishing the current pipeline composition from broader market expectations regarding hyperscaler engagement.

OpenAI is already putting that speed to work. The company last week previewed an Ultrafast tier powered by Cerebras that runs its flagship GPT-5.6 Sol model at up to 750 output tokens per second, up to 14 times faster than standard processing.

Market Reaction

The launch comes during a rough stretch for AI stocks. Cerebras shares fell 12.7% Tuesday, erasing Monday’s 15% rally, as surging bond yields pressured high-growth technology stocks. Prediction markets suggest demand for Nvidia compute could remain firm. Kalshi traders put a 64% chance on H200 rental prices ending the year above $6.69 an hour and a 61% chance on H100 prices staying above $3.38.

What the Numbers Show

The combination of doubled compute density and halved latency positions the CS-4 as a specialized inference engine rather than a general-purpose training accelerator. With memory bandwidth jumping from 21.6 PByte/s to 129.6 PByte/s, the bottleneck shifts from data movement to raw compute capacity. This architectural shift supports the company’s claim of vastly improved data center economics, as higher throughput per watt directly reduces operational costs for large-scale token generation. However, the modest improvement in performance per watt noted by SemiAnalysis suggests gains are driven primarily by increased power delivery and cooling improvements rather than fundamental efficiency leaps in silicon design.

Availability

First shipments of the CS-4 begin this quarter. Full system specifications are available in the official datasheet.

Disclaimer: This article is AI-generated using data from ViewTrade. ScanX is not liable for any inaccuracies.

How might Cerebras' $25.4 billion revenue pipeline evolve if major hyperscalers like AWS or Microsoft begin integrating the CS-4 into their cloud infrastructure?

Will the significant power consumption of the CS-4 rack (125-135 kW) limit its adoption in data centers with strict thermal constraints compared to more efficient GPU alternatives?

Could the CS-4's specialized inference architecture disrupt Nvidia's dominance in the AI inference market, or will it remain a niche solution for specific high-throughput workloads?

like15
dislike

Cerebras stock falls 12.69% as revenue misses estimates

scanx
Reviewed by
Riya DScanX News Team
Key Highlights

Cerebras Systems Inc. (NASDAQ: CBRS) reported Q3 revenue of $180.11 million, missing the $194.20 million estimate, causing shares to fall 12.69%. However, core revenue surged 103% YoY to $209.9 million. The company launched the CS-4 AI server, claiming 30x faster inference than GPUs, and guided for Q4 core revenue of $214-$216 million with 38-40% gross margins.

powered bylight_fuzz_icon
48667043

*this image is generated using AI for illustrative purposes only.

Cerebras Systems Inc. (NASDAQ: CBRS) saw its shares decline 12.69% to close at $220.01 on Tuesday, extending to a 1.32% drop to $217.10 in after-hours trading. The sell-off followed the release of its latest earnings report, where total sales of $180.11 million fell short of the $194.20 million analyst consensus.

Despite the headline miss, core revenue demonstrated strong momentum, rising 103% year over year to $209.9 million. This divergence suggests that while overall top-line figures were pressured by specific items not detailed in the filing, the underlying business operations are expanding rapidly.

What the Numbers Show

The gap between total reported sales ($180.11 million) and core revenue ($209.9 million) indicates that non-core items reduced the reported top line by approximately $29.8 million. While the source does not specify the nature of these deductions, the fact that core revenue exceeded total sales highlights a structural difference in how the company reports its primary business performance versus its consolidated financial results. Investors appear to have penalized the stock for missing the broader consensus estimate, despite the robust double-digit growth in the core segment.

New CS-4 Server Targets Nvidia

On Tuesday, Cerebras unveiled the CS-4, a rack-scale AI system designed to challenge Nvidia Corp (NASDAQ: NVDA) in AI inference. The system is powered by three Wafer Scale Engine-3 Turbo chips and built using Taiwan Semiconductor Manufacturing Co.’s (NYSE: TSM) 5-nanometer process.

Key features of the CS-4 include:

  • Nexus Architecture: Uses upgraded networking components to improve data movement between chips.
  • Component Reduction: Contains 50% fewer components than previous designs, potentially accelerating data center deployment.
  • Performance Claims: Cerebras states the CS-4 can deliver up to 30 times faster inference than GPU-based systems under certain workloads.

The company expects the new system to become available in the third quarter. CEO Andrew Feldman stated that Cerebras plans to deliver 600 megawatts of computing capacity by the end of 2027, with performance improving fourfold and throughput increasing 20 times by then.

Strategic Partnerships and Outlook

Cerebras is focusing on AI inference, the process behind generating answers from chatbots such as Anthropic’s Claude. The company is collaborating with Amazon.com, Inc. (NASDAQ: AMZN) Web Services, with Cerebras solutions expected to become available through AWS Bedrock in the first quarter of 2027.

Looking ahead, Cerebras provided guidance for the upcoming quarter:

Metric Forecast Range
Core Revenue $214 million to $216 million
Gross Margin 38% to 40%

Benzinga Edge ranks Cerebras favorably, noting positive price trends across short, medium, and long-term horizons.

Disclaimer: This article is AI-generated using data from ViewTrade. ScanX is not liable for any inaccuracies.

How might the upcoming Q3 launch of the CS-4 server impact Cerebras' ability to capture market share from Nvidia in the AI inference sector?

What specific non-core accounting items caused the $29.8 million discrepancy between total sales and core revenue, and will this reporting structure persist?

Will the integration with AWS Bedrock in early 2027 significantly accelerate adoption rates for Cerebras' inference solutions among enterprise clients?

like18
dislike

More News on Cerebras Systems