Cerebras Systems introduced the CS-4 on August 18 at its Supernova 2026 event, a rack-scale AI system the company says delivers inference speeds up to 30 times faster than GPU-based servers, according to Cerebras’ official announcement.
The system packs three of Cerebras’ new Wafer Scale Engine 3 Turbo (WSE-3T) processors — each built on the same 5-nanometer, 4-trillion-transistor wafer as the CS-3, but clocked roughly twice as fast — into a redesigned rack the company calls the Nexus platform. Cerebras says the new design cuts the rack’s component count in half while doubling power delivery to each chip. On models with more than 10 trillion parameters, the company says the CS-4 generates over 1,000 tokens per second per user, with wafer-to-wafer latency as low as two microseconds and up to 10 times more token throughput per watt than the CS-3.
Built for split inference
Many inference pipelines now split an AI query into two stages: “prefill,” which processes the incoming prompt, and “decode,” which generates the reply token by token. Cerebras says the CS-4 is designed to plug into this split setup, pairing with AMD’s Helios systems and AWS’s Trainium chips to handle prefill while its own wafer-scale hardware, which the company says is fastest at decode, generates the output.
The launch builds on Cerebras’ push to sell inference speed as a differentiator distinct from raw training power. The company partnered with OpenAI earlier this month on an “Ultrafast” mode for GPT-5.6 Sol running on the same wafer-scale approach.
Availability
Cerebras says the first CS-4 shipments are scheduled for the third quarter of 2026, with broader availability following later in the year. The company has not disclosed pricing.
The launch adds to a summer of AI hardware announcements aimed at cutting the cost of running ever-larger models at scale, as inference compute remains one of the biggest expenses for AI labs and cloud providers.