AMD and Cerebras Systems announced a technical partnership on July 23 that splits AI inference work across two different kinds of chips, aiming to cut the cost of running large language models at scale.
How the split works
The joint setup separates the two stages of generating an AI response. AMD’s Helios rack-scale servers, built around its Instinct GPUs, handle “prefill” — reading a prompt and its context — where raw throughput matters most. Cerebras’ wafer-scale engine, a chip built from an entire silicon wafer, then takes over “decode,” generating each output token, where memory bandwidth and low latency matter more. According to the companies, pairing the two lets each system do only what it is best at, rather than running the whole task on one type of AI chip.
Performance claims
AMD and Cerebras said the combined system delivers up to five times more tokens generated per second per watt than a Cerebras-only setup, based on internal modeling by AMD Performance Labs and Cerebras using Moonshot AI’s Kimi 2.6 model at a comparable interactivity level. AMD chief executive Lisa Su called AI inference “one of the largest infrastructure opportunities” the industry has seen, while Cerebras chief executive Andrew Feldman said the partnership lets the company “bring that performance to even more customers.”
Why it matters
Inference — running an already-trained model to answer queries — is increasingly the larger cost center for AI companies, ahead of training. The deal also positions AMD and Cerebras as a joint alternative to Nvidia, which still supplies most of the chips used for both training and inference. Cerebras plans to install AMD Helios systems in its own data centers, with the combined offering reaching Cerebras Cloud customers in the second half of 2026.