What Is Disaggregated Inference — and How Does It Speed Up AI?
Disaggregated inference splits an AI model's two work stages onto separate chips built for each job — a technique now expanding beyond Nvidia to AMD and Cerebras hardware.
Read more →Every AI News story tagged with both AI Chips and AI Inference — the two topics side by side, updated as new articles publish.
8 articles
Disaggregated inference splits an AI model's two work stages onto separate chips built for each job — a technique now expanding beyond Nvidia to AMD and Cerebras hardware.
Read more →AMD and Cerebras unveiled a joint AI inference system that splits chip work between AMD's Helios servers and Cerebras' wafer-scale engine, claiming up to 5x higher efficiency per watt.
Read more →
AMD unveiled its Instinct MI400 AI chips and Helios rack systems on July 23, claiming 34x faster inference, with OpenAI, Meta and Anthropic on hand.
Read more →Chipmaker Etched is negotiating a funding round that would value it at $20 billion, quadrupling its December price, as investors chase alternatives to Nvidia for AI inference, the Wall Street Journal reported.
Read more →SambaNova closed the first tranche of a $1 billion Series F round led by General Atlantic, and said JPMorganChase will deploy its chips for on-premises AI inference.
Read more →Qualcomm on June 24 unveiled the Dragonfly data center chip family — a 250-core Arm-based CPU and a new inference accelerator — with Meta signed on to deploy the hardware in its server fleet from 2028.
Read more →An AI inference chip is specialized hardware built to run trained AI models quickly and cheaply at scale. Here is why every major tech company is now racing to design its own instead of buying from Nvidia.
Read more →OpenAI and Broadcom revealed Jalapeño on June 24 — a purpose-built inference chip designed in nine months with AI assistance, targeting deployment by end of 2026.
Read more →