AMD Acquires Taalas, an AI Inference Chip Startup
AMD has agreed to acquire Taalas, a Toronto startup that etches AI models directly into silicon, aiming to boost inference speed and efficiency across its chip lineup.
Read more →Inference is the stage where a trained AI model is actually put to use — generating text, images, or predictions from new input — as opposed to training, where the model first learns from data. It has become a major focus of the AI industry because inference, not training, is where most of the day-to-day compute cost and latency happens at scale. This hub covers the chips, software, and techniques built to make inference faster and cheaper.
AMD has agreed to acquire Taalas, a Toronto startup that etches AI models directly into silicon, aiming to boost inference speed and efficiency across its chip lineup.
Read more →OpenAI says its GPT-5.6 Sol model autonomously rewrote the GPU kernels behind its inference stack, cutting serving costs 20% and helping fund an 80% price cut for the lightweight Luna tier.
Read more →
A wafer-scale chip turns an entire silicon wafer into one giant processor instead of cutting it into hundreds of separate chips. Here's how it works, and why it's built for fast AI inference.
Read more →
Groq builds ultra-fast AI inference chips called LPUs — a hardware company often confused with xAI's chatbot Grok. Here's what actually separates them, and what Nvidia's $20 billion deal changed.
Read more →Disaggregated inference splits an AI model's two work stages onto separate chips built for each job — a technique now expanding beyond Nvidia to AMD and Cerebras hardware.
Read more →AMD and Cerebras unveiled a joint AI inference system that splits chip work between AMD's Helios servers and Cerebras' wafer-scale engine, claiming up to 5x higher efficiency per watt.
Read more →
AMD unveiled its Instinct MI400 AI chips and Helios rack systems on July 23, claiming 34x faster inference, with OpenAI, Meta and Anthropic on hand.
Read more →Chipmaker Etched is negotiating a funding round that would value it at $20 billion, quadrupling its December price, as investors chase alternatives to Nvidia for AI inference, the Wall Street Journal reported.
Read more →SambaNova closed the first tranche of a $1 billion Series F round led by General Atlantic, and said JPMorganChase will deploy its chips for on-premises AI inference.
Read more →Qualcomm on June 24 unveiled the Dragonfly data center chip family — a 250-core Arm-based CPU and a new inference accelerator — with Meta signed on to deploy the hardware in its server fleet from 2028.
Read more →An AI inference chip is specialized hardware built to run trained AI models quickly and cheaply at scale. Here is why every major tech company is now racing to design its own instead of buying from Nvidia.
Read more →Qualcomm will acquire Modular — creator of the Mojo language and MAX inference engine — in a $3.9 billion all-stock deal aimed at letting developers deploy AI on any hardware without rewrites.
Read more →OpenAI and Broadcom revealed Jalapeño on June 24 — a purpose-built inference chip designed in nine months with AI assistance, targeting deployment by end of 2026.
Read more →