Thinking Machines Releases Inkling-Small, Outperforms Larger Model
Thinking Machines Lab's second open-weight model needs a quarter of its predecessor's active parameters yet edges past it on reasoning and coding benchmarks.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Mixture of Experts (MoE) is an AI model design that holds hundreds of billions of parameters but activates only a small fraction of them for each query. This lets very large models run quickly and cheaply, and it underpins many of today’s frontier systems. This hub explains how MoE works and why it matters.
Thinking Machines Lab's second open-weight model needs a quarter of its predecessor's active parameters yet edges past it on reasoning and coding benchmarks.
Read more →DeepSeek's official V4-Flash-0731 release outperforms its larger V4-Pro model on coding and agent benchmarks, despite activating far fewer parameters per token.
Read more →Ant Group's Ling-3.0-Flash matches models several times its size on reasoning and coding benchmarks using just 5.1 billion active parameters, the company says.
Read more →Beijing-based Moonshot AI has published the full open weights for Kimi K3, a 2.8-trillion-parameter model it calls the largest open model released to date.
Read more →Poolside's new 118-billion-parameter Laguna S 2.1 matches or beats far larger rivals on agentic coding benchmarks and ships under an open license.
Read more →China's Moonshot AI released Kimi K3, a 2.8-trillion-parameter open model with a 1-million-token context window — a system the company calls the world's first open 3-trillion-class AI model.
Read more →Mixture of Experts (MoE) lets an AI model hold hundreds of billions of parameters while activating only a fraction of them per query — the main reason some huge models are also cheap to run.
Read more →