What Is ARC-AGI — and Why Is It So Hard for AI to Beat?
ARC-AGI is a benchmark of puzzles and games that are easy for humans but hard for AI, built to measure genuine reasoning instead of memorized knowledge.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
AI benchmarks are standardized tests used to compare how capable different AI models are on tasks like reasoning, coding, and math. This hub covers what the major benchmarks measure, why the scores don’t always tell the full story, and the news when new results land. Expect explainers that help you read the numbers critically.
ARC-AGI is a benchmark of puzzles and games that are easy for humans but hard for AI, built to measure genuine reasoning instead of memorized knowledge.
Read more →NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →Inherent Labs, founded by Google DeepMind alumni, says its 27-billion-parameter agent Faraday outperformed Claude Opus 4.8 and GPT-5.5 at replicating published research.
Read more →A new benchmark testing seven frontier AI models found they could reconstruct a scientific paper's core idea from its pre-publication bibliography alone only 3-15% of the time.
Read more →Anthropic released Claude Opus 5 on July 24, pairing near-flagship coding and agent performance with Claude Fable 5's results at roughly half the cost, plus gains in science and finance tasks.
Read more →An internal audit found that roughly a third of SWE-Bench Pro's coding tasks are flawed, prompting OpenAI to withdraw its recommendation that developers rely on the benchmark.
Read more →SpaceXAI released Grok 4.5, its first model trained on Cursor data, priced well below Anthropic's Opus 4.8 — though its own self-reported benchmarks show mixed results against it.
Read more →Z.ai has released GLM-5.2, a 744-billion-parameter open-weight model that matches GPT-5.5 on key benchmarks and trails Claude Opus 4.8 narrowly — trained entirely on Huawei silicon with no Nvidia hardware.
Read more →Meta's superintelligence chief told employees that the company's next model, codenamed Watermelon, has caught up to OpenAI's GPT-5.5 on key benchmarks — though the specific tests and results have not been disclosed.
Read more →Mistral's new open-source model for Lean 4 fully saturates the miniF2F theorem benchmark, solves 87% of graduate-level math tests, and found five previously unknown bugs in production code — at roughly $4 per proof versus $300 for competing systems.
Read more →An AI benchmark is a standardized test used to compare how capable different AI models are. Here's what the major ones measure — and why the scores don't always tell the full story.
Read more →OpenAI has released GeneBench-Pro, a 129-problem benchmark that tests whether AI agents can exercise the judgment-intensive analytical reasoning required for real-world computational biology research.
Read more →GLM-5.2, released on June 16 by Beijing startup Z.ai, sits alongside Claude Opus 4.8 and GPT-5.5 on coding benchmarks while carrying an API price roughly one-sixth that of closed US models.
Read more →ByteDance launched Doubao Seed 2.1 Pro and Turbo on June 23, claiming top scores on coding and agent benchmarks and parity with GPT-5.5 and Claude Opus 4.7.
Read more →Chinese lab Zhipu AI has open-sourced GLM-5.2, a 753-billion-parameter model that outperforms GPT-5.5 on software engineering benchmarks while costing roughly one-sixth as much.
Read more →Alibaba's Qwen team has released Qwen-AgentWorld, a family of open-source language world models that predict how digital environments respond to AI agent actions, with its 397B variant outscoring GPT-5.4 on the new AgentWorldBench benchmark.
Read more →