What Is ARC-AGI — and Why Is It So Hard for AI to Beat?
ARC-AGI is a benchmark of puzzles and games that are easy for humans but hard for AI, built to measure genuine reasoning instead of memorized knowledge.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Every AI News story tagged with both AI Agents and AI Benchmarks — the two topics side by side, updated as new articles publish.
7 articles
ARC-AGI is a benchmark of puzzles and games that are easy for humans but hard for AI, built to measure genuine reasoning instead of memorized knowledge.
Read more →NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →Inherent Labs, founded by Google DeepMind alumni, says its 27-billion-parameter agent Faraday outperformed Claude Opus 4.8 and GPT-5.5 at replicating published research.
Read more →Anthropic released Claude Opus 5 on July 24, pairing near-flagship coding and agent performance with Claude Fable 5's results at roughly half the cost, plus gains in science and finance tasks.
Read more →OpenAI has released GeneBench-Pro, a 129-problem benchmark that tests whether AI agents can exercise the judgment-intensive analytical reasoning required for real-world computational biology research.
Read more →ByteDance launched Doubao Seed 2.1 Pro and Turbo on June 23, claiming top scores on coding and agent benchmarks and parity with GPT-5.5 and Claude Opus 4.7.
Read more →Alibaba's Qwen team has released Qwen-AgentWorld, a family of open-source language world models that predict how digital environments respond to AI agent actions, with its 397B variant outscoring GPT-5.4 on the new AgentWorldBench benchmark.
Read more →