NVIDIA's AVO Agent Aces ARC-AGI-3 With a Perfect Score
NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Every AI News story tagged with both AI Benchmarks and News — the two topics side by side, updated as new articles publish.
14 articles
NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →Inherent Labs, founded by Google DeepMind alumni, says its 27-billion-parameter agent Faraday outperformed Claude Opus 4.8 and GPT-5.5 at replicating published research.
Read more →A new benchmark testing seven frontier AI models found they could reconstruct a scientific paper's core idea from its pre-publication bibliography alone only 3-15% of the time.
Read more →Anthropic released Claude Opus 5 on July 24, pairing near-flagship coding and agent performance with Claude Fable 5's results at roughly half the cost, plus gains in science and finance tasks.
Read more →An internal audit found that roughly a third of SWE-Bench Pro's coding tasks are flawed, prompting OpenAI to withdraw its recommendation that developers rely on the benchmark.
Read more →SpaceXAI released Grok 4.5, its first model trained on Cursor data, priced well below Anthropic's Opus 4.8 — though its own self-reported benchmarks show mixed results against it.
Read more →Z.ai has released GLM-5.2, a 744-billion-parameter open-weight model that matches GPT-5.5 on key benchmarks and trails Claude Opus 4.8 narrowly — trained entirely on Huawei silicon with no Nvidia hardware.
Read more →Meta's superintelligence chief told employees that the company's next model, codenamed Watermelon, has caught up to OpenAI's GPT-5.5 on key benchmarks — though the specific tests and results have not been disclosed.
Read more →Mistral's new open-source model for Lean 4 fully saturates the miniF2F theorem benchmark, solves 87% of graduate-level math tests, and found five previously unknown bugs in production code — at roughly $4 per proof versus $300 for competing systems.
Read more →OpenAI has released GeneBench-Pro, a 129-problem benchmark that tests whether AI agents can exercise the judgment-intensive analytical reasoning required for real-world computational biology research.
Read more →GLM-5.2, released on June 16 by Beijing startup Z.ai, sits alongside Claude Opus 4.8 and GPT-5.5 on coding benchmarks while carrying an API price roughly one-sixth that of closed US models.
Read more →ByteDance launched Doubao Seed 2.1 Pro and Turbo on June 23, claiming top scores on coding and agent benchmarks and parity with GPT-5.5 and Claude Opus 4.7.
Read more →Chinese lab Zhipu AI has open-sourced GLM-5.2, a 753-billion-parameter model that outperforms GPT-5.5 on software engineering benchmarks while costing roughly one-sixth as much.
Read more →Alibaba's Qwen team has released Qwen-AgentWorld, a family of open-source language world models that predict how digital environments respond to AI agent actions, with its 397B variant outscoring GPT-5.4 on the new AgentWorldBench benchmark.
Read more →