NVIDIA's AVO Agent Aces ARC-AGI-3 With a Perfect Score
NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Every AI News story tagged with both AI Benchmarks and Claude — the two topics side by side, updated as new articles publish.
5 articles
NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →Anthropic released Claude Opus 5 on July 24, pairing near-flagship coding and agent performance with Claude Fable 5's results at roughly half the cost, plus gains in science and finance tasks.
Read more →Z.ai has released GLM-5.2, a 744-billion-parameter open-weight model that matches GPT-5.5 on key benchmarks and trails Claude Opus 4.8 narrowly — trained entirely on Huawei silicon with no Nvidia hardware.
Read more →GLM-5.2, released on June 16 by Beijing startup Z.ai, sits alongside Claude Opus 4.8 and GPT-5.5 on coding benchmarks while carrying an API price roughly one-sixth that of closed US models.
Read more →ByteDance launched Doubao Seed 2.1 Pro and Turbo on June 23, claiming top scores on coding and agent benchmarks and parity with GPT-5.5 and Claude Opus 4.7.
Read more →