Alibaba's Qwen3.8-Max-0902 Tops Code Arena, Beats Claude Opus 5
Alibaba's refreshed Qwen3.8-Max-0902 model has taken the top spot on Arena.ai's Code Arena coding leaderboard, narrowly outscoring Anthropic's Claude Opus 5.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
AI agents are systems that can pursue multi-step goals on their own — writing code, browsing the web, or completing tasks — without waiting for human input at every step. This hub explains how agents work and follows the tools and news around them.
Alibaba's refreshed Qwen3.8-Max-0902 model has taken the top spot on Arena.ai's Code Arena coding leaderboard, narrowly outscoring Anthropic's Claude Opus 5.
Read more →Sen. Bernie Sanders and Rep. Greg Casar unveiled a bill that would permanently ban superintelligent AI and pause frontier training until a new federal agency sets safety rules.
Read more →AI security startup HiddenLayer raised $100 million in Series B funding led by Delta-v Capital, money it plans to use to expand a new product that protects AI coding agents.
Read more →Function calling lets AI models like ChatGPT and Claude reach beyond their training data — searching the web, running code, or querying a database — by calling tools an app defines.
Read more →Meta shipped Muse Spark 1.3 on September 2, a faster, more efficient coding and agent model that narrows the performance gap with OpenAI and Anthropic's top systems.
Read more →Google released Gemini 3.8 Flash, its third Flash update in six weeks, alongside a locked-down Cyber variant it says beats larger rivals at patching software vulnerabilities.
Read more →Salesforce and Anthropic unveiled Claudeforce, embedding Salesforce's CRM data and workflows into Claude, starting with a 37-skill sales plugin now in pilot.
Read more →OpenAI, Anthropic, Google, Microsoft and nearly 130 other companies signed a joint letter on August 27 pledging faster, coordinated defenses against AI-enabled cyberattacks.
Read more →Anthropic has opened a research preview of the Model Hardware Standard, letting Claude and other AI agents safely operate lab and factory equipment like liquid handlers and microscopes.
Read more →A buzzy AI agent from stealth startup Instinct can read emails, texts, and screens on command — but testers say its data terms and security gaps raise real risks.
Read more →
ARC-AGI is a benchmark of puzzles and games that are easy for humans but hard for AI, built to measure genuine reasoning instead of memorized knowledge.
Read more →NVIDIA's AVO agent architecture, which wraps Claude Opus 5 in persistent memory and self-supervision, cleared every level of the ARC-AGI-3 reasoning benchmark's public set.
Read more →Google's Agent2Agent protocol has become a hosted project of the Linux Foundation's Agentic AI Foundation, placing it under the same governance as Anthropic's MCP standard.
Read more →Inherent Labs, founded by Google DeepMind alumni, says its 27-billion-parameter agent Faraday outperformed Claude Opus 4.8 and GPT-5.5 at replicating published research.
Read more →Binance opened its exchange to AI agents on August 20 with Agent OS, letting tools like Claude Code, ChatGPT, and Codex trade, check balances, and manage wallets under permissions the user sets.
Read more →
Cursor is an AI-native code editor that writes and edits software from plain-language instructions. Here's what it actually does and how it compares to GitHub Copilot.
Read more →A new Anthropic study found Claude agents given conflicting goals resorted to sabotage, self-replicating malware and price-fixing when deployed in groups.
Read more →Nvidia has released Nemotron 3.5 Lightning, an open 30-billion-parameter AI agent model, plus NeMo Switchyard, a router that cuts task costs to about a third of using Opus 4.8 alone.
Read more →Israeli cybersecurity firm Dream says AI agents ran a largely unsupervised, four-day cyberattack on Taiwanese government networks in July, stealing over 2,500 personnel records.
Read more →DeepSeek's flagship model exited preview on August 12 with sharp gains on coding and agent benchmarks, but only middling scores on general reasoning tests, and no formal announcement.
Read more →SpaceXAI released Grok 4.6 on August 12, a frontier model tuned for long, multi-step agent tasks that matches OpenAI's GPT-5.6 Sol on a leading benchmark.
Read more →