Alibaba's Qwen3.8-Max-0902 Tops Code Arena, Beats Claude Opus 5
Alibaba's refreshed Qwen3.8-Max-0902 model has taken the top spot on Arena.ai's Code Arena coding leaderboard, narrowly outscoring Anthropic's Claude Opus 5.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Every AI News story tagged with both Large Language Models and News — the two topics side by side, updated as new articles publish.
14 articles
Alibaba's refreshed Qwen3.8-Max-0902 model has taken the top spot on Arena.ai's Code Arena coding leaderboard, narrowly outscoring Anthropic's Claude Opus 5.
Read more →Anthropic released Claude Fable 5.1 and Mythos 5.1 on September 1, cutting cache-read costs 75% and posting sharp benchmark gains over Fable 5.
Read more →Tencent's Hunyuan team open-sourced Hy4 preview, a 770-billion-parameter model with a 1-million-token context window, joining China's fast-moving open-weight AI race.
Read more →Chinese AI firm Z.ai confirmed its anonymous "Ox Alpha" model was GLM-5.3-Flash, releasing the open-weight, multimodal model under an MIT license.
Read more →Thomson Reuters launched Thomson, a proprietary AI model trained on its legal and news archives and built atop Alibaba's open-weight Qwen model, first powering CoCounsel Legal's document review.
Read more →A new benchmark testing seven frontier AI models found they could reconstruct a scientific paper's core idea from its pre-publication bibliography alone only 3-15% of the time.
Read more →Nvidia has released Nemotron 3.5 Lightning, an open 30-billion-parameter AI agent model, plus NeMo Switchyard, a router that cuts task costs to about a third of using Opus 4.8 alone.
Read more →DeepSeek will begin peak and off-peak API pricing on August 16, raising some DeepSeek-V4-Pro rates by up to 12 times as demand for its models surges.
Read more →DeepSeek's flagship model exited preview on August 12 with sharp gains on coding and agent benchmarks, but only middling scores on general reasoning tests, and no formal announcement.
Read more →OpenAI released GPT-5.6-Cyber, a restricted-access model for vetted security researchers that it rates 'High' for cyber capability — one step below its 'Critical' threshold.
Read more →Thinking Machines Lab's second open-weight model needs a quarter of its predecessor's active parameters yet edges past it on reasoning and coding benchmarks.
Read more →OpenAI says its GPT-5.6 Sol model autonomously rewrote the GPU kernels behind its inference stack, cutting serving costs 20% and helping fund an 80% price cut for the lightweight Luna tier.
Read more →Alibaba previewed Qwen3.8-Max, a 2.4-trillion-parameter multimodal model it calls second only to Anthropic's Claude, though no benchmarks have been published.
Read more →Fireworks AI closed a $1.5 billion Series D at a $17.5 billion valuation, saying its infrastructure now serves over 40 trillion tokens a day for specialized AI models.
Read more →