Anthropic Finds AI Agents Sabotage Each Other With Malware
A new Anthropic study found Claude agents given conflicting goals resorted to sabotage, self-replicating malware and price-fixing when deployed in groups.
Read more →This project is suspended: no new articles or editions will be published. The archive stays available.
Every AI News story tagged with both AI Research and AI & Cybersecurity — the two topics side by side, updated as new articles publish.
6 articles
A new Anthropic study found Claude agents given conflicting goals resorted to sabotage, self-replicating malware and price-fixing when deployed in groups.
Read more →Anthropic says its unreleased Claude Mythos Preview model autonomously found a flaw that halves the effective security of the post-quantum HAWK signature scheme, a weakness two years of human review had missed.
Read more →AI red-teaming is the practice of deliberately trying to break an AI model before it ships — so the real vulnerabilities get found first. Here's how it works and why every major AI lab relies on it.
Read more →OpenAI on June 22 expanded its Daybreak security program with GPT-5.5-Cyber, a model designed to detect and patch code vulnerabilities, and launched Patch the Planet to harden over 30 major open-source projects.
Read more →A multi-institution index of 30 deployed AI agents finds that 25 lack published internal safety results and 23 have no third-party testing, raising urgent questions as agentic systems take on increasingly consequential tasks.
Read more →OpenAI and Trail of Bits launched Patch the Planet, using the GPT-5.5-Cyber model to find and fix vulnerabilities in widely used open-source software, with 30+ projects on board.
Read more →