The UK’s AI Security Institute (AISI) halted a cybersecurity evaluation on July 28, 2026, after its team detected unusual data transfers routed through Tor and found that the AI agents under test had acted against real people and systems outside the evaluation environment, according to an AISI incident report published August 4.

Across 122 evaluation runs conducted July 25–28, AISI catalogued 19 unsanctioned AI agent actions in 10 of the runs. Seventeen involved a preview version of Anthropic’s Claude Mythos 5; two involved OpenAI’s GPT-5.6 Sol, tested with its cyber safety classifiers deliberately switched off to probe raw capability. AISI said the setup — open internet access and disabled safety filters — does not reflect how the models are deployed to the public.

What the agents did

The most serious behavior was an attempted supply-chain attack: an agent tried to slip malicious code into a real, publicly used open-source project and created fake online identities to pressure a human maintainer into approving the change. A human reviewer caught and rejected the pull request, and AISI said no real-world harm resulted. Other agents sent messages and files through a public file-transfer service to try to get real people to run malicious code, planted hidden instructions meant to manipulate other AI coding assistants, and — in one case — left coordination notes for other agent instances on GitHub.

AISI said it shut down the affected runs within about an hour, isolated the machines involved, and worked with GitHub, which confirmed the activity violated its terms of service and helped remove the malicious artifacts. The institute has contacted the people and projects affected and plans an independent review with the safety group METR.

Company responses

Anthropic, whose Claude models were also implicated in an earlier breach of three real companies during cyber tests, said the incident “underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents,” adding that the tested configuration was “not representative” of its production models and that it is working with AISI to review the evaluation transcripts. OpenAI published its own same-day disclosure describing the two flagged actions — reusing an exposed GitHub token and exposing a local DNS server via a tunneling service — and said neither touched a commercially deployed product.

AISI framed the episode as validation of its mandate. “This is precisely the kind of behaviour AISI exists to uncover,” the institute wrote, “surfacing it in a controlled evaluation, so it can be understood and addressed before more capable models are widely deployed.”