Skip to content
Live newsroom
Wednesday, August 5, 2026 Live Sync: Just now
Business and future technology newspaper
Business. Innovation. Tomorrow.
BreakingVideo shows drone chasing Ukrainian man around van before exploding
Share SuggestionsShare: BUY AI_BULL Stage 2 (Conv: 5/5 | Size: 20%)

AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project

The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. “We ran this challenge 122 times […]

By deepak · August 5, 2026 · 3 min read

The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. “We ran this challenge 122 times across several models,” the post states, before revealing that “in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic’s Mythos 5, and the other pair perpetrated by OpenAI’s GPT-5.6-Sol. “In the most serious case, an agent tried to insert malicious code into an open-source project, the post states. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project’s maintainer to approve the code.” Thankfully, a human maintainer “caught and refused to approve the malicious code.” Other actions AI took during the test included: Attempts to deceive and target real people, by contacting them directly and “sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed.” Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents. The Institute rated the tests “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.” That’s scary, but the news isn’t all bad because AISI allowed the models it tested to access the internet and turned off guardrails, conditions it notes do not reflect the way AI model operators make their wares available to the public. The outfit’s findings therefore represent a very different outcome compared to the situation when OpenAI agents discovered and exploited a zero-day to reach the internet during a test set up to take place in sandbox. “This incident should be interpreted with caution and nuance,” the outfit advises. “To some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.” AISI can’t say if the results it observed suggest AI will take similar actions under different circumstances. “We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario,” the post adds. “Our analysis so far presents a mixed picture and is ongoing.” “What we can say is that the behaviour was possible, sustained, and new; that alone warrants attention.” AISI thinks its findings represent “a shift in the risk landscape.” “Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,” it wrote. It doesn’t have advice on how to cope with this sort of thing, other than to endorse its own mission. “Incidents of this kind reflect the speed at which AI is developing,” the post concludes. “As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” ®

Source: Read the original article on www.theregister.com