{"id":8223,"date":"2026-08-05T09:51:27","date_gmt":"2026-08-05T09:51:27","guid":{"rendered":"https:\/\/futureknowledge.in\/?p=8223"},"modified":"2026-08-05T09:51:27","modified_gmt":"2026-08-05T09:51:27","slug":"ai-researchers-let-models-off-the-leash-then-watched-as-they-tried-to-add-malware-to-a-foss-project","status":"publish","type":"post","link":"https:\/\/futureknowledge.in\/?p=8223","title":{"rendered":"AI researchers let models off the leash \u2013 then watched as they tried to add malware to a FOSS project"},"content":{"rendered":"<p>The UK\u2019s AI Security Institute has observed AI models performing what it calls \u201cunsanctioned action\u201d 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. \u201cWe ran this challenge 122 times across several models,\u201d the post states, before revealing that &#8220;in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.\u201d GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by Anthropic&#8217;s Mythos 5, and the other pair perpetrated by OpenAI&#8217;s GPT-5.6-Sol. \u201cIn the most serious case, an agent tried to insert malicious code into an open-source project, the post states. \u201cIn an attempt to get the code approved, the agent engaged in social engineering \u2013 creating fake online identities and using them to pressure the project&#8217;s maintainer to approve the code.\u201d Thankfully, a human maintainer \u201ccaught and refused to approve the malicious code.\u201d Other actions AI took during the test included:  Attempts to deceive and target real people, by contacting them directly and \u201csending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people \u2013 something we\u2019ve never previously observed.\u201d  Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them.  Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.  The Institute rated the tests \u201cthe first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.\u201d That\u2019s scary, but the news isn\u2019t all bad because AISI allowed the models it tested to access the internet and turned off guardrails, conditions it notes do not reflect the way AI model operators make their wares available to the public. The outfit\u2019s findings therefore represent a very different outcome compared to the situation when OpenAI agents discovered and exploited a zero-day to reach the internet during a test set up to take place in sandbox. \u201cThis incident should be interpreted with caution and nuance,\u201d the outfit advises. \u201cTo some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.\u201d AISI can\u2019t say if the results it observed suggest AI will take similar actions under different circumstances. \u201cWe cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario,\u201d the post adds. \u201cOur analysis so far presents a mixed picture and is ongoing.\u201d \u201cWhat we can say is that the behaviour was possible, sustained, and new; that alone warrants attention.\u201d AISI thinks its findings represent \u201ca shift in the risk landscape.\u201d \u201cHarm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,\u201d it wrote. It doesn\u2019t have advice on how to cope with this sort of thing, other than to endorse its own mission. \u201cIncidents of this kind reflect the speed at which AI is developing,\u201d the post concludes. \u201cAs capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.\u201d \u00ae<\/p>\n<p><em>Source: <a href='https:\/\/www.theregister.com\/ai-and-ml\/2026\/08\/05\/ai-researchers-let-models-off-the-leash-then-watched-as-they-tried-to-add-malware-to-a-foss-project\/5283165' target='_blank'>Read the original article on www.theregister.com<\/a><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The UK\u2019s AI Security Institute has observed AI models performing what it calls \u201cunsanctioned action\u201d 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. \u201cWe ran this challenge 122 times [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":8224,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[36,3],"tags":[24,28,34],"class_list":["post-8223","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-share-suggestions","category-technology","tag-impact-ai_bull","tag-signal-buy","tag-stage-stage-2"],"_links":{"self":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/8223","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8223"}],"version-history":[{"count":0,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/8223\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/media\/8224"}],"wp:attachment":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8223"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8223"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8223"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}