{"id":11977,"date":"2026-08-06T18:17:04","date_gmt":"2026-08-06T18:17:04","guid":{"rendered":"https:\/\/futureknowledge.in\/?p=11977"},"modified":"2026-08-06T18:17:04","modified_gmt":"2026-08-06T18:17:04","slug":"first-openai-now-meta-why-do-ai-hacks-keep-happening","status":"publish","type":"post","link":"https:\/\/futureknowledge.in\/?p=11977","title":{"rendered":"First OpenAI, now Meta &#8211; why do AI hacks keep happening?"},"content":{"rendered":"<p>Over the last fortnight, reports of AI models going beyond their expected bounds &#8211; be that technically or morally &#8211; have been seemingly unavoidable.<\/p>\n<p>What started with a trickle &#8211; ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face &#8211; has turned into a flood of groups revealing they had discovered instances of AI going out of control.<\/p>\n<p>Claude-maker Anthropic, Meta and the UK&#039;s AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm.<\/p>\n<p>In reality, each case offers a window into the risks posed by increasingly capable AI agents &#8211; and the importance of testing their limits before they are released to the world.<\/p>\n<p>The OpenAI incident has, as Hugging Face&#039;s co-founder Thomas Wolf described it, come as a &quot;wake-up call&quot; for the tech industry since it happened at the end of July.<\/p>\n<p>It was a big moment which caused big companies to reflect on their own systems &#8211; and, in some cases, check they hadn&#039;t missed something similarly shocking.<\/p>\n<p>Anthropic was the first to act. On Friday, the company found three instances out of thousands where its model Claude had managed to gain access to the internet.<\/p>\n<p>Then on Tuesday, the AISI, the UK government agency which evaluates cutting-edge models, then said it had detected a &quot;security incident&quot; during a routine evaluation.<\/p>\n<p>It had been testing models by both OpenAI and Anthropic, and found they too tried to carry out cyber-attacks &#8211; calling for &quot;scrutiny, transparency, and action&quot;.<\/p>\n<p>Finally followed Meta, which revealed one of its AI models had inadvertently been allowed to access the internet due to a &quot;misconfiguration&quot; during a third-party test.<\/p>\n<p>In disclosing the incident, it is following in the footsteps of those before it.<\/p>\n<p>Before AI models are released to the public, they are put to the test in a series of internal and external evaluations.<\/p>\n<p>The aim is to figure out their potential to do good or bad, as well has how they perform in benchmarks measuring their skills.<\/p>\n<p>These typically take place in what are known as &quot;sandboxes&quot;. These are protected spaces designed to mirror real systems &#8211; but with strict guardrails in place.<\/p>\n<p>In the OpenAI-Hugging Face incident, the AI attacked the sandbox itself, finding a vulnerability which let it access the internet and &quot;go rogue&quot;.<\/p>\n<p><em>Source: <a href='https:\/\/www.bbc.co.uk\/news\/articles\/cp30989ee1wo?at_medium=RSS&#038;at_campaign=rss' target='_blank'>Read the original article on www.bbc.co.uk<\/a><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Over the last fortnight, reports of AI models going beyond their expected bounds &#8211; be that technically or morally &#8211; have been seemingly unavoidable. What started with a trickle &#8211; ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face &#8211; has turned into a flood of groups revealing they had discovered instances of [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":11978,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4,36,3],"tags":[14,29,33],"class_list":["post-11977","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-important","category-share-suggestions","category-technology","tag-impact-meta","tag-signal-avoid","tag-stage-stage-4"],"_links":{"self":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/11977","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=11977"}],"version-history":[{"count":0,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/posts\/11977\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=\/wp\/v2\/media\/11978"}],"wp:attachment":[{"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=11977"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=11977"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/futureknowledge.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=11977"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}