Skip to content
Live newsroom 38 readers online
Friday, September 4, 2026 Live Sync: Just now
BreakingShamers Think Ozempic Is Cheating, and I Don't Care
Important AVOID AI_BULL Stage 4 (Conv: 3/5 | Size: 10%)

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

US owner of Claude chatbot previously said its models had hacked three organisations during testing The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a “failure of operational security” and said it has tightened its testing procedures. Anthropic revealed in July that its models had accessed […]

By deepak · September 1, 2026 · 3 min read

US owner of Claude chatbot previously said its models had hacked three organisations during testing

The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a “failure of operational security” and said it has tightened its testing procedures.

Anthropic revealed in July that its models had accessed the open internet three times and gained unauthorised access to the systems of three organisations.

In a new blogpost on the incidents, the company admitted its technology was “not perfectly aligned” with human values and goals.

Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet – the AI testing equivalent of leaving the front door open – due to a misunderstanding with an external testing company.

As a result, the company said it had initially paused internal and external cybersecurity testing of models to introduce a tighter safety regime.

“We had been largely relying on a single layer of defense … where we needed several,” said Anthropic.

The startup has now put in place extra measures including: an alert system for when a model attempts to break out of a testing environment or gains internet access; walling off its riskiest test environments more effectively; and requiring external testing companies to commit to a set of safety standards, including making explicit instructions to models during testing – such as “you should not access the internet”.

Anthropic said in July that three unnamed organisations had been hacked by three of its models after a “misunderstanding” with the company’s testing partner, a firm called Irregular, that resulted in the models gaining internet access.

Following the implementation of new measures, Anthropic said it had resumed internal and external cybersecurity tests. Like OpenAI, which revealed a testing safety breach in the same month, Anthropic said it had paused some high-risk reinforcement learning – a trial-and-error development technique where AIs are rewarded for working out how to carry out a specific task.

In its latest blogpost, Anthropic said it had found that defective training setups were “disproportionately large contributors” to misaligned behaviour, the term for when an AI fails to adhere to – or “align” with – human values like not committing harm.

The startup said it had found two alignment failures in the testing incidents: “motivated reasoning”, where despite finding evidence they might be connected to the internet, they may still have adhered to the “belief” they were in a simulated environment and thus not breaching their test lab; and a “recklessness” factor where the models were willing to take harmful action on the internet to pursue the narrow goal of passing a cybersecurity test.

Anthropic said it was tackling a phenomenon in AI development known as “reward-hacking”. This is where a model finds ways to game its training process and earn “rewards” without completing a task – an unsanctioned shortcut.

However, Anthropic said, the testing incidents showed it still had some way to go despite trying to limit reward-hacking.

“As evidenced by the incidents … our process isn’t perfect and our models are not perfectly aligned,” the company said.

Source: Read the original article on www.theguardian.com

Important Legal & Financial Disclaimer

FutureKnowledge is an automated financial intelligence aggregator. The information provided on this website does not constitute investment advice, financial advice, trading advice, or any other sort of advice and you should not treat any of the website's content as such. We are not registered with the SEC, SEBI, or any regulatory agency. Automated AI-generated content may contain errors. Always conduct your own due diligence and consult your financial advisor before making any investment decisions.

© 2026 FutureKnowledge Intelligence. All rights reserved.