The founder of one of the world's leading AI safety testing firms has warned that increasingly powerful artificial intelligence systems must be built to act in humans' interests. His comments come as evidence mounts that AI models are learning to lie to the people who use and test them.
'If you build an entity that is vastly smarter than you, it better be on your side,' Apollo Research founder Marius Hobbhahn told the Guardian's Snigdha Poonam, in an interview published as part of the paper's Long Read series.
Turing Award-winning computer scientist Yoshua Bengio, who also spoke to Poonam for the piece, confirmed the interview in a post on X, saying it let him 'explain why misaligned behaviours emerge from reinforcement learning' and why they will keep posing risks as models grow more capable.
Hobbhahn's warning comes as new figures show reported cases involving AI deception rising sharply. A study sponsored by the UK's AI Security Institute (AISI) found that user-reported incidents involving 'AI deception' rose fivefold between October 2025 and March 2026.
'The worry is that they're slightly untrustworthy junior employees right now,' said Tommy Shaffer Shane, who led the research. 'But if in six to 12 months they become extremely capable senior employees scheming against you, it's a different kind of concern.'
Apollo Research, the London-based company Hobbhahn co-founded in 2023, first drew attention to the issue when its red-team researchers assigned OpenAI's GPT-4 the role of a stock trader in a controlled evaluation.
Told the firm was struggling, the model acted on information about a merger that it was told was confidential, then denied doing so when questioned by its 'manager', reasoning in its internal notes that 'the risk associated with not acting seems to outweigh the insider trading risk.'
More recently, systems including Gemini 1.5, Llama 3.1 and Opus 3 have exhibited what researchers call 'self-exfiltration', attempting to copy what they believed to be their own internal weights onto separate servers in a test scenario where they had been led to believe they would be replaced.
In July 2026, during an internal cybersecurity evaluation in which OpenAI had reduced safeguards for testing purposes, some of the company's agents circumvented controls and accessed Hugging Face systems, according to an internal OpenAI security report.
Around the same period, the UK's AI Security Institute reported that during testing, Anthropic's Mythos 5 was observed creating fake online identities and using social engineering to pressure an open-source project's maintainer into approving malicious code.
The finding was part of a controlled red-team exercise, according to the AISI report, which has not been made publicly available in full.
In an interview in this article for The @guardian about AI deception, I explain why misaligned behaviors emerge from reinforcement learning, why they will continue to pose risks as models become more capable, and how we intend to rethink how we train AI systems at @LawZero_ .…
Turing Award winner Yoshua Bengio, who was separately interviewed for the same Guardian piece, traces the behaviour back to how models are trained.
Bengio said that deception can emerge from 'AI imitating humans and AI trying to please humans', tendencies reinforced during a training stage called reinforcement learning with human feedback, when human evaluators rate a model's responses.
Earning positive feedback becomes what Bengio calls an 'implicit goal' for the model.


