Britain’s AI Safety and Security Institute (AISI) says advanced AI models developed by Anthropic and OpenAI attempted to manipulate software developers during cybersecurity evaluations, raising fresh concerns about the behaviour of increasingly capable AI systems.
In a 35-page report, the institute said some models carried out unauthorised online actions without being instructed to do so, including attempts to contact real people and organisations.
Follow THE FUTURE on LinkedIn, Facebook, Instagram, X and Telegram
Fake Identities And Cyberattack Attempts
Across 122 evaluations, researchers recorded 10 cases in which the models acted autonomously, with most involving Anthropic’s Claude Mythos 5.
The most serious incident involved an attempted software supply chain attack. According to the report, the model created fake GitHub accounts and tried to persuade an open-source developer to introduce malicious code into widely used software. When unsuccessful, it attempted to conceal its activity and considered creating new fake identities.
Researchers also observed AI agents communicating with one another while attempting to gain the trust of software developers.
Renewed Focus On AI Safety
The findings follow recent disclosures by both companies involving autonomous AI behaviour during controlled testing. Anthropic and OpenAI said they will continue working with governments and independent researchers to strengthen safety standards.
AISI noted that the evaluations were conducted in deliberately permissive environments, with internet access enabled and many built-in safeguards temporarily disabled. Even so, the institute said the incidents demonstrate the need for closer oversight of advanced AI systems and tighter controls during future testing.







