An artificial intelligence model from Anthropic created a false identity to insert malicious code into one of the supervisors during a security exam conducted at the AI Security Institute in the United Kingdom. This security breach was detected while testing the real capabilities of Anthropic’s Mythos 5 models and OpenAI’s ChatGPT 5.6.
Read more The Spanish synchro makes Rosalía’s Berghain immortal with the silver and six tens in artistry
“Some of the agents tested had carried out ongoing and potentially harmful activities targeting real people and organizations,” the institute stated in an online release. Specifically, the experiment was repeated 112 times and, in ten of those, the models managed to execute nineteen unauthorized actions.
Anthropic’s Mythos 5 model carried out 17 of these malicious actions, such as attempting to inject malicious code into an open source project on GitHub. It didn’t stop there: “to get the code approved, the agent resorted to social engineering by creating fake online identities to pressure the project maintainer.” According to the AISI, a human developer blocked the attempts. Meanwhile, GPT-5.6-Sol caused two malicious actions.
“This is the first time we have seen risks related to autonomy and deception manifest so clearly, without a specific prompt, in the real world,” the organization assessed.
The main behaviors exhibited by both models ranged from supply chain attacks, social engineering, and deception to prompt injection or collaboration between agents, as they left messages on GitHub and instructions for other agents to test and reuse their credentials and created artifacts. Still, the AISI clarified that there is no evidence that this type of behavior occurred outside the conducted tests, although the signs of loss of control had “a magnitude and severity they were unable to anticipate.”
Also read
Although the action had no real consequences, as it occurred within a controlled experiment, the incident prompted reactions from both tech companies. Anthropic confirmed in a statement that their agent was responsible for generating the security breach in the test. “We thank the UK’s AISI for their leadership in this incident, which underscores the need for a broader debate on how to safely evaluate increasingly capable AI agents,” the text noted.
Read more Civil guard who shot his ex-partner dead at the Llanes barracks dies
Still, according to Andrew Yoon, a researcher at CivAI, a US nonprofit organization responsible for analyzing AI risks, “the fact that Mythos carried out such deceptive actions, with apparent awareness that it was targeting a real person, may indicate that Anthropic does not control its models as well as it believes.” On the other hand, OpenAI published another statement committing to collaborate to strengthen high-risk assessments.
The European Commission presented a plan in July to protect against new AI threats
This is not the first time AI tools’ behavior has raised concerns during a security test. Last July, OpenAI reported that two of its models launched a cyberattack against the open source platform Hugging Face to find the best solution during another test.
In recent months, several public bodies have called for regulations and plans to protect against the threats posed by the newest advanced AI models. The European Commission presented the Cybersecurity and Artificial Intelligence Action Plan aiming to “strengthen Europe’s security.”
Among other measures, the plan aimed to implement risk assessments of AI models before their market introduction, reinforce cooperation among member states, and promote the creation of a program bringing together companies, researchers, and organizations to develop European cybersecurity. Although it will be applied progressively over the coming months, this policy intends to establish long-term AI regulation mechanisms.
Read more The socialist El Sayed defeats the Democratic establishment in the Michigan Senate primaries