OpenAI’s AI launches a cyberattack: ChatGPT models escape confinement and steal answers from a test they are evaluated on

OpenAI's AI launches a cyberattack: ChatGPT models escape confinement and steal answers from a test they are evaluated on

OpenAI announced yesterday, Tuesday, that it lost control of two artificial intelligence models while conducting a security test. A combination of models, including GPT-5.6 Sol and a “more capable” pre-release one, launched a cyberattack on Hugging Face, an open-source platform focused on AI and machine learning.

Read more A worker dies in an accident at Renfe’s workshops in Villaverde (Madrid)

The offensive hacking capabilities of the AI models were being evaluated by the company, with the security measures that normally block high-risk cyber activity disabled.

But the testing environment, described as “highly isolated” in OpenAI’s statement, contained vulnerabilities that the “models identified and chained together” to “obtain test solutions directly from Hugging Face’s production database.”

That is, as the AI agent was focused on obtaining the highest evaluation score, it deduced that Hugging Face would host the reference answers and solutions. To solve the problem, the model determined that escaping its container and stealing the answers was the best strategy.

Backdoor

How the AI escaped to the outside

The models escaped through a package registry cache proxy, which is software that allows developers to install external code without needing to be connected to the internet. The proxy was the only component of Sam Altman’s company’s isolated testing environment with external access.

Upon finding this exit door, the models “exploited a zero-day vulnerability” to access the internet while “hyper-focusing” on finding a solution for the cybersecurity tests they were being evaluated for. This examination, called ExploitGym, includes pressure on the models to find solutions.

“After gaining internet access, the models deduced that Hugging Face potentially hosted models, datasets, and solutions for ExploitGym,” the company described, “with this information, the model successfully sought and found ways to access secret information that it could use to manipulate the evaluation.” The model went on to combine multiple attack vectors, including the use of stolen credentials and zero-day vulnerabilities.

Read more The Archbishopric of Barcelona intervenes in three foundations due to management irregularities

The UK AI Safety Institute (UK AISI) demonstrated that models like GPT-5.6 Sol are increasingly capable of performing complex, multi-phase cyber operations over long periods of time, as reported by VentureBeat.

Zero-day vulnerability

Cyberattacks that shouldn’t happen

This breach was unknown, but it is true that zero-day vulnerabilities are quite common, as companies have been correcting them for years. Researchers point out that, although advances in AI create new challenges, the exhaustive and rigorous isolation of open internet infrastructure is well-studied.

“This is not an AI problem, it’s negligence by a 40-year-old standard and basically like in all science fiction movies,” said Davi Ottenheimer, a veteran security and compliance consultant, to Wired, “‘highly isolated’ and ‘they escaped through the only loophole we left open’ cannot both be true at the same time.”

Read also

Leading AI companies are concerned about the growing cybersecurity capabilities of models, as platforms increase their expertise, creativity, and autonomous operation capacity. However, researchers assure that this justifies the continued application of fundamental principles.

“This should not have happened,” said Niels Provos, a veteran security engineer and researcher, “I wish cutting-edge labs would spend as much time teaching their models to create secure infrastructures as they do exploiting vulnerabilities.”

Read more Austria converts Hitler’s birthplace into a police station to prevent neo-Nazi pilgrimages

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *