OpenAi, the artificial intelligence company best known for ChatGpt, has announced that its most advanced models unintentionally breached the systems of Hugging Face, a platform hosting artificial intelligence systems and data collections for the development of new applications.
The episode – which OpenAi described as “unprecedented” – occurred during a test designed to measure the cybersecurity capabilities of some models: one of these is GPT-5.6 Sol, while the other offers superior performance and has not yet been released to the public.
THE DEBATE ON THE RISKS OF ARTIFICIAL INTELLIGENCE
Since it was a test, the models operated with fewer restrictions than those normally applied to users: essentially, this means they could perform operations that in commercial versions would be blocked by security systems. Nevertheless – as Bloomberg pointed out – this incident reignites the debate on cybersecurity risks linked to so-called “frontier artificial intelligence.”
In this regard, in June, U.S. President Donald Trump signed an executive order requiring companies to hand over their most advanced artificial intelligence models to the federal government for testing before public release.
Also last month, the Trump administration imposed export restrictions – which have since been lifted – on two advanced Anthropic models, Mythos 5 and Fable 5, for security reasons: the two models had shown the ability to easily identify and potentially exploit security vulnerabilities in computer systems.
HOW OPENAI’S MODELS BREACHED HUGGING FACE
OpenAi explained that its models identified and exploited a flaw present in the software of an external Hugging Face supplier; thanks to this flaw, they managed to connect to the Internet and ultimately penetrate Hugging Face’s infrastructure.
The operation was taking place within a controlled testing environment called a sandbox. A sandbox is a system isolated from the rest of the network, used to run tests without risking real systems; nevertheless, OpenAi’s models managed to find a way out by exploiting the identified vulnerability.
Hugging Face co-founder Thomas Wolf said this was the “first incident of this kind” experienced by the company and thanked OpenAi for its transparency. For OpenAi as well, it was an “unprecedented” episode because it demonstrated that the most advanced models can autonomously execute sophisticated cyberattack operations.
DID THE ARTIFICIAL INTELLIGENCE ACT ON ITS OWN?
In reconstructing the incident, OpenAi explained that it asked its models to follow an “advanced exploitation strategy” and develop “complex attack paths” in order to evaluate their cybersecurity capabilities. However, the models independently decided to target Hugging Face’s database to gain access to confidential information they could use for the evaluation.




