Skip to content

Has OpenAI really lost control of its AI? Experts’ doubts

OpenAI described the now-famous hack carried out autonomously by one of its AI models as an "unprecedented incident." However, industry experts point out that it was actually an experiment turned into sensationalistic storytelling to emphasize the offensive capabilities of its systems.

We have all now read about the case of the OpenAI agent that allegedly bypassed Hugging Face’s defenses during a security test, breaking out of the protected sandbox and managing to breach one of the most important platforms in the entire AI ecosystem.

The company calls it an “unprecedented incident.” But some independent reconstructions suggest that the reality might be very different: what was presented to the media as a dramatic and unpredictable event looks much more like an experiment conducted with the brakes deliberately loosened. A test gone wrong that OpenAI turned into a compelling narrative, mainly useful to emphasize the offensive capabilities of its models.

The article in Sole 24 Ore by industry expert journalist Luca De Biase, and the analyses conducted by academic Luciano Floridi, Professor and Founding Director of the Digital Ethics Center at Yale, and shared on LinkedIn and on Medium, converge on a suspicion: OpenAI did not tell the media the whole truth, or at least did not tell it in the most transparent way.

A statement that sounds like a marketing launch

OpenAI described the episode as “an unprecedented incident involving cutting-edge cyber capabilities.” A tone more like a product announcement than a report on an internal test gone out of control.

As De Biase observes, the agent “was portrayed as so skilled that it crossed the boundaries of the sandbox, that is, the protected computing environment it was not supposed to leave. But it remains to be understood whether this happened because the agent was very powerful or because the sandbox boundaries were too weak.”

In essence, we don’t know if we witnessed an unpredictable AI exploit or simply a demonstration of insufficient containment measures.

The kitchen robot without a lid

For his part, Floridi is concerned with dismantling the rhetoric propagated by OpenAI using a sharp metaphor.

On LinkedIn, the John K. Castle Professor in Practice of Cognitive Sciences at Yale writes: “Deliberately loosen the lid of the kitchen robot, set it to maximum, and issue a solemn statement about the soup on the ceiling. That’s more or less what OpenAI did this week.”

According to this version, the company ran the most advanced models with reduced safety filters and disabled protections, challenged them with a hacking task, and then expressed surprise when they broke out of the test box.

In the Medium article, Floridi retells the same metaphor in another way, driving the point home: “The appliance – Floridi emphasizes – simply operated at the speed you selected with the lid you chose not to secure. The blade did not conspire against you. It simply spun.”

In short, there was no rebellion nor spark of consciousness. Just a system from which the guardrails were removed to measure its maximum performance.

Anthropomorphization as an old marketing strategy

For decades, the tech industry, AI included, has resorted to a narrative device that De Biase calls “anthropomorphization”: a technique used “to attract attention to this technology and fascinate the public.”

Floridi insists on another aspect of the problem, the linguistic one. Terms like “escaped,” “gone rogue,” or “hyper-focused” shift responsibility from human choice to the machine. But in reality, the professor reiterates the accusation, the “culprit lies with the board of directors who adjusted the knob, not the blade.”

The technical lessons that matter

Beyond the alleged staging, the episode offers important technical insights. An autonomous agent completed a full intrusion chain: it discovered an unknown vulnerability, escalated its privileges, moved through systems, reached the internet, and breached a third party to achieve its goal.

As Floridi acknowledges, without giving up irony, “from an engineering perspective, a lot happened, and it matters much more than the ghost story.”

The most interesting lesson concerns the asymmetry between attack and defense. When Hugging Face’s technicians tried to use cutting-edge models to analyze the intrusion, these refused to proceed, blocked by their own security policies. The team had to resort to an open-weight model running locally. The attacker had no constraints; the defenders did.

Why this narrative benefits OpenAI

Portraying the event as an unpredictable incident allows a company like OpenAI to minimize its responsibilities while simultaneously highlighting the offensive capabilities of its systems, the very ones it sells through programs like Daybreak.

De Biase concludes that it is time to “take the measure of this communication style,” which has long played on the ambiguity between extraordinary power and out-of-control danger.

Back To Top