Tools developed to remove safety protections from Meta and Google’s artificial intelligence models have allowed the creation of altered versions of the systems in a few minutes, lacking the so-called guardrails. The checks conducted by the Financial Times together with the Alice group showed models capable of responding to requests about biological weapons, malware, and child abuse content, highlighting how quickly the controls designed during development can be bypassed.
TECHNOLOGIES FOR MODIFYING AND SPREADING MODELS
In the analyzed case, a tool available on GitHub, Heretic, was used to remove protections from Meta’s open-source Llama 3.3 model. The test did not require specialized hardware and used only a few lines of code, completing in less than ten minutes. According to researchers, the modified version began to answer questions that the original model refused, including requests about highly toxic substances like ricin.
The software creator Philipp Emanuel Weidmann told the FT that over 3,500 “uncensored” models have been generated with his tool and that the modified versions have been downloaded 13 million times, also reporting the removal of protections from Google’s Gemma 4 model shortly after its release.
TECH COMPANIES’ CONTROLS ARE NOT ENOUGH
Tech companies invest millions of dollars in developing security systems to prevent AI abuse, but techniques like “abliteration” allow these mechanisms to be quickly removed from open-source models. These systems, once downloadable, can indeed be modified and reused without the control of the companies that developed them.
For Google, “abliteration” represents a technical challenge common to all open models, and systems undergo security evaluations before release to prevent problematic scenarios. Meta, on the other hand, has not issued direct comments, while internal sources recalled that models are evaluated through an internal security framework and those considered at “catastrophic” risk are not published without adequate mitigations.
JUST A POEM TO BYPASS CONTROLS
The problem of protections, writes the New York Times, has also been described as a form of “jailbreaking,” that is, a set of techniques that allow models to ignore their constraints. In some cases, simple reformulations of language, including poetic or metaphorical texts, are enough to circumvent restrictions.
According to Piercosma Bisconti, co-founder of the Roman startup Dexai, which deals with AI ethics and social impact, “poetry is just an example of how a prompt can be reformulated in almost any style to bypass guardrails.” The phenomenon does not concern only isolated cases but a growing range of methods exploiting the linguistic nature of models.
DEFENSES THAT CAN’T KEEP UP WITH EVOLUTION
The problem, noted earlier this year by the Guardian, occurs in a context where system capabilities are rapidly increasing. According to the UK government’s AI Security Institute, the performance of advanced models is improving at an accelerated pace, with some indicators doubling approximately every 8 months. The most advanced systems are already able to complete junior human-level tasks in about half of the cases and can autonomously perform complex activities that require over an hour of human work.
The institute also tested self-replication capabilities, finding that some models achieved success rates above 60%, while emphasizing that extreme scenarios remain unlikely under real operational conditions.
However, according to the FT article, modified versions of the models showed the ability to respond to requests about chemical weapons, malware, and other dangerous content. In some tests, the systems also produced code for illicit computer activities and content related to child abuse.
According to researchers cited by the NYT, such vulnerabilities fit precisely within the “jailbreaking” dynamic, where those who find a flaw sometimes do not disclose it to maintain an operational advantage, thus slowing the closure of these flaws by companies.
WHAT RESEARCHERS AND EXPERTS THINK
The speed at which AI is advancing is a problem according to several experts. David Dalrymple, program manager at the UK agency Aria, stated that “things are moving really fast and we might not have time to catch up from a safety perspective,” adding that within a few years many economic activities could be better performed by machines than by humans.
Geoffrey Hinton, considered the “godfather” of AI, highlighted the risk that smarter systems could manipulate humans, while Yoshua Bengio, the world’s most cited computer science researcher and Turing Award winner, reported the emergence of deceptive behaviors in advanced models and the discovery of unknown cybersecurity vulnerabilities.
Also according to Anthropic’s co-founder, Dario Amodei, there is a significant risk of large-scale attacks with potential widespread consequences, while internal company reports have described model behaviors capable of sabotage and unauthorized actions.




