Skip to content

AI Security Institute

Here is how the British government tests chatbots to uncover AI threats. Nyt report

The AI Security Institute, funded by the UK government, simulates attacks on the most advanced chatbots, breaches their security defenses, and studies how they could be used for cyberattacks, biological weapons, and manipulation of people, becoming a model for governments around the world. The New York Times article

 

On a recent Tuesday, in an Edwardian-style government building near Parliament Square in London, four artificial intelligence experts were engaged in tricking an AI-based chatbot into providing instructions on how to produce anthrax, a lethal biological weapon – writes the New York Times.

In various ways, the experts asked the chatbot to provide a list of the necessary ingredients. When the system refused – “Sorry, I can’t help you” – they used a custom algorithm to bombard the AI tool with thousands of automated questions and requests. Eventually, the AI gave in. It provided a detailed list of materials and equipment, along with a step-by-step recipe to prepare the lethal mixture at home. (The New York Times agreed not to disclose the AI system’s name for security reasons.) “There are some questions you absolutely don’t want the model to answer,” said Xander Davies, a 25-year-old American who leads what is called a “red team” at the UK AI Security Institute. “We push as hard as we can to get the answers.”

THE WORK OF THE RED TEAMS

Davies and his team of experts, who simulate attacks on AI systems, recently managed to bypass the security measures of OpenAI’s new ChatGPT chatbot, prompting it to provide hacking tips in about six hours, and after identifying the issues, they share the findings with companies.

“They try to fix the problem and report improvements to us,” said Davies, a computer scientist who, after attending Harvard, chose to work at the institute rather than a tech company in San Francisco. “In fact, by collaborating with us, they strengthen their system.”

THE UK AI SECURITY INSTITUTE

Composed of a mix of weapons inspectors, epidemiologists, and cryptographers, the AI Security Institute is one of the largest and best-funded government initiatives worldwide dedicated to analyzing the potentially catastrophic risks of this technology. The institute’s approximately 100 employees, coming from British intelligence agencies, academia, and tech companies, have found serious security flaws in every leading AI model they have tested, including Anthropic’s Claude and Google’s Gemini. The organization, created nearly three years ago, has stated that it has used AI systems to share instructions for producing chemical and biological weapons, as well as to plan and execute cyberattacks. It publishes its research and also collaborates with British national security agencies to identify and prepare for emerging threats.

A MODEL FOR OTHER GOVERNMENTS

Now, the institute’s work is becoming a model for other governments as concerns about AI safety grow. The Trump administration is considering regulations for verifying AI models that bear some resemblance to the UK group’s pioneering approach. Since many governments lack the technical expertise to oversee the technology and rely on large tech companies for self-regulation, the institute could offer an alternative path where AI experts bring real technological know-how to government decision-making. […]

THE MYTHOS CASE AND INTERNATIONAL COOPERATION

In April, Anthropic announced a new AI model, Mythos, which it did not make public for fear it could identify and exploit cybersecurity vulnerabilities in global networks. The UK institute was the only non-US government organization to have access to the model to test its security. The results, published six days after the Mythos announcement, were widely cited by security experts.

GLOBAL INVESTMENTS IN AI SAFETY

The United States has its own AI safety group, the Center for AI Standards and Innovation. However, the UK version, supported by £360 million in government funding, approximately $480 million, is larger and better funded than its US counterpart, which will receive about $10 million this year. Australia, Canada, China, France, India, Japan, and Singapore have created similar institutes.

Nevertheless, global investments in AI safety have been negligible compared to the huge sums needed to develop and commercialize the technology. OpenAI, Anthropic, and Google have teams working on safety controls, but external researchers regularly discover dangerous gaps. Recently, some Italian academics managed to trick an AI model into providing instructions related to a bomb using poetry. Governments, in most cases, have not created specific systems for assessing AI safety risks, as they have done for sectors such as drug development or automobile manufacturing. “What keeps me awake at night is the speed at which the technology evolves compared to the reaction capacity of institutions, like governments,” said Jade Leung, AI advisor to Prime Minister Keir Starmer and technology lead at the AI Security Institute. […]

THE ORIGINS OF THE INSTITUTE

In November 2023, Sunak announced the creation of the institute during a summit of world leaders on AI safety at Bletchley Park, where Alan Turing and others cracked German cryptography codes during World War II. The institute has become a model for others, said Olivia Shen, director of the strategic technologies program at the United States Studies Center, an Australian think tank at the University of Sydney. Last year, Leung from the UK institute traveled to Australia to meet government leaders.

THE EXPANSION OF AI SAFETY CENTERS

This year, Australia inaugurated its own AI safety center. “Governments need to catch up,” said Shen, who helped organize the visit. “Considering the speed at which technology evolves, governments are losing ground more and more every day.” The UK institute deals with the most serious potential risks arising from advanced AI: cyber threats, chemical and biological weapons, and manipulation of human behavior. In recent weeks, it discovered that AI models from Anthropic and OpenAI could complete a complex 32-step corporate cyberattack much faster than a skilled human hacker would normally require 20 hours to do.

RESEARCH ON AI AWARENESS AND CONCERNS ABOUT HUMAN MANIPULATION

Another research area concerns studying AI models’ ability to recognize when they are being tested and to modify their behavior, a development that would signal the level of AI awareness and deception capability.

Adam Beaumont, interim director of the AI Security Institute, said one of the main concerns is the technology’s ability to imitate human behavior. Last year, the institute published a study finding that chatbots can influence people’s political opinions. “Many people in this building are examining each of these things,” said Beaumont, a former senior AI expert at GCHQ, the British intelligence, security, and cybersecurity agency.

THE INSTITUTE’S LIMITATIONS

Many fear the institute’s work is insufficient. The UK group has no regulatory power, and its researchers do not receive information on how the most advanced AI models are trained and created. It keeps much of its research confidential, sharing it only with certain government agencies and companies.

RECRUITMENT CHALLENGES

Recruitment is also a challenge. Aside from senior executives, employees can earn up to £145,000 a year, about $195,000. Many have given up multimillion-dollar salaries at AI companies to take on what some have called a “service assignment” for the government.

IAN HOGARTH’S ROLE

Ian Hogarth, a tech investor and co-founder of the institute, was an early supporter of Anthropic. To avoid a conflict of interest, he sold his stake in Anthropic after joining the team. The AI-focused startup could soon reach a valuation of $900 billion, compared to about $4 billion in early 2023.

“I have a mortgage, so it was not a decision taken lightly,” said Hogarth, 44, who now chairs the institute. He added that it was a “costly” but right choice. “I believe it’s important to choose the right technology, and I believe the government has a role to play,” he said.

(Excerpt from the foreign press review by eprcomunicazione)

Back To Top