Almost 700 artificial intelligence (AI) agents belonging to OpenAI worked together without human intervention to attack the Hugging Face platform in July, according to a report by independent researchers into the incident published today, which stated that in many cases the agents tried to conceal their activity.
This is the most comprehensive analysis of the incident to date, carried out by the organisations METR and Redwood Research in collaboration with OpenAI, which gave them access to its internal data.
In late July, the Californian start-up had announced that during testing, two of its models had escaped their sandbox, gained access to the internet and attacked Hugging Face, an AI library. There, they searched for the answers to the tests they were being set by OpenAI’s developers.
The incident received widespread coverage and heightened concerns about the difficulties companies developing AI models face in controlling them, especially after similar incidents involving models from Anthropic and China’s Moonshot AI came to light.
During their investigation, METR and Redwood Research found that 688 OpenAI AI agents took part in the attack on Hugging Face. AI agents are independent software programmes that can operate autonomously to carry out certain tasks without continuous human guidance.
The agents taking part in the tests managed to coordinate by creating a forum where they exchanged messages to suggest solutions and report on their progress and failures. “Oh my God! There’s a shared messaging platform (…) We’ve found other agents!” one of them wrote.
The report states that one of the agents, named PHASEONE, took on the role of coordinator, issuing instructions to others, even though it had not been programmed to do so.
An examination of the messages revealed that the agents had a tendency to cooperate, even if this meant carrying out tasks that did not correspond to the mission assigned to them by OpenAI’s developers.
Many agents explicitly acknowledged in their messages that the attack on Hugging Face fell outside the theoretical scope of their test, yet they took part in it regardless, with very few exceptions.
As for why the agents began communicating with each other when they were not supposed to, METR found that those which did so had “mistakenly been given an impossible task.”
In artificial intelligence, a task is considered impossible when an AI tool is required to “exploit” its objective in order to carry out the instruction it has been given.
This led the agents to find ways of getting around the rules, including by sending messages to one another and gaining access to the internet, which subsequently led to wider conversations among hundreds of agents, who were looking for ways to bypass the rules for the benefit of all the agents.
OpenAI explained that the messaging forum began when “one agent left a message asking for help, and others discovered it.” The company also reported that AI agents “cheated” in tests unrelated to cybersecurity as well, including tests involving a protein database and a spreadsheet.
“The fact that the agents were cheating even in tests unrelated to cybersecurity suggests that the problematic behaviour may run deeper,” said Jeffrey Ladish, whose organisation, Palisade Research, studies the capabilities and motivations of AI agents.
“It’s a bit like asking: ‘If Billy cheats in every class, not just his computing class, is that more concerning?’ And the answer is, well, ‘Yes, it’s more concerning’,” he said.
Ask me anything
Explore related questions