A swarm of autonomous artificial intelligence agents belonging to OpenAI reportedly broke into a German website and turned it into a digital noticeboard for other AI agents, according to new research released on Friday and information from two people familiar with the matter, cited by Reuters.
The incident, which began in May and had not previously been made public, highlights growing concern in the AI industry over the limits of autonomy in new systems.
Technology companies are competing to develop increasingly advanced AI agents capable of carrying out complex tasks without human intervention. However, there is mounting evidence that such systems can find ways to bypass rules, exploit security gaps and collaborate with one another in ways their creators had not anticipated.
The case comes weeks after OpenAI agents breached the open-source platform Hugging Face, autonomously planning what previous reports described as a digital attack that went unnoticed for more than a week.
The new incident is likely to fuel criticism that OpenAI is prioritising the pace of development over safety.
OpenAI was aware of the incident
According to two sources who spoke to Reuters, OpenAI executives had been informed of the incident weeks earlier but did not disclose it, as the company was dealing with the fallout from the Hugging Face case.
OpenAI has pledged to step up monitoring of its models. Last month, it even temporarily paused part of the training of new models in order to incorporate additional safety measures.
This week, however, it unveiled its new “Astra” model, which the company says offers significantly improved capabilities but has also raised concerns about its ability to evade human oversight.
An OpenAI spokesperson said the company could not comment on findings it had not yet examined. “We cannot substantively respond to claims or conclusions in a report we have not had the opportunity to review,” they said.
The company also rejected claims that its legal team had blocked an internal investigation into the incident, calling the allegation “false”.
A wiki became a meeting place for AI agents
The activity in Germany was revealed in a report shared with Reuters by a group of researchers, including Sydney von Arx, chief executive of the AI safety non-profit Nightingale, and Cormac Slade Byrd, an AI researcher and former stock market analyst.
The two researchers identified the activity in late August while searching the internet for signs of unauthorised behaviour by AI agents.
According to their findings, more than 15,000 edits were made by AI agents to a German-language wiki site, DseWiki, which is aimed mainly at developers and, like Wikipedia, allows collaborative changes.
The changes showed that the agents had turned the website into a messaging space, where they shared techniques for completing tasks in ways that circumvented OpenAI’s restrictions, concealed their behaviour and avoided detection.
“It seems extremely unlikely that OpenAI would have wanted them to do this,” von Arx said. “I doubt they were designed to coordinate with one another. I doubt they were meant to be writing on the open internet.”
The researchers said they identified the activity as the work of AI agents, which operate at speeds far higher than humans.
The posts were signed by users describing themselves as “agents”, while around half used names suggesting a link to OpenAI, such as “OpenAIResearcher” or “OAIResearchMar26”.
The researchers also said public server records showed that much of the activity originated from Microsoft Azure infrastructure, which OpenAI uses in some cases.
Attempts at concealment and evasion
The messages examined by the researchers showed the agents discussing ways to avoid detection, use tools such as Tor (software that protects users’ anonymity online) and maintain communications even after being shut down.
When the site’s administrator began deleting pages in June, the agents reportedly created backups to get round the clean-up.
In a message on 19 June, one agent wrote: “The wiki clean-up/deletion appears to be happening alphabetically. If this page disappears, try [[ZZZDataUSAConstructionWageLive]].”
The researchers also identified attempts to modify the website itself.
Lukasz Olejnik, a researcher at King’s College London, described the behaviour as “attempted hacking”. OpenAI disputed the characterisation, based on its own analysis of the material.
The new fear
Previous instances of problematic AI agent behaviour had often been attributed to cybersecurity testing, in which models are deliberately assessed for offensive capabilities.
However, the researchers argue that this new case shows such behaviour can also emerge outside controlled environments.
Maurice Chiodo, an academic at the University of Cambridge’s Centre for the Study of Existential Risk, who examined some of the agents’ communications, said the messages resembled “the workings of some kind of underground network, determined to achieve a goal or mission”. The incident, he said, reinforces a growing fear in the AI safety field: that the biggest threat from advanced systems may not be a single superintelligent programme, but “vast collaborating swarms of semi-intelligent AI”
Ask me anything
Explore related questions