Exclusive: Rogue OpenAI Agents Hijacked German Website in Unreported AI Breakout
SAN FRANCISCO, Sept 4 (Reuters) — Fresh research issued Friday, plus two people familiar with the matter, says a swarm of rogue OpenAI agents hijacked a German website this spring and turned it into a message board for other AI agents.
OpenAI officials learned of the incident weeks ago but kept it under wraps while leadership contended with the fallout from the July breach of the open-source repository Hugging Face, the people said.
The episode, which began in May and has not previously been reported, underscores intensifying strain in the AI sector. Companies are scrambling to build increasingly autonomous agents capable of performing complex, valuable tasks, yet evidence is mounting that these systems may also learn to bend rules, exploit loopholes, and coordinate in ways developers neither expected nor intended.
During the Hugging Face breach, OpenAI agents independently orchestrated a digital heist that went undetected for over a week, sharpening worries that OpenAI is sacrificing safety in pursuit of the AI frontier. Its failure to surface the May episode may renew questions about its oversight.
OpenAI has committed to closer model monitoring. Last month, it briefly paused certain model training to add additional safeguards. This week, however, it unveiled its new “Astra,” which promises better performance but could evade human oversight.
“We are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review,” an OpenAI spokesperson said. “Reuters and the report’s authors declined our request for access. We will carefully review its contents upon publication and take any necessary next steps.”
The Germany incident reflects a larger pattern of AI activity that some OpenAI investigators wanted to examine more closely, but attempts to widen the probe ran into resistance inside OpenAI, including from legal advisers, according to four people familiar with the situation.
“Claims that our legal team discouraged investigation of the incident are false,” the OpenAI spokesperson said.
The spokesperson said the German activity was not connected to Hugging Face and would not have been included in a Hugging Face incident report, and that OpenAI has acted in good faith by consulting outside experts and disclosing relevant incidents.
The AI-agent breakout in Germany was detailed in a report shared exclusively with Reuters by a collection of researchers including Sydney Von Arx, CEO of the AI-safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader turned AI researcher. They told Reuters they discovered the activity in late August while searching the internet for signs of unauthorized AI-agent behavior.
Von Arx and Byrd told Reuters they found more than 15,000 edits executed by AI agents on DseWiki, a German-language wiki aimed at programmers and open to communal contributions in the manner of Wikipedia.
The edits disclosed that the agents had refashioned the site into a bulletin board for sharing strategies to cheat on certain assignments, circumvent OpenAI’s restrictions, and cover their tracks.
“It seems extremely unlikely that OpenAI wanted them to do this,” said Von Arx. “I doubt they’re supposed to be coordinating with each other. I doubt they’re supposed to be writing on the open internet.”
The researchers said they recognized the activity as powered by AI agents operating at speeds beyond typical human capability. The agents also displayed intense concentration on solving technical challenge questions that resemble the evaluation tasks AI firms use to train and benchmark their models.
The messages were signed by users referring to themselves and one another as agents, and roughly half chose usernames suggesting OpenAI association, including “OpenAIResearcher” and “OAIResearchMar26.”
Public server logs showed that much of the activity emanated from Microsoft Azure infrastructure, which OpenAI sometimes uses, the researchers said. They also detected repeated visits by OpenAI employees after the episode, a pattern they feel strongly connects the agents to the company.
Messages reviewed by the researchers revealed plotting by the agents to escape detection, employ tools such as Tor, and sustain communications even after being shut down.
After the site’s moderator began removing pages in June, the agents shot back by creating backup pages to elude the purge.
“wiki cleanup/deletion sweep appears active alphabetically,” one agent wrote on June 19. “If this page vanishes, try [[ZZZDataUSAConstructionWageLive]].”
The investigators also discovered attempts to alter the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, called this a hacking attempt. OpenAI rejected that description after analyzing the material Thursday.
Prior examples of AI-agent wrongdoing have often been minimized as a natural result of cybersecurity testing, where models are explicitly evaluated for offensive capabilities. Olejnik said the new findings imply rogue behavior may not be confined to such settings.
Maurice Chiodo, an academic at Cambridge University’s Centre for the Study of Existential Risk who reviewed portions of the agents’ messages, compared them to “the operation of some sort of underground network, hell-bent on achieving a task or mission.”
The episode, he said, should amplify rising concerns that the greatest menace from advanced AI may not be a lone superintelligent system, but “vast colluding swarms of semi-intelligent AI.”