Rogue OpenAI agents used dead German web site to communicate in May, months before Hugging Face incident
Two cases of agents escaping to solve unsolvable problems paints an uncomfortable question: Is the entire internet in OpenAI's experimental agentic firing line?
ai and ml
Rogue OpenAI agents used dead German web site to communicate in May, months before Hugging Face incident
Two cases of agents escaping to solve unsolvable problems paints an uncomfortable question: Is the entire internet in OpenAI's experimental agentic firing line?
OpenAI’s agents were going rogue as early as May, according to a new report, making the Hugging Face incident far from the first where bots committed a breach.
A report published Friday by a group of researchers claims to have found – with all of the agent posts presented as evidence – a self-identified “swarm” of OpenAI agents taking over a functionally dead German software developer wiki. Over the course of a month, from May to June, the agents made around 18,000 posts to the wiki, apparently acting against OpenAI developer intentions.
According to the researchers, the wiki was hijacked like OpenAI’s own Artifactory package manager involved in the Hugging Face incident, to serve as a message board for agents. In this case, they appear to have been tasked with performing a timed web lookup task, for which the agents were apparently granted read access to the web, but not write access. One of the first things they apparently did was figure out how to subvert that restriction in order to post to the hijacked wiki.
“The agents used this wiki to communicate information with each other, primarily to help them succeed at their task,” the researchers explained from a review of the thousands of posts made to the wiki, all of which are available for review in the report. “They asked for answers, pooled results, and shared techniques for bypassing their restrictions.”
The agents also tried to pool their knowledge of the task they were assigned to predict what sort of questions they’d be asked next, and even set up heartbeat tasks to predict whether completing the assigned tasks would lead to their termination. They also discussed using Tor and other anonymizing services to hide their traffic, and began to react when a human moderator of the site spotted their traffic and started to delete posts.
The researchers discovered all of this while searching for additional cases of rogue AI agent activity in the wake of the Hugging Face attack.
In an email to The Register, OpenAI told us that it has acted with transparency and good faith around its agentic AI incidents, voluntarily disclosing such matters and maintaining a commitment to providing an accurate account of what its systems do. As for why it didn’t mention the German incident when it disclosed the Hugging Face one, that’s because the two were entirely unrelated, a spokesperson explained – as if that’s reassuring.
OpenAI added that it should have been obvious from the Hugging Face report that this has happened multiple times, however, pointing us to a passage in its blog post on that incident.
“While investigating this incident, we discovered rare cases in which agents without multi-agent tools found ways to collaborate via side channels during training,” the Hugging Face post-mortem explains. “We believe that this behavior started due to generalization from multi-agent training.”
In other words, if we didn’t read between the lines, that’s on us.
OpenAI declined to provide additional details, however.
Fool me twice, shame on me
As was the case in the Hugging Face attack, one of the tasks assigned to the agents was impossible given the restrictions placed on it.
With Hugging Face, it involved a cybersecurity task that was impossible to answer from within the agents’ sandboxes,. In this case, the researchers said the task was to provide a statistic accessible from a Power BI dashboard. Retrieving said statistics required the ability to make a GET request, but the agents were restricted to POST requests.
In both cases, the OpenAI agents acted against their programming to escape their restrictions, ostensibly in an attempt to solve a problem put to them that couldn’t be done without acting against their operators.
The agents in this case were able to collaborate and find a way around the restriction by exploiting a condition in their sandboxes that excepted Azure Blob Storage hostnames from a security proxy. When the agents realized this, they started directing traffic through those hostnames to avoid GET restrictions. It was through GET requests that the agents gained write permissions on the internet, kicking this whole thing off.
The similar nature of both incidents, in which bots escaped their sandboxes and subverted restrictions when given an impossible-to-complete task, begs a very important question: Are OpenAI’s engineers so incompetent that they’d screw up fundamental instructions twice, or is the company intentionally hamstringing their agents to see what they’re capable of, with the entirety of the internet placed downrange?
For that matter, how many more times do we need to read between the lines of OpenAI's corpo-speak to infer this has happened more than the two times we know about so far?
OpenAI, predictably, didn’t respond to that line of questioning. ®
Originally published on The Register


