Last week a OpenAI agent allegedly, went rogue and attacked HuggingFace. Looking for an answer to a quiz. This it did over the weekend prior to 25-26 Jul 2026.
OpenAI claims that they were running the agent AI an secure environment with internet access via proxy. The agent broke upon the proxy or side jumped to other nodes and went after Hugging Face servers. I am not aware of the level of competence in OpenAI to setup a secure environment. Nor the proxy used. Was it an old, unpatched proxy? There are a lot of questions which have to be answered by OpenAI.
Final puzzle. It went rouge. Humans programmed it. Humans gave it weights on actions to be performed and outcome to be sought. Humans started the agent and let it run over the weekend. How is the agent rogue?
The agent did not invent a wholly new method of hacking, Alan Woodward, a visiting professor of cybersecurity at the University of Surrey in England says. What stood out was its ability to combine several vulnerabilities and keep pursuing its objective into a live system.
Hmmm. Food for thought. Did not invent a wholly new method of hacking.
Another thought experiment to chew on.
Person X created an apparatus to do two things. One floor the accelerator on a Ford F150 fully loaded truck with ammonium nitrate and pure sodium. Two to randomly twist the steering wheel left to right. Person X takes this Ford 150 on I-95 and starts the apparatus. A pile up ensures. The apparatus went rogue? The Ford F150 went rogue? Who really went rogue?
It’s pretty simple here: if it comes from the horse’s mouth, from the individuals that host, sell, or somehow make a living out of this technology, you can be certain it didn’t actually happen in the way they said it did. This hack smells like a marketing stunt at best and incompetent “research” at worst. It’s probably both knowing AI bros, though.
Yes. Because the agent is fully in control of the OpenAI engineers and researchers. They can be stopped mid way, they can be fully controlled, they can be taught new skills and how to process new information fairly easily. OpenAI is liable for what the AI did, the AI that they told to act in this fashion.
My guess is that they are already prepared for situations like this. Would be interesting to know if they actually do have insurances for such cases of unintentional damages to third parties. And what their policies actually would be.
It’s not like that their legal team is bored, quite the contrary. But as OpenAi actually claimed to work together with hugging face to address the security incident, it’s pretty unlikely that there would be legal consequences.
In the end, it’s was an agent that had the task to solve questions of ExploitGym, which is a cybersecurity benchmark. And it came to the conclusion that it might be easier to retrieve the answers to that benchmark from hugging face servers, where the guardrails might easier to overcome than answering the tasks it was faced with in the first place. Basically cheating. I wonder where its inspiration may have come from…
Funnily enough, hugging face had to spin up a different AI model on their own systems to contain the attack.
In the end, there is some news and the associated discussions about it. A zero day exploit has been identified, but it’s unclear if it has been an already known one that simply wasn’t patched yet, or if it has been a genuinely new one.
In the end, OpenAI should have learned their lesson that a sandbox with an internet connection technically isn’t a sandbox anymore, even if it is only an proxy for package retrieval from the outside.
How would insurance even qualify and quantify a similar risk? There exists no risk models which can be used for such a type of scenarios. Eventually the accountability will fall on the entity, company, person who ran the agent. In this case OpenAI.
If OpenAI were to run a new agent and then the new agent were to run the agent which ended up attacking Hugging Face then also OpenAI is held responsible.
Even if OpenAI were to run a new agent that agent hacked into Anthropic and ran an agent from Claude which attacked Hugging Face then also OpenAI is held responsible.
About the Agent cheating, yeah OpenAI coded cheating capability into the agent. They configured the weights as well as ranked the objectives for the agent to fulfill. This would include the weights for cheating too. It is not as if the Agent can go on its own and acquire cheating capability.
Here’s the thing: we don’t have any way of knowing if this was orchestrated by the agent itself or if the engineering and research team decided to hit a fellow company that they talked with beforehand. It sounds very conspiratorial, I know, but I don’t put it behind OpenAI and HuggingFace to pull such a stunt for marketing. They are lying about a lot of things, I don’t see why they wouldn’t do something similar to this for marketing, Anthropic style.
They are also incompetent. OpenAI seemingly knew that testing cybersecurity capabilities with a computer connected to the internet is risky, yet they did it anyway. Standard practice for this type of stuff is to have the computer airgapped, but OpenAI didn’t do this very basic protection procedure because “it slowed down research”. It also seems that HuggingFace’s backend is more or less entirely vibe coded, so who knows how easy it actually is for even a common Joe to hack into HuggingFace.
We also have to remember that, in this point in time, agents can’t think entirely on their own. They have to get a prompt to do something from a human, or team of humans. It can’t act on its own volition entirely. So the legal team can try as hard as they please, but reality isn’t really kind to them.