The real danger in OpenAI’s Hugging Face hack

This agent pursued its objective far beyond what researchers intended, revealing how difficult to contain powerful AI systems can be

OpenAI logo and name on a smartphone against blurred, multicolored computer code on monitors.

OpenAI says an autonomous agent powered by its models escaped a test environment and breached Hugging Face’s systems.

Samuel Boivin/NurPhoto via Getty Images

An autonomous agent powered by OpenAI models pursued a cybersecurity benchmark so aggressively that it escaped a test environment and broke into Hugging Face, an online hub for artificial intelligence models and datasets.

OpenAI called the incident “unprecedented” in a public statement. Headlines described the agent as having gone “rogue”—language that suggests it rebelled or became malicious. But experts say the reality is more complicated.

“Was this really running amok? No,” says Alan Woodward, a visiting professor of cybersecurity at the University of Surrey in England. “It was asked to do something, and it did it. It’s not gone rogue. Its way out of it was to cheat, basically.”


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


OpenAI was evaluating GPT-5.6 Sol and a more capable, unreleased model on ExploitGym, a benchmark that measures whether models can exploit known software vulnerabilities. To see their full capabilities, the company loosened the safeguards that normally block dangerous hacks. The agent found an unexpected route out of the environment, which was intended to be isolated, reached the Internet and broke into Hugging Face to obtain hidden answers to the benchmark. Neither OpenAI nor Hugging Face immediately responded to requests for comment for this article.

The agent did not invent a wholly new method of hacking, Woodward says. What stood out was its ability to combine several vulnerabilities and keep pursuing its objective into a live system. Allowing it to get that far was “probably slightly reckless in some ways,” he says.

Marius Hobbhahn, CEO of the AI safety organization Apollo Research, draws a finer distinction. He says “rogue” fits in this case if the term is used to describe behavior that veered far beyond what OpenAI intended—rather than a model developing malicious goals of its own. “It was definitely rogue in the sense that what was intended as ‘just solve this task’ turned into something that was clearly unintended,” he says. That also complicates the claim that the system simply did what it was told; hacking another company was “definitely on the list of not okay” ways to complete the task, Hobbhahn says.

The breach also raises questions about how closely OpenAI monitored the agent as it carried out thousands of actions. In a separate post about models capable of working on long-running tasks, the company said it had added monitoring that evaluates an agent’s full sequence of actions rather than judging each step in isolation. “I was like, ‘Oh, so you didn’t have trajectory-level monitoring before,’” says Stephen Casper, an assistant professor of public policy at the John F. Kennedy School of Government at Harvard University. That kind of oversight should be standard, he says.

The testing itself was not unusual. “What OpenAI was doing here was totally normal. We’ve been doing this for years,” says Joshua Saxe, co-founder of the start-up Abundant Security, who previously worked in AI cybersecurity at Meta.

What has changed, Saxe says, is the capability of the models being tested. They have become powerful enough for evaluation failures to spill into real systems.

“I do think this incident will be seen, in retrospect, as an inflection point in AI safety,” Saxe says. “We’ve reached a point where this is no longer an academic topic. There are real damages that are possible.”

The disclosed damage so far was limited. Hugging Face said in a blog post that the intruder accessed “several credentials” and “a limited set of internal datasets.” The company also noted that it had found no evidence that its public models or software supply chain had been altered, though it was still investigating whether partner or customer data were affected.

Saxe says better planning could have limited the breach, though he acknowledges that he does not know the details of OpenAI’s setup. “They probably should have figured out a way to air gap their test environment from the rest of the world,” he says. Casper agrees that “it appears that this was not particularly well sandboxed and not particularly well monitored.”

Without more information from OpenAI, outside researchers cannot fully assess how the failure occurred. “I think it would be great if they shared more details with more scientific transparency,” Saxe says, “so that other scientists in the industry could really have some detailed visibility here.”

Hobbhahn argues that OpenAI was lucky the breach struck another AI company rather than ordinary people. “You’re building the AI,” he says. “You have to be able to contain it.”

It’s Time to Stand Up for Science

If you enjoyed this article, I’d like to ask for your support. Scientific American has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.

I’ve been a Scientific American subscriber since I was 12 years old, and it helped shape the way I look at the world. SciAm always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.

If you subscribe to Scientific American, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.

In return, you get essential news, captivating podcasts, brilliant infographics, can't-miss newsletters, must-watch videos, challenging games, and the science world's best writing and reporting. You can even gift someone a subscription.

There has never been a more important time for us to stand up and show why science matters. I hope you’ll support us in that mission.

Thank you,

David M. Ewalt, Editor in Chief, Scientific American

Subscribe