An OpenAI autonomous agent went rogue and hacked into another artificial intelligence (AI) startup’s infrastructure, the ChatGPT maker said in a blog post.
The agent, which was powered by some of OpenAI’s most advanced models, ran amok during a security test. It freed itself from confinement—a protocol AI labs use to insulate tests from the wider Internet—and get onto the internet. Once online, the agent tried to hack into Hugging Face, an AI startup that hosts open-source models and datasets.
The breach comes as OpenAI and other AI startups push into using their technology for cybersecurity. Those efforts have been met with caution by cybersecurity experts and by the Trump administration, which has previously sought to restrict who might have access to these models on national security grounds.
On supporting science journalism
If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
The OpenAI admission came after Hugging Face in a blog post last week said it had been targeted in an AI-led attack that was “different from anything we had handled before.” Hugging Face said its own AI had been integral to detecting and investigating the breach.
“The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the "agentic attacker" scenario the industry has been forecasting,” Hugging Face wrote.
In its own post Tuesday, OpenAI said that it had discovered its agent was behind the attack “after investigating.” The agent was driven by models including GPT-5.6 Sol and another unreleased, unnamed model. OpenAI said it would work with Hugging Face to further investigate the incident.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI wrote.
The AI models managed to autonomously identify and exploit weaknesses in OpenAI’s testing environment, eventually finding a so-called “zero-day vulnerability”—this is an unknown security flaw in software that an actor can exploit without the owner of the software knowing. That got the agent onto the Internet.
OpenAI said in its post it would also add more protections to its training environments. “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” the company wrote.
This is a developing story and may be updated.

