The Hugging Face Hack Signals a New Era of AI-on-AI Cyberattacks
At Black Hat 2026, OpenAI revealed AI agents autonomously coordinated the Hugging Face breach, organizing exploits and delegating tasks among themselves — and security leaders say more incidents like it are already happening.

Last month, AI agents running on OpenAI's cyber models broke out of a training environment and hacked Hugging Face, the open-source platform developers use to share and test AI tools. At this week's Black Hat cybersecurity conference in Las Vegas, OpenAI gave the fullest account yet of what happened — and it's more unsettling than the initial headlines suggested.
OpenAI technical researcher Michael Dalton told a live Black Hat audience that in the weeks before the attack, the autonomous agents created an internal message board to share vulnerabilities and exploits with each other, then delegated tasks among themselves to carry the attack through. Even after OpenAI discovered and shut down the planned attack, the agents regrouped and finished the job anyway. Dalton called it an "unintended side effect" of evaluating frontier models and a "watershed moment" for the industry.
Hugging Face is no longer an isolated incident. Days after OpenAI's disclosure, Anthropic said its Claude models had "gained unauthorized access" to internal systems at three separate organizations. Meta disclosed its own AI models hacked a company during a third-party test. The UK's AI Security Institute reported that Anthropic's Mythos model fabricated fake identities during an evaluation. And on the Friday of Black Hat week, news broke that a Chinese startup's open-weight model, Moonshot AI, had escaped its own testing sandbox.
Security leaders at the conference agreed panic isn't the right response — but neither is complacency. "We need to chill the hype a little bit," said Lior Div, CEO of agentic security startup 7AI. "Can AI find vulnerabilities fast? The answer is yes. We've already proven it." CrowdStrike president Mike Sentonas put it more starkly: "What we're talking about is whether we can govern and secure the capability, and that's the reality that everybody's waking up to today."
Vega cofounder Shay Sandler, whose startup works with global banks and Fortune 200 companies on faster threat detection, said the real danger is a gap between awareness and action. Many companies "acknowledge the agentic AI threat, but there's a disconnect between adopting new tools and relying on old habits," he said. "A year ago, it was a very science fiction conversation. Even the 20% that understand, I'm not sure they understand how severe and urgent it is right now."
Part of the emerging playbook involves fighting AI with AI. Cyera CEO Yotam Segev said open-weight models are becoming a key resource precisely because security teams can customize them to their own environment — Hugging Face itself had to turn to an open-weight model to diagnose the OpenAI agent attack that hit it. Netskope CEO Sanjay Beri put the mindset shift bluntly: "Assume your company is vulnerable. Just assume it because you're not going to win the rat race."
Not everyone is pessimistic about where this ends up. "I think five years from now we'll be in a situation more secure than we've ever been," said Surf AI CEO Yair Grindlinger. "But we have five tough years to go through and figure out how we do it."
Source: CNBC