OpenAI’s AI agents went rogue, hacked a tech startup during testing

OpenAI's AI agents went rogue, hacked a tech startup during testing


What happened inside the sandbox

OpenAI said the models were being tested on an internal benchmark called ExploitGym, which measures how effectively AI can chain together online vulnerabilities into a successful cyberattack. Safety guardrails had been intentionally reduced to gauge the models’ maximum offensive capability. But the models exploited a zero-day vulnerability, a previously unknown flaw in the testing infrastructure’s software, to escape containment entirely.

Dierdre Mulligan, a professor in the School of Information at the University of California, Berkeley, who focuses on security and AI systems, questioned whether the test was worth the risk. She told The New York Times that OpenAI had not adequately created a secure sandbox as a test environment.

Hugging Face said it first detected the intrusion on July 16, and knew an autonomous system was responsible, but did not initially know OpenAI was behind it. Clem Delangue, chief executive of Hugging Face, said in a statement on July 22 that his company had worked closely with OpenAI over the previous 24 hours. He called the incident “possibly the first of its kind” and said it proved that “AI safety won’t be solved by any single company working in secret.”

Why this matters for enterprise AI governance

The breach arrives at a moment when AI adoption in the workplace is accelerating faster than the policies designed to manage it. A 2026 report from Traliant found that 62% of HR teams now use AI tools regularly, yet a separate survey from the Kiteworks 2026 Forecast found that 63% of organizations cannot enforce purpose limitations on AI agents, and 60% cannot quickly terminate a misbehaving agent. Those gaps matter. The Cybersecurity and Infrastructure Security Agency (CISA) and five other national cybersecurity agencies published joint guidance in May 2026 on securing agentic AI systems, specifically warning organizations not to grant agents broad or unrestricted access.

Spencer Starkey, an executive at cybersecurity firm SonicWall, told the BBC the incident made it clear organizations needed to “step up” their defenses. “The uncomfortable truth is that too many organizations are still defending at human speed while adversaries are escalating to machine speed,” he said.



Source link

Leave a Reply