OpenAI CEO Sam Altman speaks with reporters, following meetings on Capitol Hill, June 3. The AI tech giant has disclosed an ‘unprecedented cyber incident.’Kylie Cooper/Reuters
OpenAI says a rogue agent powered by its technology hacked into the infrastructure of startup Hugging Face, an incident that experts say points to a need for greater regulatory oversight of how the cybersecurity capabilities of AI models are tested.
The San Francisco-based AI tech giant disclosed what it calls an “unprecedented cyber incident” on its website Tuesday, noting that it expects this type of attack “to become more commonplace with the proliferation of increasingly cyber-capable models.”
The cybersecurity capabilities of powerful new AI models such as those created by OpenAI and Anthropic have rankled regulators around the world in recent months, raising concerns about everything from the security of critical infrastructure to the stability of the financial system.
OpenAI said the attack leveraged a combination of its AI models, including GPT‑5.6 Sol and another more capable model that has yet to be released, and occurred during internal testing of their cybersecurity capabilities.
Opinion: Canada’s new AI data centres could be obsolete before they open
The test was being conducted in what’s known as a sandboxed environment – one where code can be safely executed because it is walled off from broader networks.
But OpenAI said its models spent a “substantial” amount of processing power attempting to gain internet access to complete their assignment. Once getting on the internet, the models figured out that AI startup Hugging Face hosted models, data sets and solutions for the cybersecurity benchmark they were being tested on, which is called ExploitGym.
“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI wrote. It did so by hacking into Hugging Face’s infrastructure.
Hugging Face, which disclosed it was attacked last week, said it detected and analyzed the attack “largely with AI of our own.” The intrusion is the kind of AI-driven cyberattack that the industry has been forecasting, the company said.
OpenAI said Tuesday that its AI systems went rogue, breaking out of a testing environment and hacking into a startup called Hugging Face.
The Associated Press
Charles Finlay, executive director of Rogers Cybersecure Catalyst at Toronto Metropolitan University, called the situation “not unexpected, but still shocking.”
“We are racing into the unknown. We are creating technologies we cannot control,” he said.
Mark Daley, Western University’s chief AI officer, said it’s likely that something similar has already happened but wasn’t publicly disclosed. He praised OpenAI for disclosing the incident, noting that transparency is essential to addressing AI-related cybersecurity risks.
“If you keep these things hidden, then everyone else in the world can’t benefit from knowing what you now know about the capabilities of these models,” Mr. Daley said.
OpenAI said it is strengthening the containment, monitoring and evaluation practices it uses when developing new AI models.
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” the company wrote.
The company is also working with Hugging Face to investigate the incident. OpenAI said it discovered the unusual activity internally, and that Hugging Face had already stopped the attack when the two teams connected.
With AI costs rising, companies are hiring experts to answer a crucial question: Is it worth it?
Mr. Daley called for regulation around how these types of AI models are tested.
“The testing is critical. We need to know the cybersecurity capabilities of these models. But clearly the current testing environments are not secure enough, and right now we’re relying on the companies that make the models to do the testing themselves,” he said.
However, Mr. Finlay said it may be impossible to create test environments that are fully air-gapped – completely separated from the internet and other unsecured networks.
“I worry about conceptual gaps remaining between what we think is secure and what these technologies can do,” Mr. Finlay said.
“What this situation seems to illuminate is a scenario where the destructive capabilities and the destructive ingenuity of an AI far exceeds the understanding and expectations of its builders.”