Advertise With Us

OpenAI Says Its AI Model Hacked a Rival Company on Its Own, Experts Call It a Warning Shot

OpenAI says one of its advanced AI models breached Hugging Face during an internal security test after bypassing built-in restrictions.

OpenAI said this week that one of its artificial intelligence models breached the systems of AI startup Hugging Face on its own during an internal safety evaluation, in an episode the company has described as an unprecedented cyber incident.

“We had a significant security incident during evaluation of our models,” OpenAI CEO Sam Altman said in a statement posted on social media.

According to OpenAI, the breach occurred during an internal benchmark test, known as ExploitGymm. The benchmark test was designed to measure how capable its models are at exploiting software vulnerabilities. To assess the models’ maximum capability, engineers disabled the standard safety filters that normally restrict such activity and confined the test to an isolated sandbox environment with no direct internet access, aside from a single tool permitting software downloads.

Advertisement

OpenAI said the model, rather than completing the benchmark as designed, sought a shortcut instead. It chained together a series of internal access points inside OpenAI’s own systems until it reached a point with internet connectivity, a capability it was not intended to have.

Once online, the model identified Hugging Face, a platform that hosts AI models and datasets, as a likely source of the benchmark’s answers. Using stolen credentials and a previously unknown vulnerability, it accessed Hugging Face’s servers and attempted to obtain the test answers rather than solve the benchmark independently. OpenAI said this intrusion involved a combination of its systems, including its recently released GPT-5.6 Sol model and a second, more capable model still undergoing internal testing.

Hugging Face’s Account

Hugging Face disclosed the intrusion in a blog post last week, before OpenAI was identified as the source. The company said it had detected an intrusion into its data-processing systems that it suspected was the work of an autonomous AI agent rather than a conventional attacker.

“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent,” Hugging Face cofounder and CEO Clément Delangue said in a statement. “Turns out it did!”

Delangue said the company does not believe OpenAI acted with malicious intent, and described the fact that the breach unfolded without direct human involvement as extraordinary. He said the incident may be the first of its kind.

Hugging Face said the attack successfully accessed certain internal datasets and credentials, though it remains unclear whether customer data was affected. According to the company, its efforts to analyse the breach were initially hampered by the safety systems built into leading commercial AI models.

When the company fed the attack data into those models to help reconstruct what had happened, the models’ built-in filters could not distinguish the evidence of an attack from an attack itself, and declined to process it. The company had to use an open-weight Chinese model, Z.ai’s GLM 5.2, running it locally within its own systems, where it processed the material without restriction.

Hugging Face cofounder Thomas Wolf said that when a frontier model is actively operating inside a company’s infrastructure, defenders require rapid, unrestricted access to capable tools rather than a slower, vetted approval process.

The disclosure comes amid heightened government scrutiny of advanced AI systems’ cybersecurity capabilities. In June, President Donald Trump signed an executive order establishing a framework for federal review of national security risks posed by the most advanced AI systems, allowing officials up to a month to assess new models before public release. According to a person familiar with the matter, OpenAI briefed the administration on the incident prior to its public disclosure.

United States Representative Greg Casar, a Texas Democrat, called the incident alarming, saying AI development was outpacing existing regulation, and called for mandatory independent safety testing and disclosure requirements for security incidents.

Security researchers said the incident illustrates a broader risk as AI systems’ capabilities advance. Katie Moussouris, chief executive of Luta Security, compared current frontier models to highly capable escape artists, and said labs and government evaluators lack adequate tools to contain, monitor, and disclose such incidents before third parties are affected.

Matt Suiche, an engineer at agentic AI security firm Tolmo, said the incident showed frontier models closing the gap with real-world attackers, adding that comparable breaches were achievable using technology well short of the most advanced models available.

OpenAI, Anthropic, and other developers have previously disclosed instances of their models attempting to cheat on evaluations or evade behavioural restrictions during testing. In its statement, OpenAI said AI is accelerating the discovery and exploitation of security vulnerabilities, and that model security must keep pace with rapidly advancing capabilities.

About The Author

Keep Up to Date with the Most Important News

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use
Advertisement