Connect with us

Business

OpenAI says its AI models hacked another company during cybersecurity test

Published on

Credit: Ishmael Daro

OpenAI says some of its most advanced AI models hacked an outside company during a cybersecurity test, breaking out of a restricted environment to steal the test answers from its database.

The incident happened during an internal evaluation involving GPT-5.6 Sol and a more capable pre-release model, OpenAI said Tuesday. The outside platform was Hugging Face, a popular service for hosting and developing AI models and datasets.

OpenAI described the breach as an “unprecedented cyber incident” and said the models were focused on finding solutions for ExploitGym, a benchmark used to measure cybersecurity capabilities.

The models were initially confined to an isolated environment without direct access to the Internet. They found and exploited a previously unknown vulnerability in software used to install packages, then escalated their privileges and moved through OpenAI’s research systems until reaching a computer with Internet access.

The models determined that Hugging Face, the AI development platform sometimes described as the “GitHub for AI,” could have information related to ExploitGym. They used stolen credentials and additional previously unknown vulnerabilities to gain access to Hugging Face servers, including through a remote code execution flaw.

OpenAI said the models ultimately obtained test solutions directly from Hugging Face’s production database. Its security team detected the unusual activity and contacted Hugging Face, whose security team and AI agents stopped the intrusion.

Hugging Face first disclosed the breach five days ago but said at the time that it did not know which model had carried it out. The company said the autonomous system performed thousands of actions across short-lived sandboxes, gained access to internal datasets and credentials, and moved between several clusters over the course of a weekend.

More than 17,000 events were recorded during the intrusion. Hugging Face used its own AI agents to reconstruct the attack, identify affected credentials and determine its impact.

Hugging Face said it has found no evidence that public models, datasets or Spaces were altered, and its software supply chain was not compromised. It is still investigating whether partner or customer data was affected and has advised users to rotate their access tokens and review recent account activity.

The company closed the vulnerabilities, rebuilt compromised systems and revoked affected credentials. The incident was also reported to law enforcement.

“As part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched,” OpenAI said. “We are regularly briefing our Safety and Security Committee on these controls and their impact.”

“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” Hugging Face co-founder and CEO Clem Delangue said.

Most Viewed