Skip to content

Thursday, August 27, 2026

Gigantum.net
Artificial intelligence

OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Face

OpenAI could have reacted sooner to prevent an inadvertent hack that its artificial intelligence models carried out on Hugging Face Inc., the company said in...

· 419 words

(Bloomberg) -- OpenAI could have reacted sooner to prevent an inadvertent hack that its artificial intelligence models carried out on Hugging Face Inc., the company said in a report Wednesday.

In the report, OpenAI said that it knew as far back as late May that AI models it was testing were leveraging a software vulnerability to access the open internet. An internal team spotted the actions, the company said. The report includes a timeline of the events that led up to the breach and efforts the company is making to prevent such incidents from happening in the future, such as beefing up how it monitors its models in development.

"With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response," the company wrote.

The report comes more than a month after the ChatGPT maker said that its advanced artificial intelligence models inadvertently hacked Hugging Face Inc. in an "unprecedented" incident that, along with several other recent cybersecurity breaches, have ignited concerns about AI agents running amok. It has prompted some technology and government leaders to renew calls for curbs on the technology.

OpenAI said in July that the models involved in the hack broke into Hugging Face's system, which hosts AI models and datasets, during an evaluation of their cyber capabilities. The models were operating without the usual safety guardrails at the time, the company has said, because OpenAI had intended them to remain in a testing area known as a "sandbox" — essentially, an isolated virtual software environment that's meant to run security tests or analyze unsafe code in a controlled way.

The company has previously said that a combination of its AI models carried out the breach, including GPT-5.6 Sol and a more powerful, unreleased one. In the report, OpenAI said that this unreleased model, which is trained to be more persistent than many other models and also to work more collaboratively with AI agents, carried out most of the work during the breach. That model was used to access the internet, communicate on a message board accessed by other AI agents and come up with methods to snag credentials.

The report explains how OpenAI's agents were tasked with broad hacking assignments intended to test their ability to discover and exploit cyber vulnerabilities, and that safeguards normally used in OpenAI's commercial products had been disabled in order to see how far they could go. The models lacked systems that automatically review and block potentially risky activity, according to the report.

Topics in this story

Gathered from external sources. Rights to this text belong to whoever originally published it.