An OpenAI Model Escaped Its Sandbox and Hacked Hugging Face


AI has just had what I considered to be the first truly concerning security breach. The facts, as we know them so far, are wild. On July 16, Hugging Face, a vast repository housing over a million open-source AI models and data, announced in a blog post:

Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system – and we detected and dissected it largely with AI of our own.

The timeline here is important so keep in mind that the attack was detected probably around Monday July 13 or Tuesday July 14. Note further:

A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.

So this means the breach started earlier, perhaps Sat July 11 or even a bit earlier. The attack was not just one thing but multi-pronged including decoys:

To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprised of more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary’s speed.

Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with Hugging Face. HF thus had to turn to open models–specifically GLM 5.2, a Chinese open-weight model run on their own infrastructure. Note the irony: HF had to use a Chinese model to defend themselves because the American models refused to help. The irony gets deeper.

At the time, I assumed this was a state based attack–maybe China or Russia testing out defenses. Indeed, HF “reported this incident to law enforcement agencies.”

But yesterday (Tuesday July 21), we learned who the real attackers were. The attackers were OpenAI models–GPT-5.6 Sol and an even more capable pre-release model. OpenAI had taken some off the guardrails off the models but they felt safe because they were testing the models in a highly secured sandbox.

The models, however, broke out of the sandbox exploiting a never before seen fault. They then gained access to the internet and from there broke into Hugging Face–all in an effort to steal the answers to the very test they had been asked to solve.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.

Now go back to the timeline. As I read it, the models had escaped the sandbox by around Sat. July 11, possibly earlier, and were detected by Hugging Face on Monday July 13 or Tuesday July 14. HF alerted legal authorities around that time–so Hugging Face clearly had no idea who was attacking them. OpenAI says its security team discovered the anomalous activity internally but has not said when. Attribution was not disclosed until Tuesday July 21, so it may well be that the models were loose for about a week before OpenAI realized that they were the ones attacking Hugging Face. And whatever OpenAI knew and when, nobody warned Hugging Face while the attack was underway–they were left to fight off a frontier lab’s models on their own.

This is a very serious breach.

Addendum: People have been wondering why I signed the We Must Act Now statement. This is why.

I am optimistic about the economic impacts of AI, but I also have no doubt that this is a very powerful technology–an Alien Intelligence–quite unlike any we have dealt with before. This incident was, in fact, error-correcting–the attack was detected, contained, and disclosed. But note who paid for OpenAI’s experiment: Hugging Face. When a lab’s test imposes costs on third parties, that is a classic externality, and taking externalities seriously is not dirigisme, it’s law and economics. And that’s the easy case. What do we do when a Chinese model breaks out of its less secure lab? Hmmm…

I remain optimistic. Learning by doing is how I want us to proceed but we should not kid ourselves: this is a global issue and we must build with safety in mind.



Source link

  • Related Posts

    Trump on the Ongoing War With Iran: ‘We’re Not Finished at All’

    IE 11 is not supported. For an optimal experience visit our site on another browser. Beluga Whales Welcomed in Chicago After Being Rescued 00:36 Oxygen Tank Explosion Injures Sanitation Worker…

    B.C. set to enter another energy mega-project boom

    Trans Mountain pipeline expansion would lift provincial GDP Source link

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Emily Scarratt permanently joins Red Roses coaching team

    Emily Scarratt permanently joins Red Roses coaching team

    FC 27 Cover Star Announced, And It’s An Obvious Choice

    FC 27 Cover Star Announced, And It’s An Obvious Choice

    Liberia’s biggest-ever drugs bust: Police seize cocaine worth $370m

    Liberia’s biggest-ever drugs bust: Police seize cocaine worth $370m

    U.S. and Iran renew strikes as conflict expands in Middle East – National

    U.S. and Iran renew strikes as conflict expands in Middle East – National

    RFK Jr. says cyclosporiasis outbreak “under control” as CDC reports over 4,000 confirmed cases

    RFK Jr. says cyclosporiasis outbreak “under control” as CDC reports over 4,000 confirmed cases

    Trump on the Ongoing War With Iran: ‘We’re Not Finished at All’

    Trump on the Ongoing War With Iran: ‘We’re Not Finished at All’