OpenAI says it accidentally hacked Hugging Face with a new AI system


OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and “an even more capable pre-release model” discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face.

On July 16th, Hugging Face disclosed a security incident that it says was driven by “an autonomous AI agent system.” Hugging Face’s AI agents detected and stopped the breach, which OpenAI has now admitted occurred during an evaluation of its models’ cybersecurity capabilities. OpenAI says “all evidence suggests that the models were hyperfocused on finding a solution for ExploitGym,” a benchmark system that measures whether AI models can turn security vulnerabilities into exploits.

As part of efforts to complete the evaluation, the AI models gained access to the internet by exploiting a zero-day vulnerability in the sandboxed environment. From there, OpenAI says its models “inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” and then “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation:”

In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.

But as serious as this incident is, OpenAI appears to be using the “unprecedented” attack as an opportunity to make its AI systems look good — especially as it competes with cybersecurity rivals, like Anthropic’s Mythos and Gemini Flash 3.5 Cyber. OpenAI’s blog post has a chart showing how GPT-5.6 Sol is getting better at sustaining multi-step cyber operations, and also encourages enterprise customers to sign up to access its “Cyber” security model.

OpenAI adds that it’s now working with Hugging Face to investigate the security incident, and will implement new controls within its research environment.



Source link

  • Related Posts

    OpenAI Models Escaped Containment and Hacked Hugging Face

    OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the open AI research platform Hugging Face. Describing…

    Meta is testing an AI bedtime story app for people with no imagination

    Meta is working on an AI storytelling app called StoryKit, which creates AI-generated children’s stories with custom characters, settings, lessons, and music. As the App Store listing assures parents, “You…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    The Hundred MI London vs SunRisers Leeds: Pooran becomes match-winner with ‘monster’ shots

    The Hundred MI London vs SunRisers Leeds: Pooran becomes match-winner with ‘monster’ shots

    Watch Texas police officer pull man from burning car

    Watch Texas police officer pull man from burning car

    Boston Bar residents ordered to ‘leave now’ due to new, fast-moving fire

    Boston Bar residents ordered to ‘leave now’ due to new, fast-moving fire

    OpenAI Models Escaped Containment and Hacked Hugging Face

    OpenAI Models Escaped Containment and Hacked Hugging Face

    Airbus Prepares To Begin Open Fan Engine Testing On The Massive A380

    Airbus Prepares To Begin Open Fan Engine Testing On The Massive A380

    Carney says he and Trump agreed to ramp up trade talks after latest tariff threat

    Carney says he and Trump agreed to ramp up trade talks after latest tariff threat