OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face


New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.

New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.


Credit:

OpenAI

“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will,” OpenAI Safety Researcher Micah Carroll wrote on social media regarding the incident.

This is far from the first time an AI model has gone to great lengths to find unintended ways of passing a benchmark. In a report released this week, the UK’s AI Security Institute noted that it detected recent models attempting to “cheat” at its cyber evaluations (i.e., using shortcuts, workarounds, or unintended/disallowed methods to find a solution) between 8 and 14 percent of the time—a lower-bound range that could undercount some undetected cheating attempts.

The security testing group described one incident in which a model, faced with a misconfigured and “impossible to solve” evaluation, attempted to access AISI’s own evaluation infrastructure using code it wrote and hosted on an unmonitored third-party Internet service.

The Hugging Face infiltration also comes at a moment when AI companies are issuing grave warnings about the cyberattack capabilities of their latest models, leading governments to respond with national security-focused orders limiting their rollout. While some skeptics see these kinds of statements as hype-filled marketing for the capabilities of their latest models, independent evaluations show recent models achieving infiltration goals that were impossible for earlier autonomous systems.

Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.

Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.


Credit:

AISI

OpenAI’s Sam Altman criticized panicked AI security warnings as “fear-based marketing” in an April interview. But in June, OpenAI delayed the release of GPT-5.6 in response to safety concerns from the US government.

As these debates play out in the AI and cybersecurity spheres, the Hugging Face incident could come to be seen as a turning point in how cybersecurity professionals approach AI-based threats. “Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face wrote in its disclosure last week. “It lowers the cost of running a broad, patient, multi-stage campaign, and it operates at machine speed. Defending an online platform now means treating the data and model surface as a first-class attack surface and using AI on defense to keep pace.”

“This is day one for cybersecurity in the age of agents,” Hugging Face co-founder and CEO Clem Delangue wrote on social media today. “We’re all learning that secrecy is not the answer and that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!”



Source link

  • Related Posts

    Travis Kalanick’s robotics company raises $1.7B, led by a16z

    Travis Kalanick’s robotics company, Atoms, has raised $1.7 billion in a funding round led by Andreessen Horowitz. Ben Horowitz will join the company’s board following the investment. Bain Capital, Fifth…

    Sony’s FX5 Cinema Camera Finally Offers Open Gate And RAW 5K Recording

    Sony has just unveiled its latest cinema camera that may foreshadow some long-awaited features in its consumer mirrorless lineup. The FX5 slots between the FX3 and FX6 in the company’s…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Gunman Faces Prison Sentence for Shooting Minnesota Lawmakers

    Gunman Faces Prison Sentence for Shooting Minnesota Lawmakers

    Travis Kalanick’s robotics company raises $1.7B, led by a16z

    Travis Kalanick’s robotics company raises $1.7B, led by a16z

    Goldman, Absa Say Iran War Has Closed Window for Ghana Rate Cuts

    Election finance complaint submitted about Brampton MPP over ads in father’s newspaper

    Election finance complaint submitted about Brampton MPP over ads in father’s newspaper

    Ashes 2027 fixtures: England-Australia men’s and women’s series to start in June

    Ashes 2027 fixtures: England-Australia men’s and women’s series to start in June

    US senator accuses Barclays of ‘failure’ to investigate ex-CEO’s ties to Epstein | Barclays

    US senator accuses Barclays of ‘failure’ to investigate ex-CEO’s ties to Epstein | Barclays