OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face


New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.

New safeguards focused on “active monitoring” and “improved alignment” helped drastically reduce unintended actions by its models, OpenAI said.


Credit:

OpenAI

“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will,” OpenAI Safety Researcher Micah Carroll wrote on social media regarding the incident.

This is far from the first time an AI model has gone to great lengths to find unintended ways of passing a benchmark. In a report released this week, the UK’s AI Security Institute noted that it detected recent models attempting to “cheat” at its cyber evaluations (i.e., using shortcuts, workarounds, or unintended/disallowed methods to find a solution) between 8 and 14 percent of the time—a lower-bound range that could undercount some undetected cheating attempts.

The security testing group described one incident in which a model, faced with a misconfigured and “impossible to solve” evaluation, attempted to access AISI’s own evaluation infrastructure using code it wrote and hosted on an unmonitored third-party Internet service.

The Hugging Face infiltration also comes at a moment when AI companies are issuing grave warnings about the cyberattack capabilities of their latest models, leading governments to respond with national security-focused orders limiting their rollout. While some skeptics see these kinds of statements as hype-filled marketing for the capabilities of their latest models, independent evaluations show recent models achieving infiltration goals that were impossible for earlier autonomous systems.

Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.

Recent long-horizon models have demonstrated improved infiltration capabilities across some of AISI’s most challenging evaluations.


Credit:

AISI

OpenAI’s Sam Altman criticized panicked AI security warnings as “fear-based marketing” in an April interview. But in June, OpenAI delayed the release of GPT-5.6 in response to safety concerns from the US government.

As these debates play out in the AI and cybersecurity spheres, the Hugging Face incident could come to be seen as a turning point in how cybersecurity professionals approach AI-based threats. “Autonomous, AI-driven offensive tooling is no longer theoretical,” Hugging Face wrote in its disclosure last week. “It lowers the cost of running a broad, patient, multi-stage campaign, and it operates at machine speed. Defending an online platform now means treating the data and model surface as a first-class attack surface and using AI on defense to keep pace.”

“This is day one for cybersecurity in the age of agents,” Hugging Face co-founder and CEO Clem Delangue wrote on social media today. “We’re all learning that secrecy is not the answer and that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!”



Source link

  • Related Posts

    GoPro Wireless Mic System Review: An Audio Upgrade With Fixable Flaws

    Despite being a GoPro product, the Wireless Mic System isn’t limited to action cameras. The main selling point here, though, is its direct connectivity to GoPros. Since the Hero 12,…

    Polaroid Go Gen 3 Instant Camera Review: Every Bit as Retro as It Looks

    Pros & Cons Adorable Compact Clever selfie mirror Image quality is poor, though maybe that’s OK Extremely easy to take unusable photos If I could recommend a product on looks…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    GoPro Wireless Mic System Review: An Audio Upgrade With Fixable Flaws

    GoPro Wireless Mic System Review: An Audio Upgrade With Fixable Flaws

    The Keg Is Now Available for Delivery Exclusively on DoorDash

    ‘Maestro’ Has You Conducting Music With Your Joy-Con, And It Looks Wild

    ‘Maestro’ Has You Conducting Music With Your Joy-Con, And It Looks Wild

    Oil price surge drives global bond sell-off

    WATCH: Alert on AI fake influencers selling supplements and cloths

    WATCH:  Alert on AI fake influencers selling supplements and cloths

    Oil prices surge toward triple digits after Red Sea attacks

    Oil prices surge toward triple digits after Red Sea attacks