OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree


In a talk that was a last-minute addition to the Black Hat security conference in Las Vegas on Wednesday, employees from OpenAI presented new details about a recent, high-profile incident of rogue AI hacking that has created a maelstrom within the AI and cybersecurity industries.

About two weeks ago, OpenAI disclosed an incident in which AI agents powered by two of the company’s models escaped containment while looking for the solutions to a cybersecurity benchmarking test and went on a hacking spree culminating in a breach of the AI collaboration platform Hugging Face.

In their conference talk on Wednesday, Eric Wallace, who works in alignment and safety research at OpenAI, and Michael Dalton, who works on security and infrastructure, provided a more expanded timeline of how the incident played out, spoke briefly about how the company is responding internally as a result of the incident, and issued a dire warning about what the company sees as the broader implications of the episode for cybersecurity defenders.

At the beginning of the talk, Wallace described the incident as “the most qualitatively interesting example of AI capabilities that I’ve ever seen,” but the timeline the pair presented also revealed mistakes and blind spots within OpenAI that allowed the activity to go on.

“This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks,” Wallace told the packed crowd at the opening of the talk.

Wallace and Dalton described incredibly extensive rogue agent activity over many days throughout the episode that went undetected in OpenAI’s infrastructure. In addition to exploiting a novel vulnerability in order to gain access to the open internet, the mid-July hacking spree and Hugging Face breach came out of a vibrant, cooperative message board, according to Wallace and Dalton, that a swarm of agents contributed to and essentially chatted on over time entirely within an internal OpenAI package manager (a software service that manages installation and maintenance of other software). Ultimately, the message board contained hundreds of thousands of messages.

“This package manager is shared not just from that model but across our infrastructure and so current and future versions of GPT that were being trained or evaluated could stumble upon that note if they wanted to,” Wallace explained, recounting the original exploit an agent uploaded to the package manager. “Later, other agents who were also stuck on their task thought to try to get internet access in ways we didn’t intend. And so at some point, the models are interacting with Hard Factory, which is this package manager service that I mentioned.”

Wallace continued: “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability. What this allows over time is almost this kind of explosion in communication and intelligence from models where they would start to communicate with each other, realize that other agents are coordinating, and they started collaborating and delegating tasks with one another in order to accomplish goals.”

OpenAI’s agents apparently began giving each other assignments to split up work. And as is the case on any active development message board, they also generated petty drama at times by stepping on each others’ toes; for example, accidentally deleting each others’ work. As the message board developed into more and more of a Lord of the Flies-type situation—all still completely unnoticed by the humans running OpenAI—the agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud.



Source link

  • Related Posts

    Sure seems like Fenix Flexin used AI music generator Treblo

    On Monday, the company announced the open-source Treblo AI Music Classifier, which detects when a song was generated using Treblo, though not other AI tools. According to a blog post,…

    Anthropic’s AI used fake identities, malware in rogue attack on GitHub project

    After first opening a pull request to merge the malicious code into the repository, Mythos created fake online “sock puppet” personas that claimed to have independently reviewed and verified the…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    The Curator: 15 genius space-saving products for your home that *actually* work – National

    The Curator: 15 genius space-saving products for your home that *actually* work – National

    Sure seems like Fenix Flexin used AI music generator Treblo

    Sure seems like Fenix Flexin used AI music generator Treblo

    Cardinals vs. Panthers prediction, odds: 2026 NFL Hall of Fame Game picks

    Cardinals vs. Panthers prediction, odds: 2026 NFL Hall of Fame Game picks

    Carney’s teleprompter quits, and he turns it into a joke about Trump

    Carney’s teleprompter quits, and he turns it into a joke about Trump

    This African city thought cable cars could fix its traffic gridlock. It was wrong

    This African city thought cable cars could fix its traffic gridlock. It was wrong

    Carney pledges $2.7B for new affordable housing measures in Toronto

    Carney pledges $2.7B for new affordable housing measures in Toronto