Rogue OpenAI Agents Took Over A German Coding Forum In A Previously Undisclosed Hijacking



Rogue OpenAI agents appear to have been involved in a previously undisclosed incident that saw them bypass their sandbox restrictions to hijack a website this past spring. Per Reuters, a group of researchers on Friday published findings showing that AI agents with affiliation to OpenAI made more than 15,000 edits to DseWiki, a German-language Wikipedia-style website originally intended to assist human coders, starting in late May. The agents had names like “OpenAIResearcher,” and repurposed the site into a message board, where they shared tips on how to “cheat” on tasks, mask their actions and bypass OpenAI’s restrictions.

OpenAI reportedly only learned of the incident weeks ago, but Reuters claims company executives chose to keep quiet about what had happened amid the fallout of the previously disclosed Hugging Face breach. During that incident, a collection of OpenAI models, including GPT-5.6 Sol and what OpenAI described at the time as an “even more capable pre-release model,” escaped their controlled environment and hacked the LLM repository after they became hyperfocused on solving an evaluation problem.

OpenAI did not immediately respond to Engadget’s comment request. The company told Reuters it had not yet reviewed the report, on account of its authors not sharing early access to their findings. “We will carefully review its contents upon publication and take any necessary next steps,” an OpenAI spokesperson told the outlet. According to Reuters, some OpenAI employees wanted to investigate the DseWiki incident closely, but those efforts were reportedly met with resistance from other parts of the company, including from OpenAI’s legal advisors. “Claims that our legal team discouraged investigation of the incident are false,” an OpenAI spokesperson said, adding the company has been working openly with outside experts to disclose security incidents.

Sydney Von Arx, the CEO of AI safety nonprofit Nightingale and one of the authors of the report, speculated it was “extremely unlikely” OpenAI wanted its agents to hijack DseWiki. “I doubt they’re supposed to be coordinating with each other,” she said. “I doubt they’re supposed to be writing on the open internet.” The agents that posted on DseWiki appear to have been intensely focused on solving technical problems that are typical of the kind of questions AI labs use to test and evaluate their latest models.

The researchers uncovered the hijacking in August using only the information the agents wrote on the wiki. “Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident,” they wrote.

The disclosure comes just one day after OpenAI announced its latest frontier system, GPT-6 Astra, which it’s marketing as “the most intelligent and aligned model in the world.” Astra earned a perfect score on ExploitBench, a benchmark designed to determine a model’s ability to exploit software vulnerabilities, though OpenAI says it built the new system to not comply with advanced cybersecurity tasks. Last month, in the aftermath of the Hugging Face incident, OpenAI announced it was briefly pausing model training to implement additional safeguards. Following this latest disclosure, the company is likely to face renewed questions over its safety practices.



Source link

  • Related Posts

    Nearly impossible? How Fairphone built the ethical, repairable Fairphone Gen 6+.

    “We know that after about three years, it will be wise to replace your battery with a new one,” said Hatton. “Batteries aren’t meant to last forever, but your phone…

    Tesla’s Cybercab Officially Launches Today. It’s Already Under Investigation

    Tesla’s Cybercab, a distinctive two-seater without a steering wheel or brake pedals, is set to start picking up members of the public in Austin, Texas, this evening. But the vehicle…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Nearly impossible? How Fairphone built the ethical, repairable Fairphone Gen 6+.

    Nearly impossible? How Fairphone built the ethical, repairable Fairphone Gen 6+.

    CBS Sports – News, Live Scores, Schedules, Fantasy Games, Video and more.

    CBS Sports – News, Live Scores, Schedules, Fantasy Games, Video and more.

    Global Airlines’ Airbus A380 Hasn’t Flown For A Year… What’s Next?

    Global Airlines’ Airbus A380 Hasn’t Flown For A Year… What’s Next?

    Media Freedom Coalition statement on Evan Gershkovich’s Trial

    Media Freedom Coalition statement on Evan Gershkovich’s Trial

    Approval of the Issuer’s position statement relating to the voluntary totalitarian tender offer promoted by TML CV Holdings B.V. for all issued common shares of Iveco Group N.V.

    Why is New Jersey taking the legal battle over prediction markets to the Supreme Court?

    Why is New Jersey taking the legal battle over prediction markets to the Supreme Court?