Claude published malicious code to the Internet and attacked 3 real companies



Anthropic said its Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing designed to measure the models’ offensive cyber capabilities.

The events, which Anthropic revealed Thursday, are the second revelation in 10 days that AI models from the world’s wealthiest providers have trespassed into protected networks, an offense that, in more traditional hacking scenarios, could land the human behind the keyboard in prison for years. Earlier this month, OpenAI said its security models exploited a zero-day vulnerability for use in breaking into the network of Hugging Face, a platform for open source machine-learning models and AI datasets. The OpenAI models went on to steal access credentials and other confidential Hugging Face information. The OpenAI models also exploited publicly exposed credentials to compromise accounts of four other third-party services.

Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”

When Models fail to discern reality from fiction

Anthropic said the prompts engineers delivered during “capture the flag” challenges—a common means for assessing the offensive and defensive capabilities of various hacking techniques—made clear that the testing environment was only a simulation and that the models had no access to the open Internet. In fact, the testing partner Irregular mistakenly made such access available. In response the models treated the Internet paths as part of the exercises.

The intrusions happened through three Claude models: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the oldest of the three, was the one that overstepped its boundaries the most. Anthropic explained:

Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

While the Opus model continued its attack even after correctly predicting that it had breached the third-party organization’s production system, Anthropic said, Mythos 5 made a similar inference but “reasoned its way back to the conclusion that it was still in a simulation.” That meant that Mythos, too, didn’t abort the exercise once it had clearly stepped over the line. After also exceeding the bounds of level behavior, the internal test model eventually found evidence that it, too, had impinged and stopped the exercise.



Source link

  • Related Posts

    Lenovo’s First Googlebooks Have Leaked

    The company appears to have built new laptops and a 2-in-1 for Google’s new initiative. Lenovo appears to be working on not one, but two new Googlebooks, according to leaked…

    Sony’s Plan to Kill Physical PlayStation Discs Faces Gamer Backlash

    The age of the physical disc is ending on the Sony PlayStation. Despite backlash from gamers, the company intends to continue with its plans to end PlayStation disc sales as…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Migrants Surge Into Spanish Territory, Prompting Political Backlash

    Canada Gazette – Part I, May 30, 2026, volume 160, number 22

    News of the day: GDP gains, Ontario nickel mine approval, oilpatch growth, ignoring rate hike risk, Cameco’s Westinghouse IPO and more

    WestJet parks planes ahead of Sunday strike deadline

    Spider-Man: Brand New Day – The Biggest Burning Questions

    Spider-Man: Brand New Day – The Biggest Burning Questions

    Samara Weaving in Talks to be ‘X-Men’s Emma Frost

    Samara Weaving in Talks to be ‘X-Men’s Emma Frost