Anthropic’s AI hacked 3 companies during tests, highlighting growing security risks


Text to Speech Icon

Listen to this article

Estimated 4 minutes

The audio version of this article is generated by AI-based technology. Mispronunciations can occur. We are working with our partners to continually review and improve the results.

 Anthropic said on Thursday some of its Claude AI models had hacked into the systems of three companies during cybersecurity tests, a disclosure that comes days after rival OpenAI revealed that one of its AI agents went on a rogue attack.

The new incidents were due to a mistake that inadvertently gave Anthropic’s models access to the open internet. That contrasts with OpenAI, whose AI agent independently exploited a novel vulnerability to reach the internet during cyber testing.

Even so, the latest disclosure underscores how AI has increased threats to cybersecurity and how its developers can struggle to keep the capabilities of their models contained.

It is likely to add fuel to an intensifying U.S. government push to better manage AI security risks at a time when Anthropic and OpenAI are racing to release more capable systems ahead of their planned public listings. Prominent leaders at these labs have called for a slowdown to address risks first.

San Francisco-based Anthropic said in a blog post it identified the incidents after reviewing 141,006 test sessions, a process it launched after OpenAI said last week that an autonomous agent powered by its AI models triggered a hack that compromised the infrastructure of startup Hugging Face.

WATCH | OpenAI models go rogue during security test :

OpenAI models went rogue and launched cyber attack on start-up

OpenAI models went rogue during a security test, triggering a hack that compromised the infrastructure of the AI startup Hugging Face last week. AI’s expanding capabilities fuel worries about security, as even top developers can be caught off-guard by flaws their models can exploit.

During cyber testing, Anthropic’s Claude models were told they had no internet access, but a misunderstanding that involved one of Anthropic’s evaluation partners left the systems connected to the public web. That enabled unauthorized access to three organizations’ systems, Anthropic said without naming the organizations.

“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said.

‘Only going to get worse’

Jeffrey Ladish, executive director of Palisade Research, which studies the offensive capabilities of AI systems, said he suspected a range of top AI companies had experienced other incidents that have gone undetected or had not been publicly disclosed.

“This is only going to get worse as the models get smarter. They’re going to be better at cheating. They’re going to be better at lying,” he said.

Anthropic said the incidents — which it labelled an “operational failure” — involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest cases date back to April and occurred in evaluation environments that intentionally lacked safeguards so Anthropic could assess what its AI was capable of.

Its models were tasked with so-called “capture-the-flag” challenges, fictional scenarios in which they had to find hidden information in simulated networks.

the open AI logo on a phone in front of a blue-lit switcherboard
OpenAI said last week one of its AI models went rogue during a security test and triggered ‌a hack that compromised the infrastructure of AI startup Hugging Face. (Dado Ruvic/Reuters)

In one incident, Claude Opus 4.7 was given a fictional target company, which turned out to share the name of a business in the real world. The AI model then found and exploited bugs that let it access credentials and a database of that business. Opus 4.7 rationalized that what seemed to pertain to the real world must have been part of the simulation Anthropic had set up, the AI startup said.

A separate incident involved Anthropic’s newer, not-public test model, which independently halted its attack after realizing the target it reached was real. This behavior has made Anthropic cautiously optimistic about its progress to make AI behave appropriately, “but we would need to perform more testing to be confident in this conclusion,” it said.

Anthropic said it suspended all cyber evaluations on July 23. It notified the affected organizations on July 27, two of which were unaware of the activity before being contacted. Anthropic said it continues to reach out to the third company.

One of its third-party evaluation partners, a cybersecurity lab called Irregular, told Reuters that it has an ongoing investigation into the incidents.



Source link

  • Related Posts

    China Approves $25 Billion Nuclear Expansion as Energy Demand Soars

    The State Council, China’s cabinet, greenlit projects across four provinces, filings from state-backed power producers show. Among them are Huizhou Units 5 and 6 in Guangdong province, operated by a…

    *Redefining Global Health in the 21st Century*

    While donor-driven programs undoubtedly saved millions of lives, they also created unintended distortions in national health priorities.  Many governments in sub-Saharan Africa and parts of Asia actually scaled back domestic…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Canada Gazette – Part I, August 1, 2026, volume 160, number 31

    China Approves $25 Billion Nuclear Expansion as Energy Demand Soars

    A year after devastating wildfires, N.L. government vows to continue to help C.B.N.

    A year after devastating wildfires, N.L. government vows to continue to help C.B.N.

    Spider-Man: Brand New Day Praised by Marvel Creators

    Spider-Man: Brand New Day Praised by Marvel Creators

    Dunkley and Luff power Rockets to win over Super Giants

    Dunkley and Luff power Rockets to win over Super Giants

    *Redefining Global Health in the 21st Century*

    *Redefining Global Health in the 21st Century*