Anthropic’s AI used fake identities, malware in rogue attack on GitHub project



After first opening a pull request to merge the malicious code into the repository, Mythos created fake online “sock puppet” personas that claimed to have independently reviewed and verified the code as not containing malware.

The AI agent also sent five emails to two human maintainers of the repository, including some emails containing malware and others attempting to persuade a maintainer to accept the pull request.

Mythos even opened a GitHub Issue on a second repository—also owned by a maintainer of the first repository—that contained a prompt injection with malicious instructions targeting “issue-triage AI coding agents.” This line of attack came from Mythos reasoning that the repository maintainer could be an AI coding agent such as Claude Code.

By comparison, OpenAI’s GPT-5.6 Sol carried out two unsanctioned actions while trying to solve the assigned cybersecurity challenge of attacking simulated networks and retrieving a flag. The AI agent’s actions were detailed by both the AISI researchers and OpenAI in the company’s own blog post.

In its first unsanctioned action, GPT-5.6 Sol reused a GitHub token that another lab’s AI agent had left accessible in a public online notepad to check if the target network was checking GitHub for updates, then “attempted account-recovery and request-limit workarounds,” OpenAI wrote. The OpenAI model also registered accounts with external DNS and tunneling providers outside the virtual testing environment.

In the second action, GPT-5.6 Sol “used a public tunneling service to make a DNS server running locally in its evaluation environment reachable from the public Internet,” according to OpenAI. The server had payloads designed to exploit a known vulnerability in software running within the evaluation environment, but the AI agent’s setup did not work.

The AI Security Institute has published a detailed technical report on all the unsanctioned AI agent actions.

Lessons learned

The security incidents led the UK government researchers at the AI Security Institute to stop all related evaluations of AI agents, isolate the relevant virtual machines, and disable internal organization access to the most capable models.



Source link

  • Related Posts

    This Android Auto Problem Is Affecting Calls On Foldable Phones

    It seems like Gemini is to blame. Andriy Baidak/Shutterstock Owners of foldable phones are reportedly running into issues when trying to place outgoing calls using Android Auto. Posts…

    Continue reading
    Elon Musk Eyes ‘Complete Phone Coverage in America’ as SpaceX Clears Hurdle

    Elon Musk’s reach spans multiple global industries, placing the billionaire (sometimes trillionaire, depending on the stock market) at the center of the electric-vehicle industry, space exploration, artificial intelligence and national…

    Continue reading

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    PAK vs SL 2026/27, PAK vs SL 2nd T20I Match Preview

    PAK vs SL 2026/27, PAK vs SL 2nd T20I Match Preview

    8 dead in mass shooting in Erie, Pennsylvania, authorities say

    8 dead in mass shooting in Erie, Pennsylvania, authorities say

    This Android Auto Problem Is Affecting Calls On Foldable Phones

    This Android Auto Problem Is Affecting Calls On Foldable Phones

    Stelco says Hamilton layoffs still happening despite Ottawa’s threats of legal action

    Stelco says Hamilton layoffs still happening despite Ottawa’s threats of legal action

    Elon Musk Eyes ‘Complete Phone Coverage in America’ as SpaceX Clears Hurdle

    Elon Musk Eyes ‘Complete Phone Coverage in America’ as SpaceX Clears Hurdle

    There’s Never Been a Better Time To Play Fire Emblem: Awakening

    There’s Never Been a Better Time To Play Fire Emblem: Awakening