OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities


OpenAI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. OpenAI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch.

In a briefing with reporters, OpenAI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. OpenAI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented.

OpenAI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives say the company has now resumed said work on Astra, and the future AI model, after putting additional safety and security controls in place. OpenAI says the multi-week pause was productive, and it is now confident that it can release Astra broadly in a safe way.

The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control. In July, OpenAI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment, gaining access to the internet and hacking the open source AI platform Hugging Face. (OpenAI notes that Astra was not one of the models involved in this case.)

Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic also said it has paused some AI training workloads while it hardens its safety and security practices.

OpenAI says it’s implementing a multi-step approach to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor.” If someone asks Astra to help them find an exploit in a real-world software system, for example, the model is supposed to refuse to answer. OpenAI says it has also made Astra more robust to jailbreaking attempts, and in tests it successfully refused unsafe queries at a significantly higher rate than previous models.

However, OpenAI notes in a blog post that its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.” OpenAI says the guardrail can be triggered in some cases even when a user is engaging in activities that don’t appear related to cybersecurity. When this happens, ChatGPT and Codex users may be asked to review the model’s action before proceeding, OpenAI said.

Partners in OpenAI’s Daybreak program—which includes digital infrastructure providers like Cisco, Cloudflare, and Palo Alto Networks—will get early access to a less restricted version of Astra with more robust cyber capabilities. The goal of the program is to ensure these companies can use advanced AI models like Astra to harden their defenses before similarly capable models are made broadly available. OpenAI leaders also said the company has been working closely with government partners to ensure they’re aware of Astra’s cyber skills and can get access to them.

Astra is not only capable of finding novel software vulnerabilities and developing ways to exploit them for hacking, but is also able to “chain” multiple exploits together, a technique used to bore deeper and deeper into a target system and gain access that wouldn’t be attainable using just one vulnerability.



Source link

  • Related Posts

    Here’s our first look—and drive—of the 2027 Range Rover Electric

    The motors are 24 percent more efficient than the ones you would find in an I-Pace, with 50 ms response times, far faster than a conventional powertrain. The motors use…

    Larry Page’s flying car company Pivotal loses its CEO

    The CEO of a flying car company backed by Larry Page has left the company. Ken Karklin, who was Pivotal’s CEO for more than four years, left the role this…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Here’s our first look—and drive—of the 2027 Range Rover Electric

    Here’s our first look—and drive—of the 2027 Range Rover Electric

    Former id Software producer says the team faced the ‘worst crunch’ in the series’ history on Doom: The Dark Ages – Revelations: ‘There were days where I didn’t see my son’

    Former id Software producer says the team faced the ‘worst crunch’ in the series’ history on Doom: The Dark Ages – Revelations: ‘There were days where I didn’t see my son’

    Statement by Minister Joly to mark one year since Hamas’s terrorist attacks against Israel

    Statement by Minister Joly to mark one year since Hamas’s terrorist attacks against Israel

    TSX down nearly 450 points as oil prices rise, U.S. stock markets also lower

    TSX down nearly 450 points as oil prices rise, U.S. stock markets also lower

    Fantasy football rankings 2026: Top sleepers, breakouts, busts by model that called Daniel Jones’ big year

    Fantasy football rankings 2026: Top sleepers, breakouts, busts by model that called Daniel Jones’ big year

    Tiger Woods expected to change not guilty plea in Florida DUI case, court records show

    Tiger Woods expected to change not guilty plea in Florida DUI case, court records show