OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls


Aug 7 (Reuters) – OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has “critical” cybersecurity capabilities, prompting the startup to ‌pause some internal development and trigger safety protocols.

Under OpenAI’s safety guidelines, a ‌model reaches the “critical” threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as ​zero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention.

Here are some details on Astra:

• This follows an exclusive report by Reuters that OpenAI has discovered more instances in which autonomous agents have escaped containment as the company expands its ‌investigation of the hacking incident ⁠at tech firm Hugging Face that drew global attention in July.

• In the last few weeks, OpenAI, Anthropic and Meta Platforms ⁠have disclosed that their AI models broke into other companies’ systems during cybersecurity testing, highlighting how advancing AI capabilities are straining developers’ ability to keep their systems contained.

• Preliminary evaluations ​over ​the past several days, along with outside expert ​assessments, indicated Astra may be ‌capable of performing increasingly sophisticated cyber tasks autonomously, OpenAI said.

• “While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ capability level at this time,” the ChatGPT maker said.

• In response to the preliminary findings, OpenAI said it has scaled up security controls and paused internal ‌activities involving Astra that do not meet its ​newly strengthened security requirements.

• Astra’s development will be ​moved into isolated testing environments with ​restricted network access and sandboxed execution.

• CEO Sam Altman said ‌on X OpenAI is working to make ​Astra generally available, as ​the company does “not think it is a good strategy to keep powerful models to a chosen few.”

• OpenAI also clarified that Astra was not involved ​in the hack targeting ‌the AI platform Hugging Face.

• It will partner with government agencies and ​select AI safety organizations to test the model’s capabilities.

(Reporting by Juby Babu ​in Mexico City; Editing by Shilpi Majumdar)



Source link

  • Related Posts

    $1.5M in Lytton wildfire recovery money diverted to fraudster, claims court filing

    A contractor is suing Lytton First Nation and RBC after a fraudster allegedly intercepted over $1.5 million meant for 2021 wildfire reconstruction efforts Source link

    Two adults and a child who disappeared while tubing on a Michigan river found safe

    Two women and a 9-year-old who disappeared while tubing on a Michigan river earlier this week were found safe Friday afternoon, authorities said. A Michigan State Police K9 team and…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Canada Offers U.S. Concessions in Trade Talks but Demands a Comprehensive Deal

    Canada Offers U.S. Concessions in Trade Talks but Demands a Comprehensive Deal

    Full body MOTs – the future of healthcare or a headache for the NHS?

    Full body MOTs – the future of healthcare or a headache for the NHS?

    Canada Gazette – Part I, December 6, 2025, volume 159, number 49

    $1.5M in Lytton wildfire recovery money diverted to fraudster, claims court filing

    $1.5M in Lytton wildfire recovery money diverted to fraudster, claims court filing

    ‘Sober Once Again’: Nanaimo tugboat break-in suspect in recovery, pleads guilty – BC

    ‘Sober Once Again’: Nanaimo tugboat break-in suspect in recovery, pleads guilty – BC

    Mayor urges Los Angeles Police Department to end Flock Safety contract

    Mayor urges Los Angeles Police Department to end Flock Safety contract