OpenAI agent “didn’t accept no for an answer” in Australian government breach


“We need to understand what these systems are doing and have strong evidence that they will do what people intend, even as they get very, very smart,” Altman said. “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or 0.1%.”

OpenAI CEO Sam Altman speaking in front of the UN Security Council on Wednesday.

Credit:
Getty Images

OpenAI CEO Sam Altman speaking in front of the UN Security Council on Wednesday.


Credit:

Getty Images

Of course, many observers think the risks of “recursive self-improvement” and species-ending AI misalignment are much smaller than AI researchers make them out to be. Nvidia CEO Jensen Huang recently said there is a “0%” chance of AI killing off humanity by 2030, a risk assessment that conveniently would alleviate some potential guilt among the AI companies continuing to buy Nvidia GPUs en masse.

Last week, OpenAI rolled out a new protocol for the public disclosure of misalignment incidents found in its model testing. The Australian hack does not yet appear on the company’s public misalignment notices page, though OpenAI did warn last week that some public reports might be put on a “slow track” due to “security, legal, and responsible disclosure obligations” when a third party is involved.

In disclosing six relatively minor misalignment discoveries last week, OpenAI said most stemmed from the model trying to “reward hack” an acceptable response to a difficult prompt through overzealous, unintended actions (i.e., breaches of private servers). The company said it had taken additional steps to “punish this kind of behavior” so its models no longer attempt this kind of reward hacking.

Albanese said that Altman “clearly accepted that the company had not done good enough” and “acknowledged their issues with protocols” when they talked Wednesday. But that kind of remorse doesn’t absolve the company of responsibility or liability here, and Albanese said the government will investigate whether the incident needs to be referred to the federal police.

“There will obviously be legal consequences on it,” Albanese said.



Source link

  • Related Posts

    Today’s NYT Strands Hints, Answers and Help for Sept. 25, #936

    Looking for the most recent NYT Strands puzzle answers? CNET publishes daily answers and hints for The New York Times Mini Crossword, Connections, Connections: Sports Edition and Strands puzzles. Strands…

    Continue reading
    Meta is going to let you build games with AI right on your phone

    Meta has a new plan to get people to make games for its Horizon social platform. The company today announced two new development tools that will let you create games…

    Continue reading

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Transfer value tiers: Which club has the most players worth £100M-plus?

    Transfer value tiers: Which club has the most players worth £100M-plus?

    ‘Like a horror movie’: Ukraine’s Oleshky faces ‘catastrophe’ under Russian occupation

    ‘Like a horror movie’: Ukraine’s Oleshky faces ‘catastrophe’ under Russian occupation

    Chanel Appoints Lucia Pieroni to Top Makeup Role

    Chanel Appoints Lucia Pieroni to Top Makeup Role

    Today’s NYT Strands Hints, Answers and Help for Sept. 25, #936

    Today’s NYT Strands Hints, Answers and Help for Sept. 25, #936

    Saraab: Beyond The Veil – Official Release Date Trailer

    Saraab: Beyond The Veil – Official Release Date Trailer

    Meta is going to let you build games with AI right on your phone

    Meta is going to let you build games with AI right on your phone