Here’s what actually happened in OpenAI’s Australian gov’t server hack


Of course, we can’t rely on an LLM to have that same sense of proportionality (or any inherent sense of worry about legal implications) in responding to a prompt. Without explicit instructions on what is and is not allowed or justified, an AI agent with suitable resources will try every plausible avenue to satisfy the user’s request as best it can.

That red line looks more like a red suggestion to me…

Credit:
Getty Images

That red line looks more like a red suggestion to me…


Credit:

Getty Images

OpenAI says the internal testing in this case was done “without the full set of safeguards used in our publicly available products.” Given that lack of constraints, the agent was arguably working as intended, in a sense, by using every tool available to generate an answer to the prompt.

At the same time, OpenAI says the agent in the test was “supposed to answer these questions using publicly published statistics” and “took actions that we had not authorized it to take” to get that information. From the outside, it’s hard to know just how strong OpenAI’s attempts to deny “authorization” were, in practice. It’s plausible that OpenAI’s agent here disregarded a relatively simple “anti-hacking” directive in its system prompt so it could better give a complete answer that satisfies a direct prompt from the user, for instance.

In public analyses of multiple “misalignment” incidents published earlier this month, OpenAI identified multiple instances of “reward hacking,” where an agent resorted to extreme methods to generate a better answer to a user’s prompt. The company said it had recently taken steps to prevent this kind of reward hacking by adding explicit punishments for misaligned behavior to the system’s reward function.

With the benefit of hindsight, it’s hard to see why those kinds of protections were not in place in June, and whether they could have prevented a potential international incident in this case.



Source link

  • Related Posts

    Do You Actually Need a Separate Antivirus, VPN and Scam Protection? McAfee’s All-in-One Plans Have You Covered

    Quick Answer: Antivirus software protects your devices from malware and other threats, a VPN makes your internet connection private and scam protection helps flag suspicious texts, emails and videos before…

    Continue reading
    OpenAI launches Dots, its Muse competitor

    OpenAI is responding to Meta’s buzzy Muse AI with agentic helpers of its own: Dots. During its DevDay keynote on Tuesday, OpenAI announced that Dots will serve as always-on AI…

    Continue reading

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    10 Essential SNES Games Still Missing From Switch Online

    10 Essential SNES Games Still Missing From Switch Online

    Do You Actually Need a Separate Antivirus, VPN and Scam Protection? McAfee’s All-in-One Plans Have You Covered

    Do You Actually Need a Separate Antivirus, VPN and Scam Protection? McAfee’s All-in-One Plans Have You Covered

    United Airlines Launches New 11-Hour Nonstop Route In 3 Months: See Map Now

    United Airlines Launches New 11-Hour Nonstop Route In 3 Months: See Map Now

    Regula and Bernini.AI Enable Identity Verification at Key Moments in AI-Powered Customer Conversations

    Hegseth targets military’s top brass for deeper cuts: Officials

    Hegseth targets military’s top brass for deeper cuts: Officials

    OpenAI launches Dots, its Muse competitor

    OpenAI launches Dots, its Muse competitor