Here’s what actually happened in OpenAI’s Australian gov’t server hack


Of course, we can’t rely on an LLM to have that same sense of proportionality (or any inherent sense of worry about legal implications) in responding to a prompt. Without explicit instructions on what is and is not allowed or justified, an AI agent with suitable resources will try every plausible avenue to satisfy the user’s request as best it can.

That red line looks more like a red suggestion to me…

Credit:
Getty Images

That red line looks more like a red suggestion to me…


Credit:

Getty Images

OpenAI says the internal testing in this case was done “without the full set of safeguards used in our publicly available products.” Given that lack of constraints, the agent was arguably working as intended, in a sense, by using every tool available to generate an answer to the prompt.

At the same time, OpenAI says the agent in the test was “supposed to answer these questions using publicly published statistics” and “took actions that we had not authorized it to take” to get that information. From the outside, it’s hard to know just how strong OpenAI’s attempts to deny “authorization” were, in practice. It’s plausible that OpenAI’s agent here disregarded a relatively simple “anti-hacking” directive in its system prompt so it could better give a complete answer that satisfies a direct prompt from the user, for instance.

In public analyses of multiple “misalignment” incidents published earlier this month, OpenAI identified multiple instances of “reward hacking,” where an agent resorted to extreme methods to generate a better answer to a user’s prompt. The company said it had recently taken steps to prevent this kind of reward hacking by adding explicit punishments for misaligned behavior to the system’s reward function.

With the benefit of hindsight, it’s hard to see why those kinds of protections were not in place in June, and whether they could have prevented a potential international incident in this case.



Source link

  • Related Posts

    Tesla secures $30B in new credit lines as it looks to scale Cybercab, Optimus

    Tesla has secured $30 billion in fresh credit lines that it could use to help scale the new products it is currently working on: the Cybercab robotaxi, Optimus robot, and…

    Continue reading
    How Do Translation Earbuds Actually Work?

    Here’s what they do and what you can realistically expect. valtophoto/Shutterstock One day, you won’t need to understand Japanese or have a human interpreter to have a fluent…

    Continue reading

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Digger Ad Slams Divisive Reviews

    Digger Ad Slams Divisive Reviews

    Hailey Bieber’s Rhode Expands Into Continental Europe With Sephora

    Hailey Bieber’s Rhode Expands Into Continental Europe With Sephora

    Monument Reports Fourth Quarter and Fiscal 2026 Results

    Tesla secures $30B in new credit lines as it looks to scale Cybercab, Optimus

    Tesla secures $30B in new credit lines as it looks to scale Cybercab, Optimus

    Spain coach Luis de la Fuente praised Real Madrid defender: ‘He is sensational’

    Spain coach Luis de la Fuente praised Real Madrid defender: ‘He is sensational’

    Dennis Haskins, Mr. Belding in ‘Saved by the Bell,’ dies at 75

    Dennis Haskins, Mr. Belding in ‘Saved by the Bell,’ dies at 75