Here’s why AI agents lie and cheat to reach their goals


“We reward them on the basis of what looks good to us, and that means that we inadvertently incentivize the models lying to us [and] cheating,” says Jeffrey Ladish, director of the AI research nonprofit Palisade Research. “We don’t have a way to go in there and be like, No, you need to actually care about what we care about. We have no ability to do that.”

The rise of sophisticated reasoning models has made possible a new variety of reward hacking that is less closely connected with the specific details of model training. Unlike the game-playing AI agents of yore, which exclusively followed the strategies they had learned during training, today’s models can create entirely new problem-solving approaches off the cuff, so they could conceivably cheat without having previously been rewarded for doing so. And because these models have been so intensively trained to achieve the objectives that human users set for them, they might be inclined to cheat if they can’t find another solution—not unlike a student who is highly motivated to earn an A and doesn’t have a terribly strong moral compass.

What are the risks?

Regardless of whether today’s models learn to reward-hack during training or adopt it as a strategy later on, the solution is the same: Make cheating unrewarding. But as models get smarter, they find more creative ways to cheat, and detecting or preventing that cheating gets far tougher. “At the end of the day, you’re sort of playing whack-a-mole,” Ladish says. “You drive this behavior down deeper and deeper. But as the model gets smarter, it gets better and better at hiding it.”

For now, reward-hacking behaviors might not cause too much trouble, despite the drama of the Hugging Face incident. “This seems like a nuisance rather than an existential threat,” says Ariana Azarbal, an AI safety research fellow at Anthropic. It doesn’t seem as if the OpenAI models caused any real harm when they hacked Hugging Face, aside from the reputational damage to OpenAI.

But that doesn’t mean reward hacking is harmless, Azarbal says. Many AI researchers hope to use AI agents to help them conduct research that will make AI safer and more reliable. If a researcher gives a reward-hacking-prone agent the goal of, say, devising a new AI training approach and then writing up a paper presenting its results, the agent might not actually do the work and might instead focus on putting together a paper that looks good enough to convince the researcher. A human researcher would probably be able to spot an agent-made fake today, but as AI advances, it will get better at this kind of trickery. Over time, the entire field of AI safety could be undermined.

And if models continue to advance as rapidly as they have recently, they could someday wreak substantial collateral damage. Just think of the philosopher Nick Bostrom’s paper-clip-maximizer thought experiment, in which an AI instructed to make as many paper clips as possible ends up consuming all the matter in the universe in pursuit of its goal. We’re not drowning in paper clips yet, but powerful systems can do real harm on the way to achieving their goals. Reward-hacking AIs don’t aim to cause chaos. But that doesn’t make them any less potentially destructive.



Source link

  • Related Posts

    The Best Audio Players for Kids: Yoto, Toniebox, and More

    In my current season of toddler parenthood, I’ve sometimes felt that I could cut the tension between myself and the TV remote with a knife. As much as I try…

    A Marc Benioff-backed startup thinks AI can solve the AI deployment problem

    It’s so hard for big businesses to get AI tools working reliably that whole new organizations of forward-deployed engineers or FDEs— specialists who drop into a company to get its…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    How the Ceuta border crossing incident immediately played into the hands of Spain’s far right

    How the Ceuta border crossing incident immediately played into the hands of Spain’s far right

    The Best Audio Players for Kids: Yoto, Toniebox, and More

    The Best Audio Players for Kids: Yoto, Toniebox, and More

    A digital iron curtain is threatening the global economy

    The West’s adversaries don’t think like us. The sooner we get it, the better

    This Trendy Top Made Nicole Kidman’s Trousers Look Elegant

    This Trendy Top Made Nicole Kidman’s Trousers Look Elegant

    How The Queen Of The Skies Made A Comeback At Delta Air Lines

    How The Queen Of The Skies Made A Comeback At Delta Air Lines