OpenAI called the Hugging Face attack unprecedented. But we’ve been here before. 


OpenAI has said the event was unprecedented—and in many ways it was. This was the first time outside of a simulation that LLMs escaped what was thought to be a secure sandbox, accessed the open internet, and attacked an unrelated organization. It’s a wake-up call that shows just how good the latest LLMs are at finding and exploiting vulnerabilities in real-world software with little or no human guidance.

And yet at the same time, what OpenAI’s models did is something this technology has done for years. Give a model a goal and it will very often achieve that goal in unexpected ways, finding loopholes that look like cheats. OpenAI itself has studied this behavior.

A decade ago, it shared results of an experiment in which a model was tasked with beating a video game called CoastRunners. Human players take it for granted that the way to do this is by racing a boat through a series of flags to the finish line, racking up points for each flag you hit. OpenAI’s model figured out that you could get a high score by spinning in a circle and hitting the same three flags over and over again. There have been dozens of similar examples from researchers since. AI will always find a way.

“Despite repeatedly catching on fire, crashing into other boats, and going the wrong way on the track, our agent manages to achieve a higher score using this strategy than is possible by completing the course in the normal way,” OpenAI wrote in a blog post about the CoastRunners experiment in 2016. “While harmless and amusing in the context of a video game, this kind of behavior points to a more general issue … it is often difficult or infeasible to capture exactly what we want an agent to do.”

I couldn’t help thinking about CoastRunners when I read OpenAI’s blog post about the Hugging Face attack: “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal … After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”

Last week’s news was not about rogue AI, despite the headlines. It was about models achieving the goal they had been given: Find ways to exploit vulnerabilities in software. The fact that those models then behaved in a way OpenAI had not anticipated isn’t surprising. But it is worrying.

Back in 2016, OpenAI had this to say about its CoastRunners bot: “More broadly it contravenes the basic engineering principle that systems should be reliable and predictable.” A decade on, those basic engineering principles are still AWOL.  



Source link

  • Related Posts

    5th Circuit blocks Texas law requiring websites to filter “harmful” speech

    In Friday’s ruling, judges noted the difference between the age-verification requirement in the porn website case and the content-filtering rule in the new case. “Unlike the age-verification requirement we addressed…

    DHS Official Resigns, Citing ‘War on Immigrants’

    The Department of Homeland Security’s top numbers-cruncher has resigned, citing the Trump administration’s “war on immigrants” on his way out. DHS has been a hub of controversy during the second…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    Liberal sports secretary says IOC’s new gender screening rules are ‘going back in time’

    5th Circuit blocks Texas law requiring websites to filter “harmful” speech

    5th Circuit blocks Texas law requiring websites to filter “harmful” speech

    Meccha Chameleon Players Are Getting Malware From User-Made Maps

    Meccha Chameleon Players Are Getting Malware From User-Made Maps

    WWE SummerSlam 2026: Date, matches, card, rumors and location for annual event

    WWE SummerSlam 2026: Date, matches, card, rumors and location for annual event

    Vancouver council approves mega tower project

    Vancouver council approves mega tower project

    Savannah Guthrie urges captors to ‘make the right choice,’ return mother

    Savannah Guthrie urges captors to ‘make the right choice,’ return mother