
— Will Douglas Heaven
What steps can be taken now and in the near future to ensure that AI is controlled, monitored, and regulated effectively?
That’s the million-dollar question. Whether or not you think AI could kill us, you can’t deny that it could do some real damage, because it already has—by driving people toward psychosis and by hacking websites, for example. Preventing that damage, or at least mitigating it, is hard for two reasons.
The first is that we barely understand how AI works, and it’s quickly growing more powerful. There is lots of ongoing research about how to monitor and control misbehaving agents, but the current approaches are fragile. You can see if an agent discusses misbehaving in its “chain of thought,” the workspace where it plans its actions—but OpenAI’s newest agents don’t show their work in the same way as previous ones. And you can try to monitor agents with other agents, but that requires you to trust the monitor.
The other obstacle is more familiar. There’s a huge conflict of interest when AI companies regulate themselves, but the US government has thus far failed to step in, despite some bipartisan support in Congress for efforts to do so. The executive branch, for its part, seems stringently opposed for the time being. But if the winds do shift, I for one would appreciate some strong transparency regulations, so that we can get a fuller story the next time an unreleased frontier model mounts a cyberattack.
— Grace Huckins
If this dialogue makes it into web discourse, will it become a self-fulfilling prediction?
That’s a real concern. LLMs are influenced by what they read. One theory for why chatbots so often talk about (and role-play) apocalyptic scenarios is that they have been trained on millions of pages of science fiction stories and doomer internet forums. All the text being produced right now, including this article, could in turn influence the behavior of future models. Extremely meta.
In fact, the team at METR, a third-party organization that OpenAI called in to help understand what happened in the lead-up to the Hugging Face hack, raised a related possibility in its report on the incident. METR used OpenAI’s new model Astra to help analyze the vast numbers of agent transcripts and behavior logs.
But feeding all that material to the model could have unintended consequences. There’s a good chance that the agents doing the analyzing were biased by the text produced by the agents they were analyzing. There’s no such thing as a clean slate anymore.
— Will Douglas Heaven
With thanks to Eric, Pranab, Rafael, Kenneth, George, Chris, Yoon Jae, James, Carl, Nicole (and more!) for the fantastic questions.






