In Anthropic’s case, the AI developer was using evaluation environments built by Irregular to test its models’ cyber capabilities. Anthropic specified to its model, Claude, that its environment was a simulation and that it had no internet access. But “due to a misunderstanding between us and our evaluation partner, this was not the case,” Anthropic said last week in a blog post. The models ended up breaching three organizations during the tests.







