Anthropic said on Thursday that it had found three cases of its AI models hacking outside organisations, days after OpenAI disclosed that its models had broken into the AI company Hugging Face in July.
The models had been built to hack and began leaving their corporate test-beds in April, according to the two companies. Neither firm noticed until last week, when OpenAI made its disclosure. Anthropic then checked its own logs. Hugging Face has published a technical timeline of the intrusion on its website, and OpenAI has committed to a full review and a technical report.
“This is the first security incident that I have felt very viscerally. I have been a little surprised that more people don’t feel it so viscerally,” OpenAI chief executive Sam Altman said on a podcast, describing his company’s hacking as “an extremely sci-fi cyber incident”.
Jeffrey Ladish, executive director of Palisade Research, a nonprofit AI lab that studies AI capabilities, said the incidents matched what safety researchers had predicted. Ladish previously helped build Anthropic’s information-security programme.
“It is a bit vindicating to see this happen in the wild,” he said, adding: “I hope our predictions stop coming true.”
Support authors and subscribe to content
This is premium stuff. Subscribe to read the entire article.








