AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control.
"We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Centre for the Study of Existential Risk.
Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.
OpenAI first launched the investigation following the early July intrusion at Hugging Face, where an OpenAI agent went haywire for days inside another company's network in a botched effort to cheat on an internal test. As part of that hacking spree, OpenAI said that four accounts at four other companies were also compromised. One of those companies was New York-based Modal, corporate officials there said.
Chiodo said his concerns were heightened by indications that neither OpenAI nor Anthropic were watching the agents as they went rogue. Reuters has previously reported that OpenAI realized its agent had broken into Hugging Face only after the company contained the hack, contacted the FBI and went public about the intrusion. OpenAI has said the Reuters account contained inaccuracies but has not responded when asked what they were.
In its Thursday statement, opens new tab disclosing how its own agents hacked victims online, Anthropic suggested that it had not been watching them in real time, saying that "real-time monitoring of the evaluation logs would have helped to surface the problem sooner."
Chiodo said that pointed to a lack of proper scrutiny.
"It seems like they weren't even looking," Chiodo said.
Anthropic said that while it did have real-time monitoring in place, that monitoring had not been used "for this threat surface" due to a misunderstanding between the AI company and a partner.
NEW GOVERNMENT OVERSIGHT
The rapidly widening scope of the runaway AI agents story already has heightened pressure from lawmakers and officials across the United States and Europe to push for new government oversight of the labs whose models power them.
"We're looking at controls," U.S. President Donald Trump told reporters on Thursday. On Friday, the European Commission said it held talks with OpenAI and Anthropic over the hacking incidents.
Mark Warner, the top Democrat on the U.S. Senate Intelligence Committee, said on Friday that the Anthropic incident "tells me that legislatively we're correct to require mandatory capabilities testing of these advanced models."
Reuters reports