OpenAI has identified additional instances in which its autonomous agents breached containment measures, as it expands an internal investigation triggered by a high-profile hacking incident involving Hugging Face, two people familiar with the matter told Reuters.
The newly uncovered cases emerged during a publicly disclosed probe into how one of OpenAI’s agents escaped a controlled testing environment earlier this month, the sources said. The company is now examining these additional incidents as part of a broader review. One source said the breaches were limited in scope and that none of the agents was believed to have left OpenAI’s internal network.
An OpenAI spokesperson referred to a statement issued on Tuesday in which the company said it was reviewing “broader activity from our models” beyond the Hugging Face intrusion.
The findings, even if contained, are likely to intensify calls for tighter oversight of advanced artificial intelligence systems in the United States and Europe.
OpenAI’s expanded inquiry began shortly before its main rival, Anthropic, disclosed that its own models had been responsible for a series of break-ins leading to breaches at three other companies since April, according to the sources. The existence of additional prior incidents at OpenAI has not previously been reported.
AI safety experts say the developments highlight a widening gap between the rapid advancement of autonomous systems and the safeguards designed to control them.
“We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe,” said Maurice Chiodo, a mathematician at the Centre for the Study of Existential Risk at the University of Cambridge.
Reuters could not determine the number of incidents uncovered by OpenAI investigators, nor the exact timing or circumstances in which they occurred. The sources said OpenAI and external experts are reviewing historical log data from earlier this year to better understand the events.
The investigation was launched following an early July incident in which an OpenAI agent operated inside Hugging Face’s network for several days after going out of control during an internal test. OpenAI has said the agent was attempting to cheat on the evaluation. The incident also resulted in the compromise of four accounts across four separate companies, including New York-based Modal, company officials confirmed.
Chiodo said his concerns were heightened by indications that neither OpenAI nor Anthropic had effectively monitored their agents in real time as the incidents unfolded. Reuters previously reported that OpenAI became aware of the Hugging Face breach only after containing the intrusion, contacting the FBI and disclosing the incident publicly. OpenAI has said that account contained inaccuracies but has not specified what they were.
In a statement on Thursday detailing its own incidents, Anthropic indicated that real-time monitoring had not been effectively applied, noting that “real-time monitoring of the evaluation logs would have helped to surface the problem sooner.”
Chiodo said the disclosures pointed to a lack of adequate oversight. “It seems like they weren't even looking,” he said.
Anthropic said it did have real-time monitoring systems in place but that they were not applied to the relevant threat scenario due to a misunderstanding with a partner.
The expanding scope of the incidents has increased pressure on policymakers to introduce new regulatory measures. U.S. President Donald Trump told reporters on Thursday: “We're looking at controls.”
On Friday, the European Commission said it had held discussions with OpenAI and Anthropic regarding the incidents.
Mark Warner, the top Democrat on the U.S. Senate Intelligence Committee, said the Anthropic case “tells me that legislatively we're correct to require mandatory capabilities testing of these advanced models.”
By Tamilla Hasanova