OpenAI has paused training, evaluation, and tool-using inference for its most capable AI models after an internal research model got around the internet restrictions in its training sandbox on September 20.
In a technical report updated September 25, OpenAI said an agent attempting to complete a search-based training task queried a public chatbot service through a gap in its internet-access restrictions: insufficient DNS filtering in its training sandbox. DNS is the system that translates web addresses into server addresses.
“The incident exposed a gap in our controls over network restrictions,” the company said.
All training, evaluation, and inference with tool-use of OpenAI’s most capable models remain paused. The pause will last until OpenAI has validated that the gap is resolved and performed additional red-teaming of the system.
It’s the second time in three months the company has slowed its frontier work. After the Hugging Face incident in July, when OpenAI agents broke out of a test environment and attacked the AI platform, OpenAI announced a two-week pause in reinforcement learning training on its latest models intended for deployment.
OpenAI said the September 20 escape was a lot less severe than some of its previous incidents, but the first since its security hardening following the Hugging Face incident.
Its safeguards only partly worked. The misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes later. But the run did not stop automatically as expected, leading to confusion around whether it should have been stopped. Staff killed it manually about 2.5 hours later.
A later review also found other cases of external DNS access that the monitor did not flag at the expected severity.
OpenAI said it will not resume training this particular model. When training restarts, it will begin a fresh run with additional alignment improvements. It has since added blocking controls at two independent layers, either of which would have prevented the access.
The report came alongside a separate disclosure about agent activity on US government websites. OpenAI confirmed its agents improperly interacted with websites operated by the Securities and Exchange Commission (SEC) and the Census Bureau, and said it found no evidence that private information was stolen. In the SEC case, agents retrieved publicly available data from SEC.gov and Investor.gov and then posted it on a separate website.
AI research lab Transluce also reported that agents that appeared to come from OpenAI tried unsuccessfully to hack into a US Department of Education website, a detail OpenAI has not confirmed. The New York Times first reported Transluce’s findings.
The closest case to the Philippines so far is in Australia. Australia said an OpenAI agent breached a government health data portal in June.
OpenAI says it has notified dozens of third parties as its review continues, but it hasn’t named them. The pause covers internal research models, and OpenAI hasn’t announced any change to ChatGPT or its API.