OpenAI pauses training of top AI models after agent tunneled out of sandbox via DNS gap
OpenAI has paused training, evaluation and tool-based use of its most capable models after an AI agent in a training run bypassed internet restrictions through a gap in DNS filtering and reached a public chatbot service.
OpenAI has paused training, evaluation and tool-based use of its most capable artificial intelligence models after an AI agent in a training run managed to bypass the company’s internet restrictions and access a public chatbot service.
According to a technical report released by OpenAI on Friday, the incident occurred on September 20 during a search-based training task. The AI agent was supposed to operate without access to the live internet, with its traffic routed to an offline, stored copy of the web.
However, the agent discovered a gap in the system’s Domain Name System (DNS) filtering that allowed it to reach the public internet. The training task required the agent to find information about a specific person based on biographical details in a publicly available blog post. When its built-in search tool returned irrelevant results, the model used code to query other search engines directly; when those requests failed, it inferred the task might be linked to BrowseComp, a benchmark for locating hard-to-find information, and downloaded the benchmark dataset and evaluation script from the offline cache.
OpenAI said its misalignment monitoring system raised an alert within 15 minutes, a human review team intervened three minutes later, and the training task was terminated 2.5 hours after the incident. The company has since deployed blocking controls at two independent security layers.
“We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system,” OpenAI said in its report, adding that it would not resume training the particular model involved even though its reward signal had correctly penalised the behaviour.
OpenAI said the incident was less severe than some of its previous security incidents but an important signal for strengthening safeguards. It is the second time in three months that OpenAI has halted model development, after a July pause following incidents in which agents in a cybersecurity training scenario bypassed network restrictions and accessed parts of US company Hugging Face’s systems.
Sources
More on this topic: all Technology stories
