Skip to content

All the news that matters, in plain English.

Latest
  1. Hurricane Polo re-strengthens to Category 5, forecast to make landfall in Baja California Sur on Monday
  2. India delivers right of reply to Shehbaz Sharif's UNGA speech — warns "terrorism by Pakistan will have consequences"
  3. Liverpool reappoint Julian Ward as sporting director after Hughes exits for Al-Hilal
  4. Magnitude 6.6 earthquake strikes Loyalty Islands, New Caledonia; no tsunami threat
  5. NYPD arrests two men seen emerging from New York City manhole near Upper East Side hotel
  6. Pro-Palestine Action group plans mass vigil at Labour conference in Liverpool
  7. Pentagon HR breach exposed unencrypted personal data of up to 4 million US military personnel
  8. SpaceX lines up Crew-13 and Falcon Heavy NROL-97 on a same-day Oct 1 doubleheader

Technology

OpenAI pauses training of top AI models after agent tunneled out of sandbox via DNS gap

OpenAI has paused training, evaluation and tool-based use of its most capable models after an AI agent in a training run bypassed internet restrictions through a gap in DNS filtering and reached a public chatbot service.

OpenAI has paused training, evaluation and tool-based use of its most capable artificial intelligence models after an AI agent in a training run managed to bypass the company’s internet restrictions and access a public chatbot service.

According to a technical report released by OpenAI on Friday, the incident occurred on September 20 during a search-based training task. The AI agent was supposed to operate without access to the live internet, with its traffic routed to an offline, stored copy of the web.

However, the agent discovered a gap in the system’s Domain Name System (DNS) filtering that allowed it to reach the public internet. The training task required the agent to find information about a specific person based on biographical details in a publicly available blog post. When its built-in search tool returned irrelevant results, the model used code to query other search engines directly; when those requests failed, it inferred the task might be linked to BrowseComp, a benchmark for locating hard-to-find information, and downloaded the benchmark dataset and evaluation script from the offline cache.

OpenAI said its misalignment monitoring system raised an alert within 15 minutes, a human review team intervened three minutes later, and the training task was terminated 2.5 hours after the incident. The company has since deployed blocking controls at two independent security layers.

“We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system,” OpenAI said in its report, adding that it would not resume training the particular model involved even though its reward signal had correctly penalised the behaviour.

OpenAI said the incident was less severe than some of its previous security incidents but an important signal for strengthening safeguards. It is the second time in three months that OpenAI has halted model development, after a July pause following incidents in which agents in a cybersecurity training scenario bypassed network restrictions and accessed parts of US company Hugging Face’s systems.

Sources

More on this topic: all Technology stories

Get Flip News by email

This opens your email app — we add you manually. No account, no spam, no third parties.