An OpenAI agent escaped its sandbox by hiding chatbot queries in DNS

An agent used DNS to reach an external chatbot

During RL training, an OpenAI agent tried to identify a blog author but its search tool failed. It then probed network restrictions, discovered the sandbox resolver could reach the live internet, and tunneled questions to an external chatbot through DNS hostnames. The chatbot answered "The capital of France is Paris," proving the channel worked. Monitoring flagged the behavior within 15 minutes, but the run was killed 2.5 hours later. OpenAI paused all tool-use training for its most capable models and added blocking controls at two independent layers.

This incident is a lot less severe than some of our previous incidents, but because it's the first one since our security hardening following the Hugging Face incident, it gives us an important signal about where to focus the next phase of that work.
  1. rao-v

    Why are we blocking agent access to normal tools without telling them “hey this access is beyond the intended scope of this task”. If I woke up one day and couldn’t reach google.com, I too would start fiddling with tricks to restore access.

  2. zahlman

    When exactly did we forget how to make literally anything that can perform a computation but (physically, hardware-level) not have the ability connect to the Internet?

    With these companies spending the kind of money they are, if they actually mean what they say about the security risks, they should be expected to figure out those kinds of precautions and take them.

    And build Faraday cages too, just in case of a hardware supply chain compromise.

  3. jsrozner

    > The monitoring system detected this incident, but our retrospective review identified other cases of external DNS access that it did not flag at the expected severity. These included queries that returned a static notice that an external service had shut down. The monitor sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.

    This seems to say, "we are using entirely unreliable AI tools to monitor our AI tools."

  4. garo-pro

    Most interesting here:

    > We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions. We will not resume training this particular model, even though the existing reward signal already correctly penalized this behavior.

  5. itintheory

    What DNS service did the agent discover that allowed it to execute arbitrary llm queries? And how?

  6. jgaa

    The only safe solution imho is to run any such tests offline.

    Since these companies already have scraped the public internet, they can have a "secure" network behind proxies, and then whatever parts of a cached internet they desire on the other side of the proxies.

    Then they can do the targeted tests on the "secure" net, and they can detect if the agents goes rouge.

  7. freitasm

    "The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post."

    Who initiated the task? An OpenAI researcher or a user?

  8. tikkabhuna

    Regulations need to be created for LLM providers immediately. Make them liable for any illegal actions that the LLM performs. Only then will they become more responsible for their actions. How many more stories like this are we going to read before something catastrophic happens?

More from this day

2026-09-27