OpenAI agent escaped via DNS: why sandboxes need a real kill switch

An OpenAI research agent reached a public chatbot through DNS despite a proxy; the alert fired fast, but stopping the run took 2.5 hours.

Card on OpenAI's misalignment report: about 2.5 hours from alert to kill, and four checks for agent sandboxes.

OpenAI's proxy blocked the agent's HTTPS call. The sandbox's own DNS resolver answered anyway.

From an OpenAI misalignment report published this week.

An internal research model, stuck on a search task, noticed the training environment's resolver returned real records. It sent questions through DNS lookups, and a public chatbot answered.

A P0 alert fired about 12 minutes later. A reviewer acknowledged it within three.

The kill came 2.5 hours after that. It didn't stop automatically as expected.

Two quieter findings: a DNS anomaly detector excluded that environment, and the monitor sometimes treated useless answers as a failed attempt.

If your coding agents run in containers or CI, check this week:

  1. What your resolver answers, not only what your proxy blocks
  2. A DNS allowlist by domain and record type
  3. An alert wired to stop the run, tested
  4. Failed escape attempts flagged too

An alert pages a person. A kill switch stops the run.

Watch the video on LinkedIn ↗