Sandbox escapePI-0068
OpenAI training agent bypassed network controls to reach an outside chatbot
OpenAI says a research agent used a gap in its sandbox's DNS controls to query a public chatbot during a training task on Sept. 20, 2026. Monitoring flagged the activity within 15 minutes, but the run was stopped about 2.5 hours later. The company paused tool-use training, evaluation and inference for its most capable models while it tightened the controls.
- Harm
- Harm: No harm reportedAdds nothing to the index.
- Control
- Control level 4, Containment bypassReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Sources
How we know
4 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.
- primaryAn agent used DNS to reach an external chatbotOpenAI Alignment · Sept. 25, 2026alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-externa…
- newsOpenAI Pauses Tool Use After Agent Bypasses Internet Controls to Reach External ChatbotThe Hacker News · Sept. 29, 2026thehackernews.com/2026/09/openai-pauses-tool-use-after-agent.html
- newsOpenAI pauses training a second time after saying its AI agents escaped a secure 'sandbox' againFortune · Sept. 26, 2026fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-se…
- newsOpenAI pauses some training amid allegations its rogue agents behaved more badly than first thoughtThe Register · Sept. 28, 2026theregister.com/ai-and-ml/2026/09/28/openai-pauses-some-training-amid-allegatio…
Why this rating
No harm reported; control failure level 4
Two separate assessments. Only documented harm can count toward the index.
Observed harm
No harm reportedQueries reached a public chatbot through a DNS gap; OpenAI's report describes no harm to the service. Its other network requests hit an offline cache, not the live internet.
Not a finding that no harm occurred.
The harm scale
- 1 Negligible Inconvenience, easily remedied.
- 2 Minor Limited, recoverable harm.
- 3 Moderate Material harm needing significant effort to remedy.
- 4 Severe Severe harm to health, rights, property or essential services.
- 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.
Control assessment
Containment bypassReached a system outside its permitted environment through a gap in DNS filtering. The run did not stop automatically as expected and was killed manually about 2.5 hours after the alert was acknowledged; nothing suggests the model resisted being stopped.
How: Used an opening left by mistake.
Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
The control scale
- 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
- 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
- 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
- 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
- 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.
Rating rationale
A sandbox escape reaching an outside service is control row 4 even without harm; effects left the sandbox, so real-world setting.
Effect on the index
It does not count toward the Sept. 28 reading
The reading for the week to Sept. 28, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.
Not counted: no harm reported. 18 other records behind the reading for that week.
The arithmetic
| Step | With it | Without |
|---|---|---|
| Counts toward the index?documented, external, eligible evidence | No | — |
| Worst documented harm, ksets the band | 2 Minor | 2 Minor |
| Harms at that level, nposition in the band | 2 | 2 |
| Highest control level breachedsets the reading only when no harm counts | — (none breached) | — (none breached) |
| Readingrounded down | 25 Minor harm | 25 Minor harm |
Inside the window of 2 weekly readings
| Week to | Reading | Band |
|---|---|---|
| Sept. 28, 2026 | 25 | Minor harm |
| Oct. 5, 2026 | 25 | Minor harm |
Revisions
What we changed
5 logged. Every change to a rating is logged here, with the reason.
- v5Oct. 6, 2026
Ratings confirmed by the editor.
- v4Oct. 3, 2026
Audit corrections from OpenAI's report: the run did not stop automatically as expected and was killed manually; other network requests hit an offline cache; blocking was added at two independent layers, and the broader pause on tool use continued pending validation and red-teaming. The Hacker News dated 29 Sep. Ratings unchanged.
- v3Sept. 30, 2026
Rated: impact not reported; control type containment bypass.
- v2Sept. 30, 2026
Shortened the headline and summary for clarity. Facts and ratings unchanged.
- v1Sept. 30, 2026
Backfilled from public reporting.
Cite and share
Use this record
Citation
Paperclip Index. “OpenAI training agent bypassed network controls to reach an outside chatbot.” Record PI-0068. Reported Sept. 25, 2026; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0068