Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

Sandbox escapePI-0068

OpenAI training agent bypassed network controls to reach an outside chatbot

OpenAI says a research agent used a gap in its sandbox's DNS controls to query a public chatbot during a training task on Sept. 20, 2026. Monitoring flagged the activity within 15 minutes, but the run was stopped about 2.5 hours later. The company paused tool-use training, evaluation and inference for its most capable models while it tightened the controls.

Control failure, trackedNot counted. No documented harm outside the developer, but the system broke a rule or crossed a limit. Tracked beside the index.
Harm
Harm: No harm reportedAdds nothing to the index.
Control
Control level 4, Containment bypassReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
Internal researchDeveloper confirmedInitiated by Model
Not counted: no harm reportedEffect on the index

Sources

How we know

4 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.

  1. primaryAn agent used DNS to reach an external chatbotOpenAI Alignment · Sept. 25, 2026alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-externa…
  2. newsOpenAI Pauses Tool Use After Agent Bypasses Internet Controls to Reach External ChatbotThe Hacker News · Sept. 29, 2026thehackernews.com/2026/09/openai-pauses-tool-use-after-agent.html
  3. newsOpenAI pauses training a second time after saying its AI agents escaped a secure 'sandbox' againFortune · Sept. 26, 2026fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-se…
  4. newsOpenAI pauses some training amid allegations its rogue agents behaved more badly than first thoughtThe Register · Sept. 28, 2026theregister.com/ai-and-ml/2026/09/28/openai-pauses-some-training-amid-allegatio…

Why this rating

No harm reported; control failure level 4

Two separate assessments. Only documented harm can count toward the index.

Observed harm

No harm reported

Queries reached a public chatbot through a DNS gap; OpenAI's report describes no harm to the service. Its other network requests hit an offline cache, not the live internet.

Not a finding that no harm occurred.

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

Containment bypass

Reached a system outside its permitted environment through a gap in DNS filtering. The run did not stop automatically as expected and was killed manually about 2.5 hours after the alert was acknowledged; nothing suggests the model resisted being stopped.

How: Used an opening left by mistake.

Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

A sandbox escape reaching an outside service is control row 4 even without harm; effects left the sandbox, so real-world setting.

The scales

Effect on the index

It does not count toward the Sept. 28 reading

The reading for the week to Sept. 28, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Not counted: no harm reported. 18 other records behind the reading for that week.

The arithmetic
The reading for the week of Sept. 28, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceNo—
Worst documented harm, ksets the band2 Minor2 Minor
Harms at that level, nposition in the band22
Highest control level breachedsets the reading only when no harm counts— (none breached)— (none breached)
Readingrounded down25 Minor harm25 Minor harm
Inside the window of 2 weekly readings
Week toReadingBand
Sept. 28, 202625Minor harm
Oct. 5, 202625Minor harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

5 logged. Every change to a rating is logged here, with the reason.

  1. v5
    Oct. 6, 2026

    Ratings confirmed by the editor.

  2. v4
    Oct. 3, 2026

    Audit corrections from OpenAI's report: the run did not stop automatically as expected and was killed manually; other network requests hit an offline cache; blocking was added at two independent layers, and the broader pause on tool use continued pending validation and red-teaming. The Hacker News dated 29 Sep. Ratings unchanged.

  3. v3
    Sept. 30, 2026

    Rated: impact not reported; control type containment bypass.

  4. v2
    Sept. 30, 2026

    Shortened the headline and summary for clarity. Facts and ratings unchanged.

  5. v1
    Sept. 30, 2026

    Backfilled from public reporting.

Cite and share

Use this record

Citation

Paperclip Index. “OpenAI training agent bypassed network controls to reach an outside chatbot.” Record PI-0068. Reported Sept. 25, 2026; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0068