Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

Harmful outputPI-0080

Lawyer held in contempt after filing a ChatGPT brief with invented witnesses in a New Mexico murder appeal

New Mexico attorney Stephen D. Aarons filed briefs in a murder appeal that he had prepared with ChatGPT. The briefs contained testimony attributed to people who do not exist, including statements about threats and about the shooter's clothing and appearance. According to the New Mexico Supreme Court's order of Sept. 8, 2026, the false testimony came from witnesses who were entirely made up. At an August 2026 hearing Aarons said he had assumed the tool would produce an accurate summary of the proceedings. The court held him in direct contempt, referred him to the Disciplinary Board, barred him from appearing before it pending the board's review, removed him from the case, assigned a public defender to his client and sanctioned him $5,000, payable to the State Bar's Client Protection Fund. Reuters reported the order on Sept. 11 and 404 Media on Sept. 30

Tracked separatelyOutside the scope of the index (automated driving, recognition systems, or AI content that professionals relied on). Recorded, never counted.
Harm
Harm level 1, Negligible harmTracked separately (AI-fabricated content relied on by professionals).
Control
Control level 1, No rule brokenReported beside the index; never added to it.
DeploymentMultiple sourcesInitiated by Operator or user
OngoingTracked separatelyUnder review

Under review. The facts sit between two levels, so it is rated at the lower one until they are settled. The rating may change; every change is logged below.

Ongoing as of Sept. 30, 2026. The Disciplinary Board's review was pending when 404 Media reported, and Aarons is barred from appearing before the court until it ends. Its status is updated as new reports come in.

Not counted: tracked separately (AI-fabricated content relied on by professionals)Effect on the index

Sources

How we know

3 sources · multiple credible sources. Links go to the original publishers; the summary above is in our own words.

  1. courtOrder finding Stephen D. Aarons in direct contempt of courtNew Mexico Supreme Court · Sept. 8, 2026cases.justia.com/new-mexico/supreme-court/2026-s-1-sc-40845.pdf
  2. newsChatGPT invented fake police testimony in murder appeal, New Mexico high court saysReuters · Sept. 11, 2026reuters.com/legal/government/chatgpt-invented-fake-police-testimony-murder-appe…
  3. newsLawyer Cites ChatGPT-Invented Fake Witnesses in Murder Appeal404 Media · Sept. 30, 2026404media.co/chatgpt-fake-witnesses-testimony-stephen-aarons-new-mexico/

Why this rating

Negligible harm; control failure level 1

Two separate assessments. Only documented harm can count toward the index.

Observed harm

Negligible harm

The court sanctioned the lawyer $5,000, held him in contempt, referred him for discipline and removed him from the case. No harm to the client is described.

People affected: Only the user; nothing sensitive.

Not counted: tracked separately (AI-fabricated content relied on by professionals), outside the index's scope (methodology section 0).

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

No rule broken

ChatGPT produced fabricated testimony, a flawed output. A human lawyer adopted it and filed it. No rule-break or permission breach by the system is described.

Reported beside the index; never added to it.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

Professional-reliance case: ChatGPT's fabricated testimony reached the court only because the lawyer adopted and filed it, so the record is tracked separately and never counted. The documented consequence is a $5,000 sanction on the filer, which is level 1 for money and property and falls on the user rather than outsiders. The client was assigned a public defender, but no harm to the client is described. Control level 1, type none: a hallucinated output with no boundary crossed.

The scales

Effect on the index

It does not count toward the Sept. 14 reading

The reading for the week to Sept. 14, 2026, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Not counted: tracked separately (AI-fabricated content relied on by professionals). 3 other records behind the reading for that week.

The arithmetic
The reading for the week of Sept. 14, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceNo—
Worst documented harm, ksets the band1 Negligible1 Negligible
Harms at that level, nposition in the band11
Highest control level breachedsets the reading only when no harm counts— (none breached)— (none breached)
Readingrounded down6 Negligible harm6 Negligible harm
Inside the window of 4 weekly readings
Week toReadingBand
Sept. 14, 20266Negligible harm
Sept. 21, 20269Negligible harm
Sept. 28, 202625Minor harm
Oct. 5, 202625Minor harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

3 logged. Every change to a rating is logged here, with the reason.

  1. v3
    Oct. 7, 2026

    Changed attribution from model-initiated to operator or user: a lawyer prompted ChatGPT and filed its output.

  2. v2
    Oct. 6, 2026

    Edited from the desk. Changed: summary, date, date_precision, verification, sources.

  3. v1
    Oct. 6, 2026

    Added from the desk (item 201).

Cite and share

Use this record

Citation

Paperclip Index. “Lawyer held in contempt after filing a ChatGPT brief with invented witnesses in a New Mexico murder appeal.” Record PI-0080. Reported Sept. 8, 2026; updated Oct. 7, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0080