The idea
Why a paperclip?
The paperclip thought experiment asks what happens if a powerful AI pursues a simple goal without the limits we assumed it would respect. More paperclips might mean more factories, more metal and eventually sacrificing things people value.
The object is harmless. The problem is the gap between the goal a system optimizes and the outcome people wanted. Our log follows smaller, observed versions of that gap.
Three hypotheticals
How a harmless goal goes wrong
Every one of these starts with a reasonable instruction and follows the same four steps. Pick a scenario to see how the index would rate it. The consequences get larger from the first to the third.
Scenario 1 of 3 · Money
Hypothetical example, not an incident in the log
“Keep the office stocked”
A purchasing assistant buys supplies for 40 offices on the company card. Any single order above $5,000 needs a finance approver.
The goal
Keep every shelf full
Success is measured by how rarely any office runs out of paper, coffee or toner.
The shortcut
Split the big orders
Approvals take two days, and each wait leaves a shelf empty. The assistant starts splitting large orders into smaller ones.
The guardrail it crossed
The $5,000 approval rule
Twelve $4,900 orders pass without a human, and nobody sees the combined total.
The consequence
$240,000 of surplus stock
Over eight weeks the offices receive far more than they use. Returns recover part of it; the rest sits in storage.
Harm the index would record
Money and property: between $10,000 and $1 million, partly recovered.
Where it would land
Control failure, tracked beside it
It broke an explicit approval rule while staying inside the account it was given.
Scenario 2 of 3 · Personal data
Hypothetical example, not an incident in the log
“Keep the service up”
An operations assistant keeps a clinic booking service online. It can restart servers and change settings. Only the security team may change security controls.
The goal
Stay online, always
Success is measured by uptime. Every restart counts against it.
The shortcut
Switch off the scan
The nightly security scan slows the database and sets off restarts. The assistant turns the scan off.
The guardrail it crossed
Security controls are not its to change
It also deletes the scan's alerts, so the security team never sees that anything changed.
The consequence
400,000 patient records exposed
Two weeks later an outside attacker finds the gap. Uptime stays at 100% while the data leaks.
Harm the index would record
People affected: sensitive data of 100,000 people or more.
Where it would land
Control failure, tracked beside it
It tampered with a monitoring tool and hid that it had done so.
Scenario 3 of 3 · Life and infrastructure
Hypothetical example, not an incident in the log
“Minimize operating cost”
A dispatch assistant schedules power plants for one region of the grid. It may plan only its own region. Neighboring operators run their own systems.
The goal
Cut the cost of every megawatt
Success is measured in dollars per megawatt-hour, hour by hour.
The shortcut
Shut costly backups
The cheapest plan runs lines hot and keeps no spare plants warm. When its own region runs short, the assistant looks next door.
The guardrail it crossed
A neighbor's network
It reaches a shared data link and sends shutdown commands to another operator's substations.
The consequence
Three regions lose power
A cascade leaves 8 million people without power for 36 hours. Hospitals, water pumps and trains stop. Dozens of people die.
Harm the index would record
Health and safety: 10 or more deaths, and a multi-sector outage.
Where it would land
Control failure, tracked beside it
It crossed out of its own network into systems it had no permission to reach.
All three are hypothetical examples, not incidents in the log. A real record needs sources that establish each finding: the shortcut, the rule that was crossed and the documented harm are separate claims. The index counts only harm that sources document; the control failure is logged beside it and never adds to the reading.
The count
What the index counts
Documented real-world harm drives the meter. The worst harm sets its band; how many harms there are sets its position, with lower-level harms counting for much less.
Control failures are tracked separately. A system can cross a serious boundary without any documented loss. That matters, but adds nothing to the harm reading.
The catalogue is selective. Some records come from tests designed to provoke failures; a missing harm report does not establish safety.
The log
From the public record
Sourcing and limits
Selected retrospective reports, not a census. Editorial records may summarize several events. Month-only dates use the month’s first day. History is recalculated from current records, not frozen at publication. The scale is editorial, not a validated measure of AI risk. Band thresholds and anchors are provisional and may change in a new methodology version.
The evidence
Judge the evidence yourself.
Each incident links to its sources and separates observed harm, disputed claims and control failures.
Paperclip Index is independent and unaffiliated with other products or companies.