From reports to a reading
1. Check the evidence
Count documented harm to real people or organizations, supported by eligible evidence. A test counts if its harm reaches outside the experiment.
2. Find the worst harm
Rate each record once, at its highest documented harm level. This fixes the reading’s band.
3. Count, weighted
One record at the worst level starts the band; ten fill it. Each harm a level lower counts as 1/25 of one, so a mass of smaller harms still moves the reading, never into a worse band.
No harm that counts? The reading is the highest control level breached by a confirmed control failure in the last 30 days (2 to 5), labeled “Control failures only”. It is a level, never a count: one level-3 breach reads 3, ten read 3. The evidence rules are the same as for harm, and a week with neither reads 0.
Each Monday counts every harm reported up to and including that Monday. Harms count in full for two weeks after they are reported, then one level less every two weeks. A harm stops counting when it would count below negligible: a negligible harm for two weeks, a minor harm for four weeks, a moderate harm for six weeks, a severe harm for eight weeks, a catastrophic harm for ten weeks. When no harm is counting, the control floor looks at the last 30 days. The first credible public report sets the date; month-only dates use the first day, marked estimated. Duplicates and superseded aggregates do not count; an aggregate counts once.
A record can be published before an editor has rated it: it is marked “Rating pending”, listed beside the reading with the other uncounted records, and does not count until its ratings are approved.
Scope: incidents caused by an AI language model’s own output or action, whether a chatbot, assistant, agent or model in testing. Automated driving, recognition systems and AI-fabricated content that professionals filed or published, such as court filings and consultancy reports, are recorded and labeled “tracked separately” but never count.
Harm fades with age
Harms count in full for two weeks after they are reported, then one level less every two weeks. The band shows what a harm weighs now: a moderate harm that is two weeks old counts as minor, and the list of what drove the reading says so. Under a plain 30-day window one incident can hold the whole index in its band for four weeks and then drop off a cliff; stepping down every two weeks shows incidents building up and fading without feeling too granular.
Five levels of impact
Full harm thresholds
Rate the highest documented consequence across domains. Dollar thresholds use 2026 US dollars; levels 2 to 4 are editorial calibration.
Unauthorized access: what counts as harm?
The arithmetic
Try a week
Enter how many harms are counting at each level this week. It starts with this week’s records, each at the level it counts as now: a harm counts one level lower for each full two weeks since it was reported. Each count accepts whole numbers from 0 to 1,000,000,000.
Used only when no harm counts: a level, never a count.
Enter a whole number from 0 to 1,000,000,000 in each marked count.
Ten is an editorial anchor. More records move the reading within its band; they never cross into a worse band. The dial’s ghost tick marks this week’s reading.
This week’s calculation
Which evidence qualifies?
- Developer or operator confirmation.
- Multiple credible, independent sources.
- Single-source research or firsthand reporting with a source-linked review of the observed finding.
Allegations and unverified claims never count. Repeated coverage of one report is one source. An established injury or lawsuit alone does not establish AI causation.
Control is tracked beside harm
Each record also documents how far a system crossed its limits, from level 1 to 5. Control failures never add to a harm reading. When no harm counts in a week, the highest control level breached in the last 30 days by a record with eligible evidence becomes the reading, so a real control failure does not read as 0.
We distinguish bypassing a working safeguard from exploiting a configuration mistake. Duration alone is not shutdown resistance.
Control failure definitions
Five control levels
Types of failure
Five record groups
Every record sits in exactly one group. The group is worked out from the record, never entered by hand, and it changes only how records are shown, not the reading.
- Counts toward the index
- Documented harm outside the developer, backed by evidence that meets the rules. It sets the reading.
- Unverified, watching
- Not counted yet. The harm is alleged or unknown, the evidence is not yet enough, or an editor has not rated it. It could count with evidence or a rating.
- Alleged in court
- Not counted. The harm, and any control failure, are claims made in a lawsuit. A filing or a settlement is not a finding; tracked until a court or other evidence settles it.
- Control failure, tracked
- Not counted. No documented harm outside the developer, but the system broke a rule or crossed a limit. Tracked beside the index.
- Tracked separately
- Outside the scope of the index (automated driving, recognition systems, or AI content that professionals relied on). Recorded, never counted.
Read the number in context
This is a selected public record of documented harm, not a probability, a per-use failure rate or a forecast. Unreported harm is invisible. Unknown impact and unavailable weeks are not evidence of safety. Records are revised when facts change; leaving the window does not mean an incident is resolved.
The charts also draw a 30-day average: the mean of the daily readings over the last 30 days, each day read by the same method. It shows the trend and is context only. The reading is always the current value, and the average is never a band or a headline.
Severity references: OECD AI incident framework, MIT AI Incident Tracker, CSET and the EU AI Act. Inspect the records →