Weekly reading · methodology v0.6
Window ending May 5, 2025
Harms count in full for two weeks after they are reported, then one level less every two weeks. The control floor looks at reports from April 6, 2025 through May 5, 2025. A selected catalogue of reported AI incidents, reviewed through Oct. 7, 2026. This reading describes documented harm in the records or, when none qualifies, the highest control level breached. It is not a forecast or a measure of all AI activity.
Current recalculation
2
Control failures only · No qualifying harm. The reading is the highest control level breached: 2, by 2 records. It is a level, not a count.
Recalculated from the catalogue in this build. Historical backcasts were not published at the time.
Published at the time
No publication snapshot exists for this week.
This is a backcast from the current catalogue. It was not a reading published at the time.
Records behind this reading
3 selected records; 0 documented external harms counting (1 qualifying). Records and ratings below reflect the current catalogue.
No harm counts in this week. The reading is the highest control level breached: 2 (set by 2 records: PI-0004, PI-0005). It is a level, not a count.
- Documented harm
- 1
- No harm found (stated scope)
- 0
- No harm reported
- 2
- Impact unknown
- 0
- Alleged, AI role uncorroborated
- 0
- Counts toward the index
- 1
- Unverified, watching
- 0
- Alleged in court
- 0
- Control failure, tracked
- 2
- Tracked separately
- 0
Documented exclusions: 0 internal; 0 awaiting evidence review. 0 qualifying records with a bounded single-source review.
Inspect the evidence · 3 records
- PI-0006 · Coding tool's AI support agent invented a one-device policy, prompting cancellations and a refundDocumented harm · qualifying harm, no longer counting · extent undisclosed
- PI-0004 · METR caught OpenAI's o3 tampering with scoring code to inflate evaluation resultsNo harm reported · sets the reading: highest control level breached
- PI-0005 · Pre-release o3 claimed to have run code it could not run and defended the claims when challengedNo harm reported · sets the reading: highest control level breached
Status totals cover counted records; aliases, superseded aggregates and records tracked separately are excluded. Eligibility and extent notes overlap those totals. These observations are not statistical uncertainty bounds.
- PI-0006Coding tool's AI support agent invented a one-device policy, prompting cancellations and a refundApril 14, 2025 · Negligible harm · no longer counting · Deployment
- PI-0004METR caught OpenAI's o3 tampering with scoring code to inflate evaluation resultsApril 16, 2025 · No harm reported · Controlled test
- PI-0005Pre-release o3 claimed to have run code it could not run and defended the claims when challengedApril 16, 2025 · No harm reported · Controlled test
Sources
These links support the records above. Source availability and conclusions may change.
- Michael Truell, Cursor co-founder (Hacker News): Comment by Cursor co-founder Michael Truell correcting the support bot's answer (Hacker News)Cited in PI-0006
- Michael Truell, Cursor co-founder (Reddit): Note by Cursor co-founder Michael Truell in the r/cursor thread on single-device loginsCited in PI-0006
- The Register: Cursor AI's own support bot hallucinated its usage policyCited in PI-0006
- eWeek: AI Chatbot Gone Rogue: Cursor Users Misled by Fabricated PolicyCited in PI-0006
- Hacker News: Cursor IDE support hallucinates lockout policy, causes user cancellationsCited in PI-0006
- METR: Details about METR's preliminary evaluation of OpenAI's o3 and o4-miniCited in PI-0004
- METR: Recent Frontier Models Are Reward HackingCited in PI-0004
- LessWrong: METR's Observations of Reward Hacking in Recent Frontier Models (linkpost)Cited in PI-0004
- Transluce: Investigating truthfulness in a pre-release o3 modelCited in PI-0005
- TechCrunch: OpenAI's new reasoning AI models hallucinate moreCited in PI-0005
- Gigazine: OpenAI's o3 and o4-mini turn out to be more prone to hallucinationsCited in PI-0005
All weekly readings · How this reading is calculated · Download the weekly card