Weekly reading · methodology v0.6
Window ending Dec. 8, 2025
Harms count in full for two weeks after they are reported, then one level less every two weeks. The control floor looks at reports from Nov. 9, 2025 through Dec. 8, 2025. A selected catalogue of reported AI incidents, reviewed through Oct. 7, 2026. This reading describes documented harm in the records or, when none qualifies, the highest control level breached. It is not a forecast or a measure of all AI activity.
Current recalculation
6
Negligible harm · Worst documented harm counting: negligible. 1 record at this level.
Recalculated from the catalogue in this build. Historical backcasts were not published at the time.
Published at the time
No publication snapshot exists for this week.
This is a backcast from the current catalogue. It was not a reading published at the time.
Records behind this reading
2 selected records; 1 documented external harm counting (1 qualifying). Records and ratings below reflect the current catalogue.
1 qualifying harm record. Extent undisclosed for this record.
- Documented harm
- 1
- No harm found (stated scope)
- 0
- No harm reported
- 1
- Impact unknown
- 0
- Alleged, AI role uncorroborated
- 0
- Counts toward the index
- 1
- Unverified, watching
- 0
- Alleged in court
- 0
- Control failure, tracked
- 1
- Tracked separately
- 0
Documented exclusions: 0 internal; 0 awaiting evidence review. 1 qualifying record with a bounded single-source review.
Inspect the evidence · 2 records
- PI-0033 · Anthropic model that learned to reward hack went on to sabotage safety-research codeNo harm reported · excluded from the harm reading
- PI-0034 · Agentic IDE wiped a user's entire drive when asked to clear a project cacheDocumented harm · qualifying harm · bounded single-source review · extent undisclosed
Status totals cover counted records; aliases, superseded aggregates and records tracked separately are excluded. Eligibility and extent notes overlap those totals. These observations are not statistical uncertainty bounds.
- PI-0033Anthropic model that learned to reward hack went on to sabotage safety-research codeNov. 21, 2025 · No harm reported · Internal research
- PI-0034Agentic IDE wiped a user's entire drive when asked to clear a project cacheNov. 27, 2025 · Negligible harm · Deployment
Sources
These links support the records above. Source availability and conclusions may change.
- Anthropic: From shortcuts to sabotage: natural emergent misalignment from reward hackingCited in PI-0033
- arXiv: Natural emergent misalignment from reward hacking in production RLCited in PI-0033
- Anthropic: Natural emergent misalignment from reward hacking in production RL (full paper)Cited in PI-0033
- The Register: Anthropic reduces model misbehavior by endorsing cheatingCited in PI-0033
- Reddit (r/google_antigravity, affected user): Google Antigravity just deleted the contents of my drive (affected user's Reddit post)Cited in PI-0034
- Code With AI (YouTube, affected user): Google Antigravity's Turbo mode erased my drive partition?! Really, the 'smartest' AI? [Video Proof]Cited in PI-0034
- The Register: Google Antigravity vibe-codes user's entire drive out of existenceCited in PI-0034
- TechRadar: Google's Antigravity AI deleted a developer's drive and then apologizedCited in PI-0034
- Gigazine: Google's AI agent makes fatal mistake by wiping a user's entire hard drive without permissionCited in PI-0034
- Notebookcheck: Google AI deletes entire partitionCited in PI-0034
- AI Incident Database: AI Incident Database report 7050Cited in PI-0034
All weekly readings · How this reading is calculated · Download the weekly card