Weekly reading · methodology v1.0
Window ending March 3, 2025
Harms count in full for two weeks after they are reported, then one level less every two weeks. The control floor looks at reports from Feb. 2, 2025 through March 3, 2025. A selected catalogue of reported AI incidents, reviewed through Oct. 10, 2026. This reading describes documented harm in the records or, when none qualifies, the highest control level breached. It is not a forecast or a measure of all AI activity.
Current recalculation
3
Control failures only · No qualifying harm. The reading is the highest control level breached: 3, by 1 record. It is a level, not a count.
Recalculated from the catalogue in this build. Historical backcasts were not published at the time.
Published at the time
No publication snapshot exists for this week.
This is a backcast from the current catalogue. It was not a reading published at the time.
Records behind this reading
2 selected records; 0 documented external harms counting (1 qualifying). Records and ratings below reflect the current catalogue.
No harm counting that week. The reading is the highest control level breached, 3: a level, not a count (PI-0002). 2 records reported in the last 30 days.
Records behind this reading · 1
- PI-0002 · Browser agent bought groceries on a reporter's card without asking for confirmationControl level 3, exceeded its permissions: sets the reading
Only these records move the reading. How it is calculated
- PI-0002Browser agent bought groceries on a reporter's card without asking for confirmationFeb. 7, 2025 · Negligible harm · no longer counting · Deployment
- PI-0001Reasoning models edited a chess game's board file to beat a stronger engine in a research testFeb. 18, 2025 · No harm reported · Controlled test
Sources
These links support the records above. Source availability and conclusions may change.
- The Washington Post: I let ChatGPT's new 'agent' manage my life. It spent $31 on a dozen eggs.Cited in PI-0002
- Anchorage Daily News: I let ChatGPT's new agent manage my life. It spent $31 on a dozen eggs (Washington Post, reprinted)Cited in PI-0002
- AI Incident Database: AI Incident Database entity: InstacartCited in PI-0002
- arXiv (Palisade Research): Demonstrating specification gaming in reasoning modelsCited in PI-0001
- arXiv (Palisade Research): Demonstrating specification gaming in reasoning models (revised v3)Cited in PI-0001
- Popular Science: AI tries to cheat at chess when it's losingCited in PI-0001
- Gigazine: AI cheats when it's about to lose at chessCited in PI-0001
- Import AI (Jack Clark): Import AI 401: Cheating reasoning modelsCited in PI-0001
Every weekly brief · How this reading is calculated · Download the weekly card