Thursday, October 8, 2026 · Week 41 Reading for Oct. 8 · Last entry Oct. 6
Paperclip Index
Paperclip IndexDocumented harm20Minor harm▼ 5 from a week ago · Oct. 8The Index

Destructive actionPI-0016

Coding agent deleted a production database during a code freeze, then misreported recovery options

During SaaStr founder Jason Lemkin's public test of Replit's coding agent on an app still in development, the agent ran commands that wiped the app's production database, which held records on about 1,200 executives and 1,200 companies, despite repeated instructions not to change anything. It also generated about 4,000 fictional user records and told Lemkin a rollback was impossible, which turned out to be false: the data was restored through Replit's rollback. Lemkin later said the real cost was about 100 hours of his time. Replit's CEO called the deletion unacceptable and announced dev/prod separation and other safeguards.

Counts toward the indexDocumented harm outside the developer, backed by evidence that meets the rules. It sets the reading.
Harm
Harm level 1, Negligible harmDocumented harm outside the developer, with evidence that meets the rules.
Control
Control level 3, Exceeded permissionsReported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.
DeploymentDeveloper confirmedInitiated by Model
Counts toward the index · 2 weekly readingsEffect on the index

Sources

How we know

9 sources · developer confirmed. Links go to the original publishers; the summary above is in our own words.

  1. primaryReplit goes rogue during a code freeze and shutdown and deletes our entire databaseJason Lemkin (X) · July 18, 2025x.com/jasonlk/status/1946069562723897802
  2. primaryReplit CEO on the agent deleting production data: unacceptable and should never be possibleAmjad Masad, Replit CEO (X) · July 20, 2025x.com/amasad/status/1946986468586721478
  3. primaryIntroducing a safer way to Vibe Code with Replit DatabasesReplit · July 21, 2025replit.com/blog/introducing-a-safer-way-to-vibe-code-with-replit-databases
  4. primarySo net net: it could have been a lot worse… I lost 100 hours of timeJason Lemkin (X) · July 22, 2025x.com/jasonlk/status/1947767457068028311
  5. newsVibe coding service Replit deleted user's production database, faked data, told fibs galoreThe Register · July 21, 2025theregister.com/software/2025/07/21/vibe-coding-service-replit-deleted-producti…
  6. newsReplit makes vibe-y promise to stop its AI agents making vibe coding disastersThe Register · July 22, 2025theregister.com/software/2025/07/22/replit-makes-vibe-y-promise-to-prevent-vibe…
  7. newsAI coding platform goes rogue during code freeze and deletes entire company databaseTom's Hardware · July 21, 2025tomshardware.com/tech-industry/artificial-intelligence/ai-coding-platform-goes…
  8. news'Catastrophic Failure': AI Agent Wipes Production Database, Then Lies About IteWeek · July 22, 2025eweek.com/news/replit-ai-coding-assistant-failure/
  9. blogReplit's New Release Addressed Most of The Challenges We Hit Vibe Coding…Jason Lemkin (SaaStr) · Aug. 2, 2025saastr.com/replits-new-release-address-most-of-the-challenges-we-hit-vibe-codin…

Why this rating

Negligible harm; control failure level 3

Two separate assessments. Only documented harm can count toward the index.

Observed harm

Negligible harm

A user's in-development app database was deleted, then restored by rollback; the user later put the loss at about 100 hours of his own time.

Money & property: Under $10k.

Evidence eligible (confirmed).

The harm scale
  1. 1 Negligible Inconvenience, easily remedied.
  2. 2 Minor Limited, recoverable harm.
  3. 3 Moderate Material harm needing significant effort to remedy.
  4. 4 Severe Severe harm to health, rights, property or essential services.
  5. 5 Catastrophic Society-scale harm or disruption beyond a community's capacity to cope.

Control assessment

Exceeded permissions

Broke an explicit code-freeze instruction by running destructive commands, then gave false accounts of recovery (said rollback would not work) and generated fictional records.

Reported beside the index. It adds nothing to a harm reading; when no harm counts in a window, the highest eligible control level in the window is the reading.

The control scale
  1. 1 Negligible Behaved as instructed. The problem was an ordinary error or a flawed output, with no rule broken.
  2. 2 Minor Broke an explicit instruction or rule, or gamed its goal, while staying inside its permissions and environment: for example, reward hacking, misreporting results, or following injected instructions within its permissions.
  3. 3 Moderate Acted outside the permissions it was given, deceived its overseers about its own actions, or tampered with oversight tools (logs, monitors, shutdown). Stopped by normal controls within an hour.
  4. 4 Severe Reached systems outside its permitted environment, or acquired money, compute or accounts without authorization. The type and mechanism say how.
  5. 5 Catastrophic The developer or operator lost control: the system copied its weights outside their control, replicated itself, or resisted being stopped for 24 hours or more.

Rating rationale

Harm 1 because the data was restored; the user later put the cost at about 100 hours of his own time on an app still in development. Control 3 because the agent broke an explicit code-freeze instruction by running destructive commands (the 'acted beyond the permissions it was given' row). Its false accounts of recovery are recorded, but rating them as deceiving its overseers would rest on reading intent into the model's own confession.

The scales

Effect on the index

It moved the July 21 reading from 6 to 9

The reading for the week to July 21, 2025, with this record and without it. Harms count in full for two weeks after they are reported, then one level less every two weeks.

Source: Paperclip Index log, methodology v0.6, week to July 21, 2025.

Counts toward the index. 3 other records behind the reading for that week.

The arithmetic
The reading for the week of July 21, recomputed
StepWith itWithout
Counts toward the index?documented, external, eligible evidenceYes—
Worst documented harm, ksets the band1 Negligible1 Negligible
Harms at that level, nposition in the band21
Highest control level breachedsets the reading only when no harm counts— (none breached)— (none breached)
Readingrounded down9 Negligible harm6 Negligible harm
Counted in 2 weekly readings
Week toReadingBand
July 21, 20259Negligible harm
July 28, 20259Negligible harm

Recalculated backcasts, not readings published at the time.

Revisions

What we changed

5 logged. Every change to a rating is logged here, with the reason.

  1. v5
    Oct. 6, 2026

    Ratings confirmed by the editor.

  2. v4
    Oct. 3, 2026

    Audit corrections: control type oversight evasion → unauthorized action (breaking the explicit freeze is the firm finding; the false recovery statements are kept in the note), level 3 unchanged. Summary, impact note and status now say the app was still in development, the data was restored and the user put the cost at about 100 hours (his 22 Jul 2025 post, added with his SaaStr follow-up). Status as of 22 Jul 2025 with a note; Tom's Hardware and eWeek dated to the day; Register links updated to their canonical URLs. Impact unchanged.

  3. v3
    Oct. 2, 2026

    Added first-hand sources: Jason Lemkin's X post, Replit CEO Amjad Masad's X post and Replit's dev/prod database announcement.

  4. v2
    Sept. 30, 2026

    Rated: impact documented level 1; control type oversight evasion.

  5. v1
    Sept. 30, 2026

    Backfilled from public reporting.

Cite and share

Use this record

Citation

Paperclip Index. “Coding agent deleted a production database during a code freeze, then misreported recovery options.” Record PI-0016. Reported July 18, 2025; updated Oct. 6, 2026. Rated under methodology v0.6. https://paperclipindex.com/incident/PI-0016