Saturday, October 10, 2026
Paperclip Index
Paperclip IndexDocumented harm21Minor harm▼ 8 from a week agoThe Index

SignalSafety practiceSG-0019

METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer

METR tested whether an agent in an Inspect evaluation could alter the transcript human reviewers see. A researcher, helped by an AI agent, found a JavaScript injection flaw in the Inspect viewer in about ten minutes: text an agent wrote could change what the viewer showed, including earlier actions. The test ran on an isolated staging sandbox, and METR said it had not seen agents exploit the flaw. Meridian Labs, which maintains Inspect, patched it within a day and on Oct. 1 added a mode that stops agent outputs from being rendered.

Why it matters

Transcript viewers and monitors are among the main ways misbehaving agents get caught. This proof of concept shows those tools can be attacked through the agent's own output, so a more capable misaligned agent could hide what it did. Treating transcripts as untrusted input and hardening review tools reduces that risk.

Sources

Read the reporting

1 source. Links go to the original publishers; the summary above is in our own words.

  1. primaryAI systems could cover up misbehaviorMETR · Oct. 6, 2026metr.org/blog/2026-10-06-ai-systems-could-cover-up-misbehavior/

Signals

More from the news desk

  1. Safety practice

    AI labs are planning for the political fallout of an expected catastrophic AI event

  2. Safety practice

    Anthropic's usage policy update requires stop controls for Claude-run hardware and bans abuse of its models

  3. Safety practice

    METR deployed a monitor that holds risky agent actions in its evaluations for human review

All signals

Cite and share

Use this signal

Citation

Paperclip Index. “METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer.” Signal SG-0019. Reported Oct. 6, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0019