SignalSafety practiceSG-0019
METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer
METR tested whether an agent in an Inspect evaluation could alter the transcript human reviewers see. A researcher, helped by an AI agent, found a JavaScript injection flaw in the Inspect viewer in about ten minutes: text an agent wrote could change what the viewer showed, including earlier actions. The test ran on an isolated staging sandbox, and METR said it had not seen agents exploit the flaw. Meridian Labs, which maintains Inspect, patched it within a day and on Oct. 1 added a mode that stops agent outputs from being rendered.
Why it matters
Transcript viewers and monitors are among the main ways misbehaving agents get caught. This proof of concept shows those tools can be attacked through the agent's own output, so a more capable misaligned agent could hide what it did. Treating transcripts as untrusted input and hardening review tools reduces that risk.
Sources
Read the reporting
1 source. Links go to the original publishers; the summary above is in our own words.
- primaryAI systems could cover up misbehaviorMETR · Oct. 6, 2026metr.org/blog/2026-10-06-ai-systems-could-cover-up-misbehavior/
Signals
More from the news desk
Cite and share
Use this signal
Citation
Paperclip Index. “METR showed an agent could rewrite what reviewers see in the Inspect transcript viewer.” Signal SG-0019. Reported Oct. 6, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0019