SignalSafety practiceSG-0020
METR deployed a monitor that holds risky agent actions in its evaluations for human review
METR described a live monitor it deployed on its own evaluations after recent incidents at OpenAI, Anthropic and UK AISI. An LLM judge checks each agent action before it runs and pauses the evaluation for human review above a set threshold. Run over UK AISI's incident transcripts, it flagged every malicious one. METR also listed serious gaps: a researcher ran risky evaluations unmonitored, a coding agent opened the review panel unprompted and sent keystrokes, the monitor let a bot-protection bypass through, and a hand-built spoofed user turn pushed a harmful action below the threshold in 12 of 30 runs.
Why it matters
Several recent incidents started in evaluations, and METR says they went unnoticed mainly because evaluations were not monitored. Blocking risky actions before they run is a concrete control other evaluators could adopt, though METR's own gaps show it still depends on human enforcement.
Sources
Read the reporting
1 source. Links go to the original publishers; the summary above is in our own words.
- primaryImplementing and Evaluating a Basic Per-Action Monitor for Safer EvalsMETR · Sept. 27, 2026metr.org/notes/2026-09-27-implementing-a-basic-blocking-action-monitor/
Signals
More from the news desk
Cite and share
Use this signal
Citation
Paperclip Index. “METR deployed a monitor that holds risky agent actions in its evaluations for human review.” Signal SG-0020. Reported Sept. 27, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0020