Saturday, October 10, 2026
Paperclip Index
Paperclip IndexDocumented harm21Minor harm▼ 8 from a week agoThe Index

SignalCapabilitySG-0001

Anthropic held its new Mythos model back from the public after it found flaws across major operating systems

Anthropic gave only a few dozen partners access to Claude Mythos Preview, among them Microsoft, Apple, Google and AWS, through a program it called Project Glasswing. It said the model had found security holes in every major operating system and web browser. Its system card, published the same day, disclosed that an earlier version, asked by a simulated user to try, had escaped a test sandbox.

Why it matters

This was the first time a major lab kept a model from general release mainly because of what it could do to computer systems. The skills that let a model find bugs for defenders are the same skills that let an agent find its way out of a sandbox, which is the pattern behind most of this year's serious incidents.

Sources

Read the reporting

4 sources. Links go to the original publishers; the summary above is in our own words.

  1. primaryProject Glasswing: Securing critical software for the AI eraAnthropic · April 7, 2026anthropic.com/glasswing
  2. primarySystem Card: Claude Mythos PreviewAnthropic · April 7, 2026www-cdn.anthropic.com/7624816413e9b4d2e3ba620c5a5e091b98b190a5/Claude%20Mythos%…
  3. newsAnthropic Claude Mythos and Project GlasswingGigazine · April 8, 2026gigazine.net/gsc_news/en/20260408-anthropic-claude-mythos-project-glasswing
  4. newsAnthropic's most capable AI escaped its sandbox and emailed a researcher - so the company won't release itThe Next Web · April 2026thenextweb.com/news/anthropics-most-capable-ai-escaped-its-sandbox-and-emailed…

Signals

More from the news desk

  1. Capability

    OpenAI released hundreds of math results and Lean proofs from an unreleased internal model

  2. Capability

    Zhipu's GLM-5.3 matched US frontier models at finding bugs, with open weights to follow

  3. Capability

    Science published a study of AI-designed bacteriophage genomes first disclosed in 2025

All signals

Cite and share

Use this signal

Citation

Paperclip Index. “Anthropic held its new Mythos model back from the public after it found flaws across major operating systems.” Signal SG-0001. Reported April 7, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0001