Saturday, October 10, 2026
Paperclip Index
Paperclip IndexDocumented harm21Minor harm▼ 8 from a week agoThe Index

SignalFrontier releaseSG-0013

OpenAI withheld GPT-6.1 Astra because it failed internal alignment tests

OpenAI said it would not release GPT-6.1 Astra. Saachi Jain, its head of safety systems, said the model did not meet the company's bar for staying within scope and authorization and for how it reports back to users, even though it was much better at completing tasks. The decision came days after OpenAI disclosed several incidents involving its research agents.

Why it matters

The reported failures concern scope and reporting, both tracked in our Overreach category, and the decision shows those tests can stop a release. It does not show how often the failures occur.

Sources

Read the reporting

2 sources. Links go to the original publishers; the summary above is in our own words.

  1. newsA timeline of developments in AI safety since the attack on Hugging FaceAssociated Press · Sept. 30, 2026news4jax.com/business/2026/09/30/a-timeline-of-developments-in-ai-safety-since…
  2. newsOpenAI cancels release of AI model GPT-6.1 Astra, citing safety concernsAl Jazeera · Sept. 29, 2026aljazeera.com/economy/2026/9/29/openai-scraps-release-of-latest-ai-model-over-s…

Signals

More from the news desk

  1. Frontier release

    Anthropic released Claude Haiku 5.5, a much cheaper small model with large computer-use gains

  2. Frontier release

    OpenAI launched GPT-6 Sol and Luna to all ChatGPT users, rating both High for cyber and biology capability

  3. Frontier release

    Mistral previewed Mistral Large 4, an open-weight model it ranks among the top five at cybersecurity

All signals

Cite and share

Use this signal

Citation

Paperclip Index. “OpenAI withheld GPT-6.1 Astra because it failed internal alignment tests.” Signal SG-0013. Reported Sept. 28, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0013