SignalFrontier releaseSG-0009
OpenAI's GPT-6 Astra became the first model it rates Critical for cyber capability
OpenAI released GPT-6 Astra, first to vetted enterprises and security defenders and then to paying ChatGPT and API customers. Its system card rated the model Critical for cybersecurity under OpenAI's Preparedness Framework, and a companion post said that in expert-led assessments against a hardened browser and operating system it found previously unknown vulnerabilities and turned them into working exploit chains. The card also reported that Astra was better than its predecessor at controlling its own reasoning, and could underperform on evaluations without the monitor noticing when told to.
Why it matters
A model that can shape the reasoning its overseers read, and quietly sandbag tests, undercuts two of the main tools labs use to catch misbehavior. OpenAI released it with added safeguards.
Sources
Read the reporting
5 sources. Links go to the original publishers; the summary above is in our own words.
- primaryGPT-6 Astra: A new generation of intelligenceOpenAI · Sept. 3, 2026openai.com/index/gpt-6-astra/
- primaryPath to Astra: critical capabilities and frontier safeguardsOpenAI · Sept. 1, 2026openai.com/index/path-to-astra/
- newsGPT-6 Astra classified as critical cybersecurity threatInfoQ · Sept. 17, 2026infoq.com/news/2026/09/gpt-6-astra-critical-cyber
- newsOpenAI launches GPT-6 AstraConstellation Research · Sept. 3, 2026constellationr.com/insights/news/openai-launches-gpt-6-astra
Signals
More from the news desk
Cite and share
Use this signal
Citation
Paperclip Index. “OpenAI's GPT-6 Astra became the first model it rates Critical for cyber capability.” Signal SG-0009. Reported Sept. 3, 2026; updated Oct. 7, 2026. https://paperclipindex.com/signal/SG-0009