Summary
On September 1, 2026, OpenAI said its upcoming Astra model is the first to cross the 'Critical' cybersecurity threshold in its Preparedness Framework, meaning it can find previously unknown flaws and develop working exploits without step-by-step human guidance. OpenAI still plans to release Astra soon but will restrict access to its most advanced cyber capabilities and has added extra safeguards and monitoring.
What changed
OpenAI publicly designated Astra as 'Critical' for cyber capability, citing the ability to identify and develop functional zero-day exploits across hardened systems and to plan and execute novel end-to-end attacks from a high-level goal; it paused parts of development and restarted a large frontier RL run on August 28 under new safety requirements.
Why it matters
This is the first time a major lab has self-classified a model at the top cyber-risk tier, setting a precedent for gated release, restricted access, and enhanced monitoring of dual-use capabilities. It raises the bar for how frontier labs disclose and contain offensive-security capability, with direct implications for enterprise trust, regulation, and defensive tooling.
Evidence excerpt
Astra can find previously unknown security flaws and exploit them without step-by-step guidance from humans, which means the model falls under the most advanced category of its Preparedness Framework.