OpenAI releases GPT-6 Astra, the first model it rates Critical for cybersecurity under its Preparedness Framework
OpenAI published GPT-6 Astra on 3 September 2026 and said it is the first model to meet the Critical threshold for cybersecurity capability under the company's Preparedness Framework.
On benchmarks Astra saturates FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9% and ExploitBench at 100%, against 78.5% for GPT-5.6 Sol. It reaches 42.4% on ExploitGym versus 30.3% for Sol, and on SRE-Bench, which measures reverse engineering of binaries, it solved 88.0% of tasks in a single attempt and 99.2% within four, against 55.9% and 68.7% for Sol. To rule out contamination from historical vulnerabilities, OpenAI built an internal "ExploitBench (June-August 2026)" set from the previous three months; during that evaluation Astra discovered and used two previously unknown zero-days, which OpenAI says it is disclosing to the maintainers. The API model id is gpt-6-astra, standard pricing is $10 per million input tokens and $50 per million output tokens, and a Fast mode runs at up to 2x the speed for 2x the price. It rolled out first to a limited set of organisations, then to ChatGPT Plus, Pro, Business and Enterprise users and through the OpenAI API, Microsoft Azure and AWS Bedrock; for Enterprise workspaces access is off by default at launch.
Astra ships refusing the more advanced cyber tasks such as building proof-of-concept exploits, but OpenAI states plainly that it plans to expand access and relax safeguards through Daybreak in the coming weeks. If you are building agents on it, assume misalignment monitoring runs in production for Astra-class models: a suspicious action can pause a task for review in ChatGPT or Codex, and in the API the task simply stops. The speed gap matters too — in OSWorld 2.0 latency simulations Astra scored 72.6% at roughly 40 minutes per task against Sol's 65.7% at roughly 75 minutes.
