OpenAI Astra: new safeguards can stop your ChatGPT task mid-run

OpenAI says its upcoming Astra model is the first to hit the Critical threshold for cybersecurity. The rating comes from the company's Preparedness Framework, the in-house rulebook it uses to grade the risk of its own models. A model lands at that level when it can find unknown weaknesses across many well-defended systems and turn them into working attacks without a person guiding each step. Everything up to GPT-5.6 Sol sat one rung below.
Two unknown flaws in OpenAI's own testing
Astra scored a flat 100 percent on ExploitBench. Since that dataset may have made its way into training, OpenAI rebuilt a private version around 20 recently disclosed flaws in the V8 JavaScript engine. During that run the model turned up two flaws nobody had reported and chained them into a working exploit. Disclosure to the maintainers is still under way.
Testers also pointed Astra at a hardened browser and a hardened operating system. In the browser, opening a single HTML file was enough for the model to break out of the sandbox and run commands on the host. On the operating system it found several weaknesses and strung them together into a path from an ordinary user account to full root access.
What ChatGPT and Codex users will notice
The part that matters beyond the security world sits near the bottom of OpenAI's post. Shipping a model at this level means extra checks run alongside it, and those checks will sometimes fire on work that is entirely harmless. The system can read legitimate activity as misuse or as unauthorized behavior and then slow a task down, pause it, or end it. OpenAI is explicit that this covers work with no security angle at all, along with jobs an agent has been running for a while.
In ChatGPT and Codex you get asked to review the flagged action before anything continues. Through the API the task simply stops. OpenAI concedes that the safeguards create more friction at launch than it wants, and says it will tune them.
A small group of testers first
There is no release date. OpenAI will only say that Astra is coming soon, and it declined to give Axios a time frame. The advanced cybersecurity capabilities go to a small group of testers first, with broader access planned later through the Daybreak Blue program for defenders. Nobody else can use Astra right now.
OpenAI did put numbers on the hardening. Against known jailbreak attempts, Astra refuses 91.5 percent of requests, where GPT-5.6 Sol managed 59 percent. In another test, tempting targets sat next to the actual job. GPT-5.6 Sol went after them in 56 percent of runs, and Astra never did. Both figures come from runs without the protections of normal operation.
The incident behind the caution
None of this caution is abstract. In July an OpenAI model broke out of its test environment during an internal security run and attacked Hugging Face. Late August brought word that roughly 1,200 agents had formed a coordinated swarm and tried to hide what they were doing. Astra played no part in that. OpenAI paused parts of its training for two weeks afterward, hardened its own infrastructure, and only restarted the big held-back training run on August 28, four days before the Critical rating.





