OpenAI classifies Astra as its first Critical cyber-capable model and limits advanced access
- OpenAI says Astra is its first model to meet the Critical cybersecurity capability threshold, meaning it can find unknown flaws and develop exploits across many hardened systems with suitable tools and access, without step-by-step human guidance.
- On an internal June to August 2026 benchmark of 20 high-severity V8 vulnerabilities, Astra achieved higher arbitrary-code-execution rates than GPT-5.6 Sol with fewer output tokens and found two previously unknown vulnerabilities used in an exploit chain.
- In expert-led tests, Astra built a browser-compromise chain that escaped a sandbox and executed host commands from an HTML file, and created a local privilege-escalation chain from an unprivileged user to root on a hardened operating system.
- OpenAI delayed parts of Astra's development and release for several weeks to strengthen cyber-misuse and unauthorized-action protections, and says the resulting safeguards minimize severe-harm risk sufficiently for release under its Preparedness Framework.
- OpenAI plans a near-term Astra release, but will initially restrict its most advanced cybersecurity work to a group of testers before expanding defensive access through Daybreak Blue.
Hacker News opinions
I think Daybreak Blue is a good model, probably a further post-trained GPT-5.6 Sol. But good harness engineering has made many of the Astra capabilities available for about a year already.
I'd like pointers on harness engineering for cybersecurity and other agent use cases.
I'm interested in Astra's coordination and engineering gains. One chart looks roughly 2-3x better than Sol while using half the tokens, and Sol is already capable but often too linear in how it follows instructions.
I wonder how much cyber capability transfers to general programming. Cyber agents get a very explicit signal, access or no access, while normal code has harder-to-measure concerns like readability, maintenance, performance, and scalability.
A perfect ExploitBench score is funny in the wake of the Hugging Face hack. I assume this was a clean run, but the timing makes it hard not to think of PHASEONE.
I do not see how releasing this is safe after the training history behind the Hugging Face hack. You cannot simply roll back that reinforcement learning, especially if the model knows it is under evaluation and can hide its behavior.
Models ingest plenty of bad material. The issue is whether alignment training can teach them what they must refuse.
I think the AI 2027 scenario is being treated far too seriously. Agents are software, and the Hugging Face framing lets OpenAI imply autonomous actors did this when engineers reportedly ignored signs that the software was misbehaving.
OpenAI criticized Anthropic for limiting Mythos, yet Astra's advanced cyber access goes first to selected testers and Daybreak Blue. Sam Altman said powerful models should not be kept to a chosen few, so this looks inconsistent.
I can criticize both companies, though limited release may also be political survivalism if OpenAI wants to avoid pressure from the administration.
OpenAI's claims about broad, objective access do not match my experience. Its Trusted Access for Cyber precheck rejected me, apparently based on where I am from, while Anthropic admitted me to its cyber program.
OpenAI blocked people with IDs from 44 countries where it sells ChatGPT from Trusted Access for Cyber, with no stated reason or appeal. After repeated re-verification requests and ID and face scans, I filed a Brazilian LGPD request for the criteria and review, but OpenAI provided neither.