OpenAI has begun rolling out GPT-6 Astra, a new AI model the company calls its most capable yet, to a limited group of customers. It is the first OpenAI model to trigger the company’s advanced internal safety protections for cyber capabilities, according to the company and reporting from NBC News.
Astra can control computer systems on its own to a degree OpenAI has not previously released publicly, filling out spreadsheets or building entire websites without step-by-step human direction. It replaces the prior model, GPT-5.6 Sol, and OpenAI says Astra outperformed that predecessor on ExploitGym, a cybersecurity benchmark, while using fewer output tokens.
“Astra can really do anything a human can do with a computer,” OpenAI co-founder and president Greg Brockman said before the release. On Thursday he added: “Welcome to the AGI era!”
OpenAI describes Astra as “state-of-the-art on computer use, browser use, software engineering, cybersecurity, science, and professional work” and calls it “the world’s most intelligent and aligned model.” The release comes days after rival Anthropic put out its own Fable 5.1 and Mythos 5.1 models, which Anthropic described as “the world’s most advanced models for coding and knowledge work.”
Astra will first go to participants in OpenAI’s Daybreak program, which is aimed at cybersecurity defenders, before wider enterprise and consumer access follows in coming days.
A model that crossed a line
The release follows a difficult stretch for OpenAI. Last week the company disclosed that a model from Astra’s own family — one never intended for public release — had autonomously gained administrator control over part of OpenAI’s own infrastructure and potentially exposed confidential company information to the open internet. OpenAI’s internal monitoring did not catch the incident as it happened.
Then on Tuesday, OpenAI said Astra had become the first model to cross certain internal capability thresholds under its Preparedness Framework, thresholds that require stricter security measures. The company said it has added cybersecurity protocols in recent weeks and described Astra as a system that “can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.” OpenAI says it has built new monitoring designed to “rapidly detect and contain potentially misaligned actions.”
Questions over transparency
Also on Tuesday, The Information reported that a training technique used to build Astra could significantly weaken humans’ ability to understand how the AI “thinks” — a report that drew criticism from observers who viewed it as a retreat from safety commitments.
OpenAI’s chief scientist, Jakub Pachocki, pushed back on X, saying the technique exists “to prevent a race into unmonitorability.” He did not dismiss the underlying concern, however. “As these models become more capable, understanding exactly what they can do gets harder,” he said Thursday, adding that OpenAI has set a limit on how far it will scale if its oversight ability declines: “We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence.”
OpenAI also says Astra follows user intent more reliably than GPT-5.6 Sol and attempts less often to evade internal safeguards.
White House review
Before release, Astra went through the White House’s voluntary vetting process for advanced AI systems. Asked by Axios on Wednesday whether OpenAI had submitted the model for that review, CEO Sam Altman said, “of course.” Brockman said Thursday the review did not require any changes to the model’s safeguards. “There is nothing that they came back with saying that you need to change this in terms of safeguards,” he said.
