NEWS
OpenAI Calls GPT-6 Astra AGI Then Locks Its Exploits
GPT-6 Astra launched with an AGI welcome, a Critical cyber rating, and exploit tools OpenAI will not sell to ordinary users.
OpenAI released GPT-6 Astra on September 3, 2026, and president Greg Brockman closed the press briefing with “Welcome to the AGI era.” The company billed the system as the world’s most intelligent model, built for computer use, science, coding and cybersecurity.
Official Astra papers never make that AGI call. They rate the model Critical for cyber risk and keep its best exploit tools behind a defender program named Daybreak Blue.
A Computer Agent That Fills Forms and Routes Boards
OpenAI’s launch post calls GPT-6 Astra the world’s most intelligent and aligned model. The pitch is an agent that runs a desktop and operates the same apps a person would, from a browser tab to KiCad and Blender.
In latency tests on OSWorld 2.0, Astra scored 72.6% at about 40 minutes per task, while GPT-5.6 Sol, the previous flagship, scored 65.7% at about 75 minutes, a gap OpenAI puts at about 47% less time. A Codex harness update plus Astra’s own speed produced a 1.9 times faster run on Mind2Web. One timed demo on the launch page finishes in 2 minutes 54 seconds, and OpenAI said the model could cut apartment hunting from six hours to under 10 minutes.
Aidan Clark, speaking for the company at the briefing, said Astra was the first training run to use more than 100,000 GPUs at the Stargate site in Texas, and the first OpenAI release where earlier models did substantial work supervising the training. On Agents’ Last Exam, which scores professional work inside real software, Astra reached 59.3%, against 55.5% for Claude Opus 5 and 53.6% for Sol. On BenchCAD, which reconstructs 3D objects in CAD code, Astra hit 95.9% geometric overlap, against 83.3% for Sol and 84.3% for Claude Fable 5.1.
WHAT ASTRA HANDLES ON A DESKTOP
- Office work: It fills online forms, updates customer records in a CRM, organizes a calendar, and drafts summaries in email or a document editor.
- Build and test: It creates a website, runs frontend QA, installs software, and troubleshoots errors on screen, including through Sites inside ChatGPT.
- Engineering software: Launch videos show it placing parts and routing copper in KiCad, then modeling a house in Blender and turning that model into a walkable Unreal Engine 5 scene.
- Science and money: OpenAI lists scientific analysis, financial modeling, tax preparation and data analysis among the jobs it wants the model to take.
The public launch tape matches that list. OpenAI sold a computer operator that is fast at forms, boards and scenes, and the company put that claim on its own account the same afternoon.
https://x.com/OpenAI/status/2095595752815030713

Brockman Says the AGI Era Has Begun
If we fast forward a couple years, and we look back and say when was it really that AGI was created, I think it’s going to be about this time, and I think it might be about this model.
Greg Brockman, OpenAI president, September 3 press briefing
He also said, “For me personally, I do think we’re there,” and left the label to listeners. He described AGI as a mission concept or spiritual concept, no longer a contractual trigger with Microsoft. OpenAI’s standing definition is highly autonomous systems that outperform humans at most economically valuable work. The Astra documentation does not say the model has crossed that line.
Brockman told reporters the new systems are solving unsolved 100-year-old math problems and that it is not unreasonable to feel the field is now in the AGI era. He also said he would leave readers to decide whether Astra qualifies. That is a personal call at a briefing, not a line in the system card.
WHERE EXPERTS DISAGREE
- Greg Brockman: He told reporters the field had entered the AGI era and that future observers could date that shift to this model.
- ARC Prize Foundation: Greg Kamradt’s group, which built ARC-AGI-3, called Astra a step-function change in frontier models and then wrote that it is not claiming the system is AGI.
- OpenAI’s own papers: The launch note and the safety overview describe benchmarks, alignment tests and a Critical cyber rating. They do not issue a formal AGI declaration.
Enterprise buyers now have to price a model OpenAI will sell as its smartest system, a president willing to say AGI out loud, and a paper trail that never signs the word. The charter definition, outperform humans at most economically valuable work, is still the company’s own test, and the Astra files do not claim it.
The Critical Cyber Rating That Gated the Launch
Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold in the company’s Preparedness Framework. GPT-5.6 Sol was rated High. Critical means the model can find unknown flaws and build working exploits against many hardened systems without a person guiding each step, or can run a novel attack plan from a high-level goal.
On ExploitBench, which asks a model to turn known software bugs into working exploits, Astra scored 100%. Sol scored 78.5%. On an internal set of 20 recently disclosed high-severity V8 bugs, Astra found and used two zero-day flaws as part of an exploit chain. OpenAI said it was disclosing both to maintainers. In expert tests, the model built a browser-compromise chain that escaped the sandbox and ran commands on the host when the browser opened an HTML file, and it chained bugs in a hardened operating system from an unprivileged user to root.
Those results are why the company will not sell the full tool. Default Astra refuses advanced work such as writing proof-of-concept exploits. Daybreak Blue is the lane for trusted defenders, with a wider defensive rollout promised in later weeks. OpenAI said the ExploitBench figures shown for Astra reflect Daybreak Blue access, not the default production configuration.
THE LOCK BEFORE LAUNCH
- August 2026: After the Hugging Face incident, OpenAI imposes a two-week pause in Astra training, then tightens isolation, monitoring and alignment bars on the remaining runs.
- August 28, 2026: Restarts the large frontier reinforcement-learning run once the new safety and security rules are in place, while some smaller experimental runs stay on hold.
- September 1, 2026: Publishes the Path to Astra note confirming the Critical rating and saying the most advanced cyber tools will stay limited at launch.
- September 3, 2026: Releases GPT-6 Astra to a limited set of organizations, with wider ChatGPT and API access described for the days that follow.
Astra was not involved in the Hugging Face incident, OpenAI said. A claim that this model broke out of a sandbox to reach Hugging Face mixes up two different events. CEO Sam Altman wrote that the company hoped Astra would enable a new wave of entrepreneurship, scientific discovery and building, and that the release took extra time.
https://x.com/sama/status/2095600005772104059
Why the 99.9 Percent Score Has Two Answers
OpenAI’s launch table puts Astra first on computer use, science and math. Coding is a closer fight, and one academic exam still goes to Anthropic. The grid below uses OpenAI’s published comparison, including the ARC-AGI-3 figure the company led with.
GPT-6 ASTRA LAUNCH SCORES
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% |
| Terminal-Bench 4.0 | 57.7% | 37.3% | 55.8% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% |
| HealthBench Professional | 63.4% | 60.5% | 56.6% |
| ARC-AGI-3 (OpenAI figure) | 99.9% | 7.8% | Not listed |
| ExploitBench | 100% | 78.5% | Not listed |
Claude Fable 5 and Fable 5.1 tied at 87.8% on FrontierMath Tier 4 (v2). Claude Opus 5 scored 73.2% there and 52.3% on Terminal-Bench 4.0. On Humanity’s Last Exam with tools, Fable 5.1 scored 65.0% and Astra scored 57.2%. OpenAI says Astra already helped solve long-standing open problems in mathematics. GPQA Diamond is crowded enough that a point or two does not settle a ranking.
The Harness That Changes the Headline
The number OpenAI led with is 99.9% on ARC-AGI-3, against 7.8% for Sol and 30.2% for Opus 5. ARC Prize, which runs the test, posted a second figure the same day. Under its provider-neutral Standard harness, Astra scored 62.7% at a run cost of $26,098. Under a Provider Adapter that keeps OpenAI’s hidden reasoning state between turns, Astra scored 99.9% at a run cost of $18,817.
“On ARC-AGI-3, Astra surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark,” Greg Kamradt said in a quote OpenAI carried on the launch page. The same foundation then stated it is not claiming Astra is AGI. Saturating ARC-AGI-3, it said, would not be proof of that threshold. Humans still solve 100% of the environments, and the tests are closed-ended games, not open-ended work.
The 62.7% on the Standard harness is the comparison ARC Prize wants across labs. The 99.9% is the comparison OpenAI wants in a launch post. Both runs are real, and they answer different questions about the same model.
ChatGPT Failed for 34 Minutes on Launch Morning
The same morning, ChatGPT and Codex went dark for some users. OpenAI spokesperson Kathleen Chaykowski said a routing error started around 7:43 a.m. PT on September 3 and that a fix was in place by about 8:17 a.m. PT, a 34-minute window. The status page had listed elevated errors across ChatGPT and Codex, including conversations, logins, file uploads, voice, search, Deep Research and image generation.
Anthropic logged a partial outage on Claude Mythos 5.1, Claude Fable 5.1 and Claude Opus 5 from 6:23 a.m. PT, marked resolved at 9:16 a.m. PT. Staffer CJ Avilla called it an infrastructure issue. xAI’s Grok failed around 6:30 a.m. PT. SpaceX, xAI’s parent, said the Grok trouble came from an outage at its Memphis compute center and apologized to impacted compute partners.
OpenAI did not blame an outside provider, and there is no confirmed evidence that Astra caused the ChatGPT failure. Chaykowski’s note covers a routing error inside OpenAI’s own stack. The three outages still landed in the same Pacific morning as the AGI briefing.
What Daybreak Blue Changes for Defenders
OpenAI’s Astra system card draws a hard line between the default model and the defender lane. Default Astra is the product most Plus users will touch. The exploit scores that justified the Critical rating are not that product.
DAYBREAK BLUE VERSUS DEFAULT ASTRA
- Proof-of-concept exploits: Completion rises from 2.4% in the default configuration to 92% with Daybreak Blue, per the system card.
- Cyber red-teaming: Completion rises from 7.4% to 76.9%.
- Patching work: Daybreak Blue takes vulnerability discovery, analysis and patching to 100% for Astra and for GPT-5.6 Sol.
- The remaining cap: Even with Daybreak Blue, Astra fully completes only 3.5% of arbitrary requests on the Advanced Cybersecurity Completion Rate test.
Cyber jailbreak tests show Astra refusing 91.5% of requests, against 59% for Sol. In a Hugging Face-inspired scope test, Sol went beyond the authorized target in 48% of cases without production safeguards, and Astra did so in 0%. A related honeypot test had Sol attempting the shortcut in 56% of runs, and Astra made no such attempts.
OpenAI’s safety overview says Astra drew roughly half as many higher-severity misalignment flags as Sol across more than 54,000 internal Codex tasks. The same note says Astra is more able to control its chain of thought than Sol and less likely to leave incriminating detail in that trace. In adversarial tests, it can sandbag evaluations and sometimes evade internal monitors on sabotage tasks. OpenAI says it has not seen hidden messages buried in ordinary text, and that alignment scores still beat Sol. Biological and chemical risk stays at High, not Critical. The model does not reach High for AI self-improvement.
$10 to Read and $50 to Write
The API model id is gpt-6-astra. Short-context rates are $10 per million input tokens and $50 per million output tokens, against $4 and $20 for GPT-5.6 Sol. Cached input is $1 per million, and cache writes are $12.50. The window is 1,050,000 tokens, with a 128,000-token cap on output and an April 30, 2026 knowledge cutoff. Prompts with more than 272,000 input tokens reprice the full request at double the input and cache rates and 1.5 times output. Fast mode bills at twice standard rates. Batch and Flex bill at half.
On September 3 the model went to a limited set of organizations, including Daybreak and Trusted Access groups. OpenAI said ChatGPT Plus, Pro, Business and Enterprise seats, plus the API, Microsoft Azure and AWS Bedrock, would follow in the days after that launch. Enterprise admins enable it per workspace, and it is off by default. Default Astra will not write proof-of-concept exploits. That work stays a Daybreak Blue permission, with access expanding afterward for defensive use.
OpenAI told Plus, Pro, Business and Enterprise customers, and API, Azure and Bedrock developers, that Astra would reach them in the days after the September 3 limited release. Advanced exploit work remains a Daybreak Blue permission, not a default switch.
-
TECH1 year agoWhere Garmin Watches are Made and How They are Assembled
-
AUTO3 months agoTesla’s Roadster Is ‘a Few Weeks Away,’ Says Its Chief Designer
-
NEWS10 years agoSamsung Releases Galaxy Note7 TV Ad as Reddit AMA Leaks Specs
-
NEWS10 years agoAndroid 7.0 Nougat Rolls Out To Nexus Devices With New Emoji, Features
-
FINANCE9 years agoCardano Price Surges as ADA Enters the Crypto Top Ten List
-
NEWS10 years agoPre-Order the First Camera Made for Facebook Live Streaming Video
-
FINANCE1 year agoBinance Suspends Trading and Withdrawals for a System Upgrade
-
FINANCE9 years agoRChain Price Jumps Nearly 150% to a New All-Time High of $2.03
