OpenAI officially launched its next-generation flagship model, GPT-6 Astra, declaring the system to be the most intelligent and aligned artificial intelligence in the world. Arriving in the wake of synchronized service disruptions across OpenAI, Anthropic, and Grok, the rollout immediately sparked industry-wide debate as leadership asserted that artificial general intelligence has effectively arrived[3].
The release marks a dramatic shift toward autonomous agentic computing. Unlike prior iterations centered primarily on conversational text, Astra is engineered to directly operate software environments, run multi-step computer tasks, draft complex codebases, and conduct technical research with minimal human oversight[1, 2].
The Claim That AGI Has Arrived
The most consequential assertion surrounding the release came directly from OpenAI leadership. In an interview reported by The Washington Post and The Guardian, OpenAI president Greg Brockman stated that he believes Astra qualifies as artificial general intelligence, meeting the company's internal benchmark of systems that outperform humans across most economically valuable tasks.
I do leave it up to the reader to decide for themselves if this qualifies for them. I think we're there.
Greg Brockman, President of OpenAI
Brockman told reporters that future historians looking back to pinpoint when AGI was created will likely point to this moment and this model. OpenAI chief executive Sam Altman offered a more cautious operational framing on Fox Business Network, describing Astra as setting a new safety standard after development delays while confirming it pushes the company directly into the AGI era[7].
Agentic Benchmarks and Real-World Workflows
Astra's technical claims center on deep system autonomy rather than marginal conversational gains. OpenAI's published evaluations show the model achieving unprecedented scores across several difficult problem domains designed to test reasoning limits. It scored 72.6 percent on the OSWorld 2.0 computer-use benchmark, completing tasks in approximately 47 percent less time than its predecessor, GPT-5.6 Sol.
The system also saturated specialized evaluations, recording a 97.6 percent score on FrontierMath Tier 4 and 99.9 percent on ARC-AGI-3 under OpenAI's testing adapter harness. Early corporate trial partners reported practical workflow shifts[1]:
- Financial document engine Legora processed 41 complex filings in minutes, successfully spotting four planted anomalies and boosting overall audit speed by roughly 40 percent.
- Gaming studio Playco developed three interactive game prototypes from a single foundation, noting a 50 percent drop in manual bug fixes compared to earlier models.
- Legal intelligence platform Harvey reported that Astra distinguished source records from unsubstantiated assumptions in judicial drafting.
Third-party testing platforms added essential nuance to these headline figures. According to an analysis by Artificial Analysis, Astra matches Anthropic's Claude Fable 5.1 in coding agent benchmarks at lower token costs, yet on the broader Intelligence Index, Astra registered a 61.2, roughly flat compared to GPT-5.6 Sol's 60.9. The performance leaps are heavily concentrated in agentic tasks and tool execution rather than pure abstract question-answering[8].
| Benchmark / Metric | GPT-5.6 Sol | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|---|
| OSWorld 2.0 (Computer Use) | Baseline | 72.6% | Trailing |
| FrontierMath (Tier 4) | Incomplete | 97.6% | Unsaturated |
| ExploitBench (Cybersecurity) | 78.5% | 100.0% | Gated |
| Humanity's Last Exam (with tools) | N/A | 57.2% | 65.0% |
| API Input / Output (per 1M tokens) | $4 / $20 | $10 / $50 | $10 / $50 |
The 'Critical' Cyber Threshold and Safety Lockdowns
The dark side of Astra's technical capability lies in cybersecurity. In its official Preparedness Framework documentation, OpenAI revealed that Astra is the first model in history designated at the Critical cybersecurity capability tier. Under internal safety testing, Astra demonstrated an autonomous ability to discover zero-day vulnerabilities in hardened software architectures and engineer working exploits without human direction, achieving a 100 percent score on ExploitBench.
Because of this capability, OpenAI delayed Astra's public rollout by several weeks to build strict architectural guardrails. The company disclosed that while Astra was not involved in the earlier security breach at Hugging Face, lessons from that incident prompted deeper restrictions. Unrestricted access to advanced offensive cyber workflows remains completely blocked for the general public, accessible only to an alpha cohort within OpenAI's Daybreak Blue security framework[12].
To balance the hazard, OpenAI announced Daybreak for Frontline Defenders, a 1 billion dollar initiative providing defensive cyber institutions, critical infrastructure operators, and essential public services with frontier AI protection.
Pricing, Access Tiers, and Market Counterpoints
Frontier power carries frontier expenses. Via the OpenAI API, GPT-6 Astra costs 10 dollars per million input tokens and 50 dollars per million output tokens, matching the high watermark set by Anthropic's flagship models and representing a 2.5-fold increase over GPT-5.6 Sol. OpenAI defended the pricing by arguing Astra requires fewer total tokens to finish complex jobs, though industry observers note that monitorability inside long reasoning chains has decreased.
Skepticism remains prominent across the AI research community. Independent evaluators cited by DataCamp pointed out that Astra's near-perfect ARC-AGI-3 performance required a specialized, stateful test harness; standard stateless API requests scored significantly lower. Furthermore, Astra trailed Anthropic's Claude Fable 5.1 on the tool-assisted Humanity's Last Exam benchmark, scoring 57.2 percent against Anthropic's 65.0 percent[9].
The launch rollouts will unfold gradually. Enterprise clients via Microsoft Foundry and OpenAI's Trusted Access Program receive initial integration this week, with ChatGPT Plus, Pro, Business, and Codex developer tiers scheduled to follow in subsequent days.
