On September 3, 2026, OpenAI unveiled GPT-6 Astra, describing it as its most intelligent and aligned AI model yet.
Within hours, one number began dominating the conversation: 99.9%.
That is the score OpenAI reports for Astra on ARC-AGI-3, a benchmark designed to test how well AI systems can solve novel, unfamiliar environments and demonstrate what ARC Prize calls agentic intelligence.
OpenAI also highlighted a statement from ARC Prize Foundation’s Greg Kamradt saying Astra surpassed the benchmark’s human action-efficiency baseline on 96% of levels and effectively reached human parity on the benchmark.
That sounds like a major milestone. But there is an important detail behind the 99.9% figure.
ARC Prize reports that Astra’s best observed score increased from 62.7% to 99.9% when it was run with OpenAI’s Provider Adapter harness. OpenAI says its Responses API harness changes two settings intended to better match real-world performance, and that those changes were not specifically designed to target ARC-AGI-3.
So the story is not simply “Astra scored 99.9% and therefore AI has become as intelligent as humans.”
The reality is more interesting. GPT-6 Astra represents a substantial jump in AI capability, but the benchmark result needs to be understood in the context of how the evaluation was run and what ARC-AGI-3 actually measures. Here’s what we know.
What Is GPT-6 Astra?
GPT-6 Astra is OpenAI’s newest flagship AI model, designed for much more than traditional chatbot conversations. OpenAI says Astra is state-of-the-art across areas including:
- Computer use
- Web browsing
- Software engineering
- Cybersecurity
- Scientific research
- Mathematics
- Professional knowledge work
- Multi-step automation
The major shift is that Astra is designed not merely to explain how something can be done, but to carry out the work itself. It can interact with software, browse websites, fill forms, work with documents and spreadsheets, write and test code, analyze data and complete multi-step workflows.
That makes Astra much closer to an AI agent than a conventional question-and-answer chatbot — the same shift we broke down in our earlier piece on what agentic AI actually is.
The 99.9% ARC-AGI-3 Score Explained
This is the part that has generated the most confusion online. ARC-AGI-3 is developed by the nonprofit ARC Prize Foundation. Its goal is to measure forms of intelligence that remain difficult for AI systems, particularly the ability to deal with novel environments rather than simply relying on memorized patterns.
OpenAI reports that GPT-6 Astra achieved a 99.9% score on ARC-AGI-3. But ARC Prize’s published material adds important context. The foundation reports that Astra’s best observed score on the semi-private ARC-AGI-3 evaluation increased from 62.7% to 99.9% when using OpenAI’s Provider Adapter harness.
In other words, the testing configuration matters. OpenAI says its Responses API harness modifies two settings to better reflect real-world model performance. OpenAI also explicitly notes that those changes were not designed specifically for ARC-AGI-3.
That means the most accurate way to describe the result is:
“GPT-6 Astra achieved a reported 99.9% on ARC-AGI-3 when evaluated with OpenAI’s Responses API/Provider Adapter configuration, while ARC Prize reports a 62.7% result under the other observed configuration. Both numbers are part of the published evaluation story.“
Does 99.9% Mean Astra Has Reached Human Intelligence?
No — not by itself. This is one of the most important distinctions to make. A benchmark measures a particular set of capabilities under particular conditions. A near-perfect score on one benchmark does not demonstrate that an AI system can perform every type of intellectual task humans can perform.
ARC Prize itself describes ARC-AGI as a benchmark intended to reveal gaps between what is easy for humans and difficult for AI. It also says that new ideas are still needed to reach AGI.
So it would be misleading to turn the 99.9% figure into: “AI is now officially as intelligent as humans.”
A more defensible interpretation is that Astra has demonstrated an extraordinary level of performance on a benchmark designed specifically to challenge AI systems with novel environments. That is significant. It just isn’t the same thing as proving AGI.
Why the 99.9% Result Still Matters
The benchmark controversy should not obscure how impressive the underlying result is. OpenAI reports Astra at:
- 99.9% on ARC-AGI-3
- 98% on FrontierMath Tier 4
- 96% on GPQA Diamond
- 72.6% on OSWorld 2.0
- 100% on ExploitBench
- 88% on SRE-Bench in a single attempt
These results span very different areas. And Astra’s improvements aren’t limited to theoretical benchmarks.The model is designed to perform actual computer-based work.
Astra Can Actually Use a Computer
This may ultimately be more important than the 99.9% headline. OpenAI says Astra can perform tasks such as:
- Filling out online forms
- Updating customer records
- Organizing calendars
- Conducting online research
- Drafting documents
- Creating spreadsheets and presentations
- Analyzing scientific data
- Generating plots
- Creating websites
- Testing websites
- Installing and testing software
- Troubleshooting problems on screen
OpenAI’s demonstrations also show Astra working on tasks involving PCB design, game development, Blender, Power BI, legal-document formatting and other professional workflows.
This represents a fundamental change in how people can interact with AI.
Instead of: “Tell me how to do this.“
The goal becomes: “Do this for me.“
Astra Is Faster at Computer Tasks Too
OpenAI reports that Astra scored 72.6% on OSWorld 2.0, compared with 65.7% for GPT-5.6 Sol.
More importantly, OpenAI says Astra achieved its performance in approximately 40 minutes per task, compared with about 75 minutes for GPT-5.6 Sol in the reported latency simulations.
That’s where the “agentic” part becomes particularly important. If an AI can reason through a task and operate a computer, but takes hours to finish something, its usefulness is limited. If it can perform the same work substantially faster, the economic impact becomes much larger.
Astra Is Also a Major Coding Upgrade
OpenAI describes GPT-6 Astra as its best software-engineering model to date. Its published results include:
- 57.9% on Terminal-Bench 4.0
- 74.1% on DeepSWE v1.1
- 53.3% on FrontierCode 1.1 Main
- 63.9% on internal database migration tasks
The important development isn’t simply that Astra can generate code. It can operate within a software-development environment, inspect existing code, run tests, troubleshoot problems and continue working across long tasks.
OpenAI is also introducing improvements to Codex that allow Astra to preserve and retrieve context across context windows, helping it retain important details during lengthy coding sessions. That could make long-running AI coding agents considerably more useful.
The Cybersecurity Story May Be Even Bigger
While the internet has focused on the 99.9% benchmark, another Astra statistic could have much bigger real-world consequences. OpenAI says GPT-6 Astra is the first model to meet its Critical cybersecurity capability threshold.
During testing without production safeguards, Astra achieved:
- 100% on ExploitBench
- 42.4% on ExploitGym
- 88% on SRE-Bench in one attempt
- 99.2% within four attempts on SRE-Bench
Even more significantly, Astra discovered and successfully used two previously unknown zero-day vulnerabilities during evaluation.
OpenAI says it is disclosing those vulnerabilities to the relevant maintainers. The model was also able to use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits against hardened operating systems during expert-led testing.
That’s an extraordinary capability. And it explains why OpenAI isn’t simply releasing every cybersecurity capability without restrictions.
This connects directly to what we covered in our piece on AI-powered cyberattacks in 2026 — Astra’s Critical classification is essentially confirmation that the trend we described there has crossed a real, measurable threshold.
Why OpenAI Is Restricting Some Cybersecurity Capabilities
There is a difficult trade-off here. The same AI that can help security researchers discover vulnerabilities faster could potentially help attackers do the same thing.
OpenAI says the publicly released version of Astra is intended to support defensive work such as secure code review and patching, while more advanced cybersecurity capabilities are being introduced through controlled access programs.
That creates an unusual situation: Astra is powerful enough to discover serious security vulnerabilities, but that very capability is one reason access to its strongest cybersecurity functions has to be restricted.
This may become one of the defining challenges of increasingly capable AI agents.
Astra Is Making Progress in Mathematics Too
Astra’s scientific capabilities are another reason this release matters. OpenAI says Astra helped produce new results concerning prime gaps. One result establishes that infinitely many pairs of primes occur within a gap of 186, improving the previous bound of 240.
OpenAI also says Astra improved a term in a bound concerning unusually large gaps between primes — a term that had remained unchanged for more than 80 years.
Importantly, OpenAI has published proofs and supporting research materials for these results. That makes these claims different from simply reporting an internal benchmark score. Mathematical claims can ultimately be examined independently by researchers.
Is GPT-6 Astra AGI?
This is where the debate becomes philosophical as well as technical.
OpenAI President Greg Brockman has described Astra’s arrival as the beginning of the “AGI era” and has said that he personally believes people may eventually look back at Astra as the point when AGI was achieved.
But that is not the same as saying there is universal agreement that AGI has been achieved. There is no single universally accepted benchmark that settles the question.
The ARC Prize Foundation, for example, has its own definition of AGI centered on human learning efficiency, while OpenAI’s broader discussion of AGI concerns systems capable of performing economically valuable work at or beyond human levels.
Astra clearly moves the conversation forward. But whether it should be called AGI is still a matter of definition, evidence and interpretation.
How Much Does GPT-6 Astra Cost?
For developers using the OpenAI API, Astra is priced at:
- $10 per million input tokens
- $50 per million output tokens
OpenAI also offers a Fast mode that can provide up to twice the speed at twice the Standard processing price. That makes Astra significantly more expensive than many conventional AI models, but its value proposition is different.
The goal is not simply generating text more cheaply. The goal is completing complex work.
For businesses, the relevant question may therefore become:
How much does it cost to complete a task with Astra compared with paying a human to perform the same workflow?
Who Can Use GPT-6 Astra?
OpenAI says Astra is rolling out in stages. It is becoming available to:
- ChatGPT Plus users
- ChatGPT Pro users
- ChatGPT Business users
- ChatGPT Enterprise users
- OpenAI API developers
- Microsoft Azure customers
- Amazon Bedrock customers
Enterprise administrators can enable Astra for their workspaces, although access is off by default at launch. The strongest cybersecurity capabilities are subject to additional controls.
What Does GPT-6 Astra Mean for Ordinary Users?
For most people, the most important change isn’t going to be a benchmark score. It is the possibility of delegating complete digital workflows. Imagine asking an AI to:
- Research several products.
- Compare prices and specifications.
- Create a spreadsheet.
- Summarize the best options.
- Draft an email.
- Build a presentation.
- Create a simple website.
- Test that website.
- Revise it based on the results.
Instead of using several applications manually, the AI becomes the layer connecting those applications together. That’s the bigger story behind Astra.
The Real Question Isn’t Whether Astra Is “Human”
The 99.9% number makes for a great headline. But it may not be the most important thing about GPT-6 Astra.
The more consequential development is the combination of: reasoning + computer use + coding + browsing + long-context memory + professional workflows + scientific capability.
An AI that scores highly on a benchmark is interesting. An AI that can take a vague objective, navigate software, make decisions, write code, analyze information and deliver a finished result is potentially much more transformative. And that is what makes Astra different.
GPT-6 Astra: Final Verdict
GPT-6 Astra is unquestionably a major step forward in AI capability.
The 99.9% ARC-AGI-3 score is real, but it should not be presented as a simple universal measurement of “human intelligence.” The score depends on the evaluation configuration, and ARC Prize’s published results show a substantial difference between the observed configurations.
That nuance doesn’t make Astra less impressive. If anything, the broader evidence makes the release more interesting.
Astra is substantially stronger at computer use, coding, professional workflows, mathematics, science and cybersecurity than its predecessor. It can perform tasks rather than merely discuss them, and OpenAI says it has reached a level of cybersecurity capability serious enough to trigger its highest risk classification.
So is GPT-6 Astra AGI?
We don’t have a universally accepted answer yet. But one thing is becoming increasingly difficult to dispute: AI is moving from systems that answer questions toward systems that can actually do the work.
And GPT-6 Astra may be one of the clearest demonstrations yet of where that transition is heading.
Frequently Asked Questions
GPT-6 Astra is OpenAI’s latest flagship AI model, designed for advanced reasoning, computer use, browsing, coding, cybersecurity, scientific research and professional workflows.
Yes. OpenAI reports a 99.9% ARC-AGI-3 score. ARC Prize also reports that Astra’s observed score increased from 62.7% to 99.9% when using OpenAI’s Provider Adapter configuration. The testing setup therefore matters when interpreting the result.
No. A benchmark score — even an exceptional one — does not by itself prove that an AI system has achieved artificial general intelligence.
OpenAI says Astra reached its Critical cybersecurity threshold after demonstrating very strong exploit-development and reverse-engineering capabilities, including discovering and using two previously unknown zero-day vulnerabilities during testing.
OpenAI’s API pricing is $10 per million input tokens and $50 per million output tokens. Fast mode is available at twice the Standard processing price.
OpenAI says Astra is rolling out to ChatGPT Plus, Pro, Business and Enterprise users, as well as through the OpenAI API, Microsoft Azure and Amazon Bedrock.
According to OpenAI’s published evaluations, Astra substantially outperforms GPT-5.6 Sol across several computer-use, coding, scientific and cybersecurity benchmarks. However, individual performance varies by task and evaluation setup.
Astra is designed to operate computers and complete multi-step workflows. Instead of simply telling users what to do, it can interact with software, browse the web, work with documents, write and test code, and perform other actions needed to complete a task.

Leave a Reply