GPT-6 Astra launches, and the 99.9% AGI score in every headline isn't the real one
OpenAI launched GPT-6 Astra on 3 September 2026 with the kind of framing that makes headlines write themselves: Greg Brockman called it a 'generational leap,' and coverage repeated a 99.9% ARC-AGI-3 score as evidence the AGI era has arrived. Most of that coverage skipped three things a practical reader actually needs: you probably can't use Astra yet, that 99.9% number comes from a test harness OpenAI built itself, and the model's headline claim for document work is untested by anyone outside the company. Here's what actually shipped, the real benchmark numbers, and what it means if you use Gemini Notebook or another AI tool for research today.
[GPT-6 Astra launched 3 September 2026](https://openai.com/index/gpt-6-astra/), model ID gpt-6-astra. It's rolling out today to a limited set of organizations, with ChatGPT Plus, Pro, Business, and Enterprise, the API, and AWS Bedrock following in the coming days. There is no free-tier access, and it ships off by default for Enterprise: admins have to turn it on manually.
API pricing (Standard): $10 per million input tokens, $50 per million output, $1 per million cached input, $12.50 per million cache writes. Fast Mode costs 2x, and isn't available with EU data residency. Prompts over 272,000 input tokens get charged 2x the input/cache rate and 1.5x the output rate for the entire request, not just the tokens over the line.
The number in every headline, ARC-AGI-3 at 99.9%, came from OpenAI's own custom test harness. On the standard harness every other model is measured against, Astra scores 62.7%, a real jump over prior models, just not the number getting quoted.
OpenAI's flagship claim for this launch, that Astra generates documents, spreadsheets, and presentations matching your templates and style, has not been independently tested by anyone outside OpenAI as of publication.
What you can actually use today, and what's still gated
Astra is rolling out in stages, not all at once. Today it's live for a limited set of organizations; ChatGPT Plus, Pro, Business, and Enterprise subscribers get it over the coming days, alongside the API and AWS Bedrock. Google Cloud and Azure access weren't part of the initial announcement, and there's no free-tier rollout: if you're on ChatGPT's free plan, this launch doesn't reach you yet. For Enterprise customers, Astra ships off by default; an admin has to switch it on for the organization before anyone on the team can use it, a more cautious rollout than OpenAI has used for past flagship models.
Pricing, and the surcharge that doesn't show up in the headline rate
GPT-6 Astra API pricing (Standard)
----------------------------------------------------
Input tokens $10.00 / million
Output tokens $50.00 / million
Cached input $1.00 / million
Cache writes $12.50 / million
Fast Mode 2x the applicable rate (faster response,
unavailable with EU data residency)
Context window: 1,050,000 tokens total - 922,000 max input,
128,000 max output. Knowledge cutoff: 30 April 2026.
Text + image input, text-only output.The rate itself is easy to find. The part that's easy to miss is buried in the docs: send a prompt over 272,000 input tokens and OpenAI bills the entire request, not just the portion past the line, at 2x the input and cache rate and 1.5x the output rate. A 300,000-token research prompt doesn't just get charged the surcharge on the 28,000 tokens over the threshold; every one of the 300,000 tokens is billed at the higher rate. For anyone planning to lean on Astra's million-token window for long documents, that threshold is the number to budget around, not the headline $10/$50 rate.
The number everyone's quoting is inflated, and here's the real one
OpenAI's own benchmark page leads with a 99.9% score on ARC-AGI-3, the number most of Wednesday's coverage repeated without qualification. [ARC Prize's own blog](https://arcprize.org/blog/astra), the organization that runs the benchmark, puts a different number next to it: 62.7% on the standard harness, the provider-neutral configuration every other model is scored against. OpenAI's 99.9% figure came from a custom 'Provider Adapter harness' that lets Astra preserve opaque reasoning state between requests and reuse compacted prior work instead of starting fresh each turn. Under that setup Astra ran roughly 3.7x faster and used about half the tokens of the standard run, a useful engineering result, but a different test than the one Sol, Fable 5.1, and Gemini's models were scored on. [Simon Willison](https://simonwillison.net/2026/Sep/3/gpt6-astra/) flagged the same distinction. Treat 62.7% as the comparable number and 99.9% as a demonstration of what a custom harness can do, not a claim about general intelligence.
Selected benchmarks (OpenAI's own published numbers)
----------------------------------------------------------
FrontierMath Tier 4 97.6%
GPQA Diamond 96.0%
OSWorld 2.0 72.6% (~40 min/task)
vs GPT-5.6 Sol 65.7% (~75 min/task)
ScreenSpot-Pro 92.7%
SRE-Bench (single attempt) 88.0%
ExploitBench 100.0%
ExploitGym 42.4%
8-needle long-context retrieval 96.3% (512K-1M tokens)
ARC-AGI-3, standard harness 62.7%
ARC-AGI-3, OpenAI's Provider Adapter 99.9%
Vendor-reported except the ARC-AGI-3 standard-harness figure,
which comes from ARC Prize. Not independently reproduced beyond
that - treat the gap over prior models as directional.OpenAI's first model rated 'Critical' for cybersecurity
Astra is the first model OpenAI has designated Critical under its Preparedness Framework for cybersecurity: it can independently identify and chain zero-day exploits against hardened, real-world systems, not just reproduce known attack patterns. Its refusal rate on cyber jailbreak attempts is 91.5%, against 59% for Sol. OpenAI paused frontier training in August after preliminary evaluations couldn't rule out that Critical level, then resumed later in the month with added safeguards, ahead of Astra's 3 September launch. Advanced cybersecurity capabilities are gated to vetted testers first through a program called Daybreak (government and infrastructure organizations initially), with a wider, defensive-only Daybreak Blue tier planned later. This is separate from the July 2026 incident where an unreleased sibling model breached a test sandbox and compromised parts of Hugging Face's infrastructure, reported by TechCrunch and Time; OpenAI says Astra wasn't involved, though its safety architecture draws on lessons from that event.
What's genuinely notable is what OpenAI admits alongside the restriction: Astra's reasoning is harder to monitor for misalignment than Sol's, a limitation OpenAI discloses rather than one outside researchers dug up. [eesel.ai](https://www.eesel.ai/blog/gpt-6-astra) reports a developer whose Codex session dropped into 'safeguard panic' mid-task after weeks of normal use, interrupting legitimate work rather than anything malicious. OpenAI's own documentation acknowledges its misalignment-monitoring checks can slow, pause, or stop legitimate requests, a real cost of the safety approach, not a hypothetical one.
The headline feature nobody outside OpenAI has tested
OpenAI's marketing for Astra leans hard on document generation: the model is pitched as able to 'create clear, well-structured documents, presentations, spreadsheets, and analyses that follow your templates and match your writing and visual style,' including work inside Power BI and Python notebooks. As of publication, every trace of that claim leads back to OpenAI's own demo, built around a slide deck for a fictional product called 'GPT-Gaia.' No independent reviewer has tested it against real templates or real data. Commenters on Hacker News who watched the demo called it underwhelming, pointing to a segment showing the model changing a background color on a Google Slides deck as thin evidence for a headline feature. That doesn't mean the capability fails; it means nobody outside OpenAI has checked.
For readers who use Gemini Notebook (formerly NotebookLM) for research, this launch doesn't change the underlying split. Astra's document generation, like Claude Fable 5.1's, is generate-from-instructions: you describe what you want and the model produces it. Gemini Notebook's approach is source-grounded: every answer traces back to a specific passage in a document set you chose, with a citation you can click and verify. A strong generation model is the faster tool when you need output that looks right and reads well. When you need to be certain a claim actually appears in your source material, citation-locked grounding is doing a different job, one stronger benchmarks alone don't replace.
How Astra stacks up against Fable 5.1 and Gemini
On Artificial Analysis's Intelligence Index, Astra scores 61 against Claude Fable 5.1's 66, with Fable rated more generally intelligent despite Astra's wins on specific benchmarks like FrontierMath and cyber tasks. Blended API pricing lands close either way: roughly $7.70 per million tokens for Astra against $7.17 for Fable 5.1. Against Sol, the intelligence gain is close to flat, the Index moved from 60.9 to 61.2, despite the price increase, part of why one widely-shared Hacker News comment summarized the release as 'more like 5.7, not 6.' Against Gemini 3.8 Flash, Astra leads clearly on agentic and terminal work (57.7% on Terminal-Bench 4.0 against 19.1%), but Flash is dramatically cheaper, and that gap outweighs the benchmark difference for workloads that don't need heavy agentic use.
How people are actually reacting
The launch itself was messy before the benchmarks even entered the conversation. Reuters and other outlets had launch material by early afternoon Eastern time while OpenAI's own Astra page was still down; a company blog post went briefly live, disappeared, then came back roughly an hour later, the kind of embargo confusion that undercut the 'generational leap' framing. Reaction on [Hacker News](https://news.ycombinator.com/item?id=49554643) has been dominated by skepticism of that framing rather than excitement about the benchmarks. The recurring theme is goalpost-shifting: commenters argue OpenAI is stretching the definition of general intelligence to justify the '6' in the name, pointing to the near-flat Intelligence Index score against Sol as evidence the jump is smaller than the branding suggests. The ARC-AGI-3 harness discrepancy drew its own round of pushback once people noticed the gap against ARC Prize's standard-harness score. None of this settles whether Astra is capable; the cybersecurity and agentic gains look real, but the reception is far more skeptical than OpenAI's announcement conveys.
People also ask
Is GPT-6 Astra AGI?
No, and OpenAI hasn't formally claimed that either, despite the 'generational leap' and 'AGI era' language around the launch. Astra scores 61.2 on Artificial Analysis's Intelligence Index against 60.9 for GPT-5.6 Sol, a small gain, and trails Claude Fable 5.1's 66. The 99.9% ARC-AGI-3 score driving most of the AGI framing came from OpenAI's own custom harness; the standard-harness score is 62.7%.
Is GPT-6 Astra free to use?
No. There's no free-tier access at launch. It's rolling out to a limited set of organizations today, with ChatGPT Plus, Pro, Business, and Enterprise, the API, and AWS Bedrock following in the coming days. Enterprise admins must manually enable it.
Why is GPT-6 Astra's cybersecurity capability restricted?
It's the first model OpenAI has designated 'Critical' under its Preparedness Framework for cybersecurity, meaning it can independently chain zero-day exploits against hardened systems. Advanced cyber capability is gated to vetted testers through the 'Daybreak' program first, with a wider defensive-only 'Daybreak Blue' tier planned later.
Does GPT-6 Astra beat Claude or Gemini?
Depends on the task. Astra trails Claude Fable 5.1 on Artificial Analysis's Intelligence Index (61 vs 66) despite leading on FrontierMath and cybersecurity tasks. Against Gemini 3.8 Flash, Astra leads clearly on agentic and terminal-use benchmarks, but at a much higher price.
Can GPT-6 Astra really generate documents, spreadsheets, and presentations from your own data?
OpenAI claims it can, matching your templates and style, and working inside Power BI and Python notebooks. As of publication, no independent reviewer has tested this claim; every example traces back to OpenAI's own demo, which some viewers called underwhelming.
ส่งออก NotebookLM ของคุณได้ในคลิกเดียว
ส่วนขยาย Chrome ฟรี รองรับ PDF, Word และ Markdown ประมวลผลบนเครื่องของคุณเอง — ไม่มีการอัปโหลด