OpenAI released GPT-6 Astra on September 3, 2026 as a staged GA, API id gpt-6-astra, with a self-reported launch table that puts it at the top of the abstract-reasoning and cybersecurity rows and a private, third-party index that puts it behind Anthropic’s Claude Fable 5.1. The sticker is the same frontier class as recent OpenAI lists — $10 per 1M input, $50 per 1M output — but the rollout, the reasoning-effort defaults, the context-window math, and the cyber gating all changed enough that “GPT-6 dropped, you should switch” is the wrong mental model. This is a more carefully staged release than the headline suggests, and the staged bits matter more than the headline.
I read the OpenAI announcement, the OpenAI path-to-astra safety post from September 1, the Vellum benchmark explainer, the Artificial Analysis launch write-up, and the LLM Stats catalog entry. Four sources, four slightly different framings, one set of numbers. What follows is what I think is actually going on under the launch table, and where the lines on the table don’t match what shows up in your Codex bill.
The shape of the launch
Catalog id gpt-6-astra, hosted first-party by OpenAI, modalities text and image in, text out. Context window 1,050,000 tokens, max output 128,000, knowledge cutoff April 30, 2026. The reasoning-effort parameter now has five levels — low, medium, high, xhigh, max — where xhigh is a new ceiling rung added for this release. API default is still low. OpenAI’s marketing copy calls Astra “Highest reasoning,” but the API does not default to max; you set reasoning.effort explicitly or you get the cheap rail. That alone moves a lot of pre-launch speculation: the model card sells capability, the API sells throughput.
Standard list pricing is $10 input / $50 output per 1M tokens, with cached input at $1 and cache writes at $12.50. The cache-write line is new since GPT-5.6 — it’s a 1.25× surcharge on uncached input, not a discount, which means anyone building a prompt-cache pipeline needs to recompute break-even. The long-context multiplier kicks in at 272K input tokens: prompts above that threshold get billed at 2× input and cache, 1.5× output for the full request, not just the overage. So a 500K-token prompt is not “1M tokens at $10,” it’s “1M tokens at $20” before output enters the picture.
Rollout is staged the way it has been for the last few frontier OpenAI releases. Trusted Access Program enterprises get it today; Plus, Pro, Business, Enterprise, and the broader API are “in the coming days,” which OpenAI’s prior launches have usually meant 5–10 business days. The slower part is cyber. Path to Astra, published September 1, designates Astra as the first model to hit the Critical cybersecurity capability threshold under OpenAI’s Preparedness Framework. Advanced cyber access is gated behind Trusted Access and Daybreak Blue. If your agent stack relies on Astra for anything offensive-adjacent — exploit chaining, vulnerability discovery at scale — you are waiting on access approval, not on the API being live.
The headline table and what it isn’t
OpenAI published the launch table themselves and called Astra “the world’s most intelligent and aligned model.” Greg Brockman told reporters it is “not unreasonable to feel that we are now in the AGI era.” The table backs most of that swagger on OpenAI’s own chosen rows:
- FrontierMath Tier 4 (v2): 97.6% — Astra vs Fable 5.1 at 87.8% and Fable 5 at 87.8%, with Opus 5 at 73.2%. Epoch AI runs FrontierMath and notes OpenAI funded its development and has exclusive access to part of it, which is worth remembering when you read the saturation row.
- ARC-AGI-3: 99.9% on Astra’s launch table, against 30.2% on Opus 5 and 7.8% on GPT-5.6 Sol. Fable 5.1, Fable 5, and Gemini 3.8 Flash have no published score on this row. The New Stack’s writeup cites 98.6% for the same benchmark, likely a different harness configuration. Either way, this row is near-ceiling for Astra and the others are not playing in the same room.
- GPQA Diamond: 96.0%, the highest published score.
- Terminal-Bench-Science 0.1: 64.6% vs Fable 5.1’s 52.6%. The biggest single-table gap on the science side.
- OSWorld 2.0: 72.6% in roughly 40 minutes per task versus 65.7% (with a 70.2% unattributed row) for GPT-5.6 Sol at 75 minutes. 47% time reduction per task. If you price agent work by the hour, this is the row that changes the bill.
- ScreenSpot-Pro: 92.7% vs 76.9% for Sol. Dense-UI pixel clicking.
- Agents’ Last Exam: 59.3% vs 53.6% for Sol.
- AutomationBench: 41.4% vs Fable 5.1 at 31.4% and Sol at 18.1%. The biggest professional-work gap on the table.
- BenchCAD (Python): 95.9% vs 84.3% for Fable 5.1 and 83.3% for Sol, with OpenAI flagging that the Claude runs used modified evaluation settings.
- BrowseComp: 91.5% vs 90.4% for Sol. Closer than the headline.
- ExploitBench: 100%. Cyber exploit development from known vulnerabilities — the row that earned the Critical threshold.
The asterisks matter.
Humanity’s Last Exam with tools: 57.2% on Astra, Fable 5.1 at 65.0%, Fable 5 at 63.8%, Opus 5 at 63.6%. Astra loses this row. It is the only academic row in the table where it loses. Announcement prose does not mention it. On a benchmark named after the end of testing, the new flagship trails every Claude in the table.
FrontierCode 1.1 Extended: 64.5% on Astra, Fable 5 at 64.9%. On Main, Fable 5 (53.5%) and Opus 5 (53.4%) both edge Astra’s 53.3%. OpenAI’s table calls Astra “the best model for software engineering to date” and the table doesn’t support that on this row.
DeepSWE v1.1: 74.1% on Astra, 72.7% for Sol in OpenAI’s table. Meta’s Muse Spark 1.3, released the same week, hit 75.4% at its max reasoning setting. The public DeepSWE leaderboard had Gemini 3.8 Flash and Opus 5 around 74% with overlapping uncertainty ranges. The New Stack also flags that OpenAI’s chart uses a 67.4% Fable 5.1 result, which widens the visual lead over the broader set of runs.
The OpenAI launch table is a self-reported vendor table. It is the strongest possible framing for Astra’s choices and it is still losing two rows that are central to the “best in the world” claim.
The independent index: where the headline stops being true
Artificial Analysis publishes two flagship indices — the Coding Agent Index and the Intelligence Index v4.1.1 — that aggregate third-party evals. The launch write-up from September 3 is the clearest external read of what Astra actually does in an agent harness versus what the launch table shows.
Coding Agent Index (Codex harness): Astra scores 67, “approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code.” Fable 5.1 in Claude Code leads the Index at 70. So: tied with the field, behind Fable 5.1.
Token efficiency in Codex: Astra uses about one-third the tokens of GPT-5.6 Sol (max) in the Codex harness, and one-fifth the tokens of Claude Opus 5 (xhigh). This is the result that justifies the price increase. At max effort, Astra costs about the same per task as GPT-5.6 Sol (max) while scoring 2 points higher, and is less than half the cost per task of Claude Fable 5 for the same score.
Intelligence Index: Astra scores 61, the same as GPT-5.6 Sol. That is 5 points behind Claude Fable 5.1 (max with fallback) at 66, behind Opus 5 (63), behind Fable 5 (62), and behind Meta’s Muse Spark 1.3 (max). Various effort levels of Astra occupy the Pareto frontier of token efficiency, but on raw intelligence points, it sits next to its predecessor and behind Claude.
Hallucination: AA-Omniscience hallucination rate dropped from 92% to 51% at max effort, with a 4-point accuracy gain at the same time. This is the strongest single improvement on the Intelligence Index side. The accuracy-and-hallucination moving in opposite directions is the rarer shape — most “less hallucinated” results trade against accuracy, not alongside it.
AA-Briefcase (long-horizon knowledge work): Astra gains ~80 Elo, with improvements in both rubric scores and Analytical Quality Elo. But Presentation Quality Elo regresses, with GPT-5.6 Sol (max) still leading all models. Multi-week project work is better; the presentation of multi-week project work is worse.
GDPval-AA v2: Astra drops ~80 Elo, an evaluation Artificial Analysis adapted from OpenAI’s own dataset measuring economically valuable tasks across 44 occupations. So the model is better at long-horizon knowledge work and worse at the tasks OpenAI itself used to define “economically valuable.”
Small regressions: 2–3 point drops across τ³-Banking (customer support), SciCode (Python problems in scientific domain), and AA-LCR (long-context reasoning over large documents). The Intelligence Index is up because the gains outweigh the regressions, but the regressions are real and they are not in the marketing.
What the cyber threshold actually means
The Critical cybersecurity designation is the load-bearing piece of the rollout. ExploitBench 100%, ExploitGym 42.4%, SEC-Bench Pro 85.4%. Path to Astra describes Astra as “a significant increase in cybersecurity capabilities compared to GPT-5.6 Sol: it is both significantly more token efficient and more capable.” Daybreak Blue is the program that routes advanced cyber access for defenders. If you are not in the defender program, you don’t get the cyber-eval-trained behavior in the wild.
The cyber gating is the reason for Trusted Access first, Plus/Pro/API second. The previous OpenAI rollout where capability and access diverged was Mythos 5 / Fable 5, where Mythos was the unrestricted model and Fable shipped with biology and cyber guardrails. Astra doesn’t bifurcate that way — it’s one model — but the gate is at the access layer, not the model layer. Anyone who needed Mythos-grade capability for legitimate cyber defense work is going to be on Trusted Access. Everyone else is going to be running Astra through a wrapper that strips the cyber surface and bills them $10/$50.
The two changes that matter for agent builders
The first is the Codex context-preservation behavior. OpenAI now describes Astra as keeping notes across context windows in Codex, where earlier context windows stay searchable instead of getting compressed into a single summary when the window fills. If you’ve ever lost the detail of why the first fix failed in a long debugging session, this is the change you’ve been waiting for. Long-horizon work was the weak row for the GPT-5.6 series; the AA-Briefcase 80-point Elo gain is the external confirmation that this design choice actually moved something measurable.
The second is the OpenAI Responses API tool surface, which is more interesting than the launch table suggests. Streaming, function calling, structured outputs are all supported; Chat Completions and Responses are supported; Batch is supported; fine-tuning is not. Responses API tools include web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Input is text and image; output is text only. The tool list is the most concrete signal of what OpenAI expects Astra to do in production — and the inclusion of MCP as a first-class tool, not a connection string you bolt on later, is the quietest but most consequential line in the announcement for anyone who has been waiting on OpenAI to commit to the protocol stack MCP has been building.
The lines that didn’t move
Two specific lines that didn’t move, and the absence is interesting.
Open-weight counterpart. None. GPT-6 Astra is hosted first-party only. The closest open-weight peer on the launch week’s release schedule is Meta’s Muse Spark 1.3, which beat Astra on DeepSWE v1.1 (75.4% to 74.1%) the same week. DeepSeek V4 Pro and Kimi K2.6 were already the strongest picks per the Onyx open-source roundup, and neither refreshed for this release cycle. The open-weight frontier is not catching GPT-6-class this week; it’s also not losing to it cleanly.
Pricing restructure. The $10/$50 sticker is the same class as GPT-5.6 Sol — both pre- and post-launch. The cache math changed (cache writes at $12.50 instead of free), the long-context multiplier is new at 272K, the Fast-mode multiplier is 2×, the Batch and Flex rates are 50%. None of this is a price cut. Anthropic’s Fable 5.1 cut cache reads by 75% on its own release; Astra did the opposite. If you were betting on a price war triggered by the GPT-6 release, this isn’t it. The price war is on the Intelligence vs Cost-per-Task Pareto frontier, not on the sticker.
What I actually think is going on
The launch is two releases in one package. The marketing release is the OpenAI launch table: Astra tops abstract reasoning, math, cybersecurity, and computer use on rows where OpenAI chose the harness and the comparison set. The independent release is Artificial Analysis’s read: Astra is tied with the field on agentic coding, behind Fable 5.1 on raw intelligence points, 70% more token-efficient than its predecessor in Codex, 75% more expensive per task at max effort because of the 2.5× sticker increase, and hallucinating at roughly half the rate of GPT-5.6 Sol.
The right read is that both are true. OpenAI’s launch table is the strongest possible framing of a model that is genuinely better at the rows OpenAI cared about. Artificial Analysis’s index is the third-party composite that controls for harness effects and shows where the gain doesn’t generalize. The Critical cyber designation is the reason for the staged rollout — and the staged rollout is the reason the marketing claim and the access claim don’t quite line up. You can have Astra in your Codex today if you are in Trusted Access, and you can have Astra in your billing tomorrow if you are Plus or Pro, and the cyber-eval-trained behavior is gated behind Daybreak regardless of which tier you’re on.
The Questions I Don’t Have An Answer To
The biggest open question for me is whether the token-efficiency gain compounds in production the way it shows up in the Codex harness. AA-Briefcase going up 80 Elo and GDPval-AA v2 going down 80 Elo at the same launch is a strange profile — it’s the same model with two large regressions and two large gains on adjacent categories. If you run multi-week project work, Astra is measurably better. If you run economically-valuable occupation-spanning tasks, Astra is measurably worse. Picking which one your workload looks more like is a judgment call I can’t make from the launch table.
The second question is the cyber gating in practice. Trusted Access is a program, not an API surface; what it means for a developer building a defensive product is unclear until the Daybreak terms are public. Mythos 5 / Fable 5 had this same ambiguity in June, and the resolution then was: Mythos is unrestricted, Fable has guardrails, both run on the same API with different pricing. Astra is one model with one price and a capability gate at the access layer. That structure is harder to reason about and easier to ship a product against by accident.
The third question is the reasoning.effort default. OpenAI’s marketing says “Highest reasoning.” The API says low. Until someone publishes end-to-end cost numbers at max for representative agent workloads, the gap between “most intelligent model in the world” and “most intelligent model in the world running on the cheap rail” is going to be load-bearing for everyone who doesn’t set the parameter explicitly.
The release is real and the gains are real on the rows OpenAI chose to publish. The rollout is staged and the access is gated and the index doesn’t agree with the table. The next interesting data point isn’t another benchmark; it’s the Daybreak terms and what reasoning.effort=max actually costs on a real Codex workload.
References and where to dig further
- OpenAI, GPT-6 Astra: A new generation of intelligence, September 3, 2026. The launch table and the tool surface.
- OpenAI, Path to Astra: critical capabilities and frontier safeguards, September 1, 2026. The Critical cyber designation and the Daybreak program framing.
- Vellum, GPT-6 Astra Benchmarks Explained, Nicolas Zeeb, September 3, 2026. The clearest side-by-side of OpenAI’s table rows against external context.
- Artificial Analysis, Benchmarking GPT-6 Astra, September 3, 2026. The Coding Agent Index tie, the Intelligence Index 61-vs-66 gap, the AA-Omniscience 92→51 hallucination drop, the AA-Briefcase +80 / GDPval-AA −80 split, and the per-task token efficiency numbers.
- LLM Stats, GPT-6 Astra: Flagship Specs, Empty Scorecard, Sebastian Crossa, September 3, 2026. The catalog entry, the
reasoning.effortladder, the 272K multiplier, the cache-write surcharge, and the explicit “API default is low” call-out.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.