Anthropic put Claude Opus 5.5 in the field on September 22, 2026, and OpenAI put GPT-6 Sol and Luna in the field the same day. That timing isn’t a coincidence — the two launches were within hours of each other — but the play each lab made is structurally different. OpenAI went to the bottom of the price stack (Luna at $0.10/$0.50, Sol at $2/$10 with a 90% cache discount). Anthropic went sideways: same intelligence tier as Claude Fable 5.1 on most tasks, 40% cheaper to run than Opus 5, $4 input / $20 output / $0.20 cache reads. Opus 5.5 is the first model in the new Claude 5.5 family; Sonnet 5.5 and Haiku 5.5 follow in the coming weeks with the same efficiency-first posture.
The headline pitch is blunt: nearly the same intelligence as Fable 5.1 at a fraction of the running cost. On Anthropic’s Terminal-Bench 4.0 numbers (extra-high effort, standard error ±2.6 pts), Opus 5.5 scores 66.4%, comfortably ahead of GPT-6 Astra at 57.9% (high effort), ahead of Fable 5.1 at 55.8%, and ahead of Opus 5 at 52.3%. On Artificial Analysis’ Intelligence Index (max effort), Opus 5.5 took the top spot at 58, leading on six of the ten core evaluations, including Humanity’s Last Exam at 61.4% and SciCode at 66.9%, and reaching an Elo of 1822 on the AA-Briefcase knowledge work eval. Across four of its five effort settings, Opus 5.5 sits directly on the intelligence-versus-cost Pareto frontier.
What I find more interesting than the leaderboard numbers is the pricing shape. Cache reads dropped from $0.50 to $0.20 per million tokens — a 60% cut — while base input/output dropped 20% ($5→$4 and $25→$20). Anthropic’s framing here is honest about workload shape: cache reads are the majority of agentic and coding costs because the same context gets re-read on every turn. Cutting cache reads harder than base tokens is a direct bet that the workload they most want to win is the long-running agent session, not the one-shot chat query. For teams running Claude Code or any agent harness with substantial system prompt + tool definitions, that cache math matters more than the 20% headline.
What the same-day collision actually set up
OpenAI’s pricing on Sol and Luna — $2/$10 and $0.10/$0.50 respectively, with Sol’s cache reads at $0.20 — puts Luna at the bottom of any reasonable comparison. Sol’s per-token price is half of Opus 5.5’s; Luna’s is forty times cheaper on input. But there’s a wrinkle: OpenAI’s promotional cache-read price is 50% off for three months, which means the current Sol rate ($0.20) is not the steady-state rate. Anyone pricing a year-long production workload should be planning around what happens in January 2027 when that promo lapses, not what they’re paying today.
Opus 5.5’s pricing is the steady-state rate. There’s no promotional period to discount, no sunset clause. Anthropic isn’t trying to win a race to the bottom — they’re trying to close the gap between “what Fable 5.1 costs” and “what most customers will actually pay” while still charging a premium tied to reasoning quality. If you’re migrating from Opus 5 to Opus 5.5 you save 40% on typical workloads; if you’re migrating from Fable 5.1 to Opus 5.5, you save substantially more on token price but trade off the top-end reasoning ceiling. The decision becomes: which fraction of your traffic is the Fable 5.1 ceiling actually worth?
The technical spec sheet
Configuration details Anthropic has published: 1 million token context window, 128,000 token max output, always-on adaptive thinking, new beta support for inline tool calls mid-conversation. The model runs under the identifier claude-opus-5-5 and is available now on Anthropic’s own platform, AWS Bedrock, Google Vertex AI, and Microsoft Azure — all four at launch, not staggered.
That multi-cloud simultaneous launch is the part that most enterprise teams will care about. Recent frontier launches have typically gated access by tier or region (GPT-6 Astra’s release last week was phased; Claude Mythos 5.1 stayed Anthropic-direct for weeks). Anthropic skipped that step. If you have committed cloud spend on AWS or Azure already, you can route Opus 5.5 traffic through existing billing without standing up a separate vendor relationship. That removes a class of procurement friction that hasn’t been much talked about in the model-release coverage but shows up in every enterprise RFP.
Fast mode is also available — up to 2.5× the speed, priced at $8 input / $40 output per million tokens. For latency-sensitive paths in an agent harness (the foreground generation that the user is waiting on), that’s the option you’d reach for. For the background batch path (long-running tasks, overnight migrations), you’d stay on standard pricing.
Tester claims worth weighing
Anthropic shipped a long block of tester quotes. Some are the usual marketing copy; a few are concrete enough to use.
- One tester reports completing a 680,000-line code migration in less than a day — work that “would have taken an engineering team weeks.”
- On web-app load-time optimization across every page: Opus 5.5 succeeded 39 of 40 times; Opus 5 made smaller improvements that also altered the app’s behavior.
- A complex coding task that previously took 38 prompts over four days came in at 11 prompts over three hours, with “more production-ready outputs and less rework.”
- A tester handed Opus 5.5 a large engineering task across six repositories and let it run overnight, unattended. It stayed on task for over 18 hours defining how services communicate and worked out how each one should apply that. Compared with Opus 5, “it hit milestones faster and required minimal reworking.”
- Across public command-line benchmarks, Opus 5.5 solved more tasks than Opus 5 while making about 40% fewer calls and using half the tokens.
The 18-hour unattended run is the one I’d flag. Long-horizon agent tasks — the kind that matter for actual production deployment, not benchmark theater — have historically been where frontier models fall over. If Opus 5.5 genuinely stays on task for 18 hours with no rework, that’s the capability that matters. If it’s marketing copy, we’ll find out when real workloads start running on it.
Safeguards: the most consequential part of the release
This is the section most consumer-facing coverage will skip. It shouldn’t.
Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, so Anthropic is deploying it with safeguards similar to Fable 5.1’s. Most cybersecurity tasks will be re-routed to Opus 4.8. Biology work goes through the existing Life Sciences Verification Program. A Cyber Verification Program — three tiers for increasingly permissive trusted access, including access to Claude Mythos models — is “coming in the coming weeks.”
The distillation safeguard is the more interesting one. Opus 5.5 is launching with preserved thinking: the anti-distillation protection Anthropic introduced with Fable 5.1. It stops API users from editing Claude’s prior context in an attempt to extract reasoning. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. The framing in Anthropic’s September 2026 threat intelligence report is that distillation attacks — where bad actors use thousands of fake accounts to extract a model’s capabilities at industrial scale — create safety and national security risks, and that preserved thinking is the technical answer.
What this means in practice: if you create a new Anthropic API account after August 31, you cannot edit prior context on Fable 5.1 or Opus 5.5. The conversation history is locked to the model’s view. This breaks a class of prompt-engineering patterns that depend on context rewriting (some agent harnesses do this; some eval harnesses do this). It’s a deployer-side trade-off: better protection against model exfiltration, less flexibility on the prompt-construction side.
What this looks like for someone deploying today
If you’re already on Opus 5, the migration is straightforward and the cost reduction is real. Cache-read pricing is the lever that matters most for agentic workloads — if your system prompt + tool definitions are large and re-read on every turn, you’re seeing the 60% cut directly. Output tokens dropping 20% matters less for coding workloads where the bulk of the cost is context re-reads, but it matters more for generation-heavy paths like long-form writing or batch document processing.
If you’re on Fable 5.1, the calculus is whether the top-end reasoning ceiling is worth the price premium. Fable 5.1’s value proposition has been “the hardest problems are worth paying for.” Opus 5.5 narrows the gap — Anthropic’s own claim is “same level on most tasks” — but the hardest problems aren’t most tasks. If your workload includes a non-trivial fraction of genuinely hard problems, the Fable 5.1 → Opus 5.5 migration costs you something on the long tail. For everything else, you’re overpaying.
If you’re on GPT-6 Sol, the cross-vendor comparison is murkier. Sol is cheaper per token ($2/$10 vs $4/$20) and has its own cache-read pricing structure. But Sol doesn’t have the Mythos-tier safeguards or the Life Sciences/Cyber Verification Programs. If your workload touches regulated domains — health-adjacent software, financial trading desks with code-execution paths, anything that touches synthetic biology tooling — the Fable 5.1 / Opus 5.5 + Verification Program route is the only one that has the safety wrapper around the model. Sol is a faster, cheaper raw model with no equivalent program.
If you’re choosing between vendors for a new deployment, the multi-cloud availability matters more than the leaderboard position. AWS Bedrock + Anthropic direct, Google Vertex AI + Anthropic direct, Microsoft Azure + Anthropic direct — all four routes are live. For enterprise procurement, that’s the deciding factor more often than a 5-point Terminal-Bench gap.
What I haven’t seen yet
A few things I’d want before forming a firmer view:
- Independent verification of the 18-hour unattended claim. Anthropic’s tester quotes are not benchmarks; they’re anecdotes. The Terminal-Bench 4.0 number is verifiable on the public leaderboard, but the “680K-line code migration in less than a day” is a marketing claim. Someone with a real migration will have to run it before we know whether that’s representative.
- Public pricing on preserved-thinking workflows. The safeguard is real; the cost impact on context-rewriting-heavy workflows is not documented. If your agent harness depends on editing prior context (some compaction strategies do), the deployer-side cost on the new accounts may be higher than the headline 40% reduction.
- What Sonnet 5.5 and Haiku 5.5 actually look like. Anthropic says they’ll follow in the coming weeks with the same improvements. Sonnet is the volume workhorse for most Claude Code deployments; if Sonnet 5.5 lands at the same 40% reduction, the deployer-side cost story changes substantially.
- The Cyber Verification Program rollout specifics. “Three tiers for increasingly permissive trusted access, including access to Claude Mythos models” is a teaser, not a tier definition. The actual access criteria will matter for anyone doing security work on Anthropic models.
The same-day collision with GPT-6 Sol and Luna will probably be the angle most coverage takes — “frontier-model price war accelerates.” That’s true but it’s the boring read. The more interesting read is that Anthropic and OpenAI made structurally different plays on the same day: OpenAI went down-market with a three-tier stack (Sol, Luna, plus the existing higher tier), Anthropic went sideways into a new family positioned between their existing Opus and Fable. Both are responses to the same pressure — open-weight frontier releases like DeepSeek V4.1 keeping the floor moving — but they’re not converging on a single answer. The market is bigger than one pricing shape.
The Anthropic launch page is at anthropic.com/claude-opus-5-5. The benchmark breakdowns Vellum and Artificial Analysis published are the cleanest third-party reads — Vellum has the full Terminal-Bench 4.0 scoreboard with standard errors; Artificial Analysis has the Intelligence Index scatter across all five effort settings. The September 2026 threat intelligence report (linked from the launch page) is where the distillation defense is documented in detail.
Comments
Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.