TL;DR
- Claude Opus 5.5 scores 54.4% on FrontierCode v1.1 (vendor-run, adaptive thinking at max effort) — above Fable 5.1’s 50.3% and last month’s Opus 5 at 48.0%.
- Pricing: $4 input / $20 output per MTok — 20% cheaper than Opus 5 on tokens, 60% cheaper on cache reads, netting ~40% lower cost per task when you account for fewer thinking tokens.
- Adaptive thinking is now always on and cannot be disabled — four API changes will silently break existing integrations.
Who should care: Teams running Claude Opus 5 in production coding loops, code-review pipelines, or knowledge-work automation where cost matters.
Verdict: Use it — migrate Opus 5 workloads now; the half-day migration pays back within the first week of inference spend.
Fable 5.1 Performance at Opus 5 Prices
On September 22, 2026, Anthropic released Claude Opus 5.5 — a model that lands between Opus 5 and Fable 5.1 in capability but costs 40% less than Opus 5 to run on typical agentic workloads. According to Anthropic’s release post, the model generates output more than 30% faster than Opus 5 and uses fewer tokens per task at the same effort level, which is where most of the cost reduction actually comes from.
The positioning is deliberate. Fable 5.1, released in September, remains the top of the Anthropic stack and costs $10/$50 per MTok. Opus 5.5 at $4/$20 offers comparable coding performance at less than half the price. For teams that defaulted to Fable 5.1 because Opus 5 felt inadequate, this is the model to re-evaluate first.
The model is available via the Claude API, AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure Foundry. Model ID: claude-opus-5-5. No weights are released; self-hosting is not an option.
Benchmark Breakdown
All scores below are from Anthropic’s release page — vendor-run evaluations using adaptive thinking at max effort with production safeguards enabled. Treat them as an upper bound on what you will see in production; Anthropic does not publish error bars.
| Benchmark | Opus 5.5 | Fable 5.1 | GPT-6 Astra | Opus 5 | Notes |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 57.9% | 52.3% | Vendor-run |
| FrontierCode v1.1 | 54.4% | 50.3% | 53.3% | 48.0% | Vendor-run |
| CursorBench 4.0 | 57.8% | 51.8% | — | 46.6% | Vendor-run |
| GDPval-AA v2.1 (Elo) | 1,846 | 1,735 | 1,542 | 1,708 | Vendor-run; knowledge work |
| AutomationBench | 40.0% | 31.4% | 41.4% | 26.9% | Astra leads marginally |
| Humanity’s Last Exam | 67.7% | 65.6% | 57.2% | 63.6% | With tools; vendor-run |
| OSWorld 2.1 (partial) | 81.8% | 80.7% | — | 74.0% | Computer use; vendor-run |
Source: Anthropic release post, September 22, 2026. All scores use adaptive thinking at max effort. Opus 5.5 wins or ties on every benchmark except AutomationBench and Terminal-Bench-Science (Astra leads there). No independent replication available at time of writing.
What a Code Audit Actually Costs Now
Token price alone does not tell you what you pay per task. Opus 5.5 reaches the same result in fewer thinking steps than Opus 5, which compounds the price cut. The math below uses Anthropic’s pricing page figures (Claude Platform docs) for a 200,000-line codebase security audit — a common CI use case.
| Item | Opus 5 | Opus 5.5 | Change |
|---|---|---|---|
| Input price ($/MTok) | $5.00 | $4.00 | −20% |
| Output price ($/MTok) | $25.00 | $20.00 | −20% |
| Cache read ($/MTok) | $0.50 | $0.20 | −60% |
| Estimated tokens per audit | 2.5× baseline | baseline | −30–40% |
| Estimated cost per audit | ~$40 | ~$12 | ~−70% |
Token prices from Anthropic Platform docs. Token efficiency delta from Anthropic’s release post. Cost estimates are illustrative — your actual spend depends on codebase structure, prompt length, and cache hit rate.
If your team runs daily CI audits, that delta adds up fast. At $40 per run, a team hitting the pipeline five days a week spends $200/week on Opus 5. The same workload on Opus 5.5 runs around $60. The migration pays for itself in days, not weeks.
For Swiss & EU teams
EU-based teams have two concrete wins with Opus 5.5. First, zero data retention (ZDR) is available — unlike Fable 5.1, which carries a mandatory 30-day retention requirement, Opus 5.5 is not a covered model and supports ZDR, relevant for teams operating under FADP or sector-specific data handling rules (source: innfactory.ai Claude availability page).
Second, AWS Bedrock now offers a cross-region inference profile (eu.anthropic.claude-opus-5-5) with source regions in Frankfurt, Ireland, and Stockholm. Note that there is no in-region endpoint in Europe on AWS — traffic processes cross-region. Google Cloud Vertex AI lists EU multi-region availability with ML processing in the EU from day one. Microsoft Foundry offers Global Standard and US Data Zone only; no EU data zone deployment as of September 22.
All Claude models released since August 2, 2026 include machine-readable watermarking per EU AI Act Article 50, covering transparency obligations for AI-generated content.
Four API Changes That Will Break Your Integration
According to CodingFleet’s migration analysis, four changes in Opus 5.5 silently break existing code — none return obvious errors until you check response payloads carefully.
1. Thinking cannot be disabled. Passing thinking: {"type":"disabled"} returns HTTP 400. Replace it with thinking: {"type":"adaptive"} and use the output_config.effort parameter for control.
2. Forced tool calls are gone. tool_choice: {"type":"any"} is rejected. Switch to tool_choice: "auto" with strict: true.
3. Thinking blocks are model-specific. Opus 5.5 rejects thinking blocks generated by Fable 5.1 or Mythos 5. Add the prefix_mismatch_behavior: "drop_block" beta header to multi-model pipelines.
4. Computer-use toolset changed. computer_20251124 is rejected on the Claude API and Google Cloud. Declare {"type":"computer_toolset_20260801"} instead.
There is also a silent behavioral change: text between tool calls now arrives in thinking blocks with display: "omitted" by default. Progress streams go silent without an error. Set thinking.display explicitly if you rely on streamed intermediate output.
Anthropic estimates the migration takes roughly half a day for a typical integration (source). The checklist: swap the model ID, delete thinking blocks from conversation history, replace forced tool calls, update the computer-use tool declaration, and re-run your effort sweep with the new baseline.
Routing Guide: When to Use Which Model
Based on the benchmark data above, here is a sourced routing recommendation:
- Use Opus 5.5 as the default for agentic coding, code review, codebase audits, knowledge work, and unattended production agents. It beats Opus 5 on every relevant benchmark and costs significantly less.
- Escalate to GPT-6 Astra only for scientific computing tasks where Terminal-Bench-Science matters, or for authorized offensive security work — Astra leads on both.
- Keep Fable 5.1 only for the hardest multi-step reasoning tasks where you have benchmarked it against Opus 5.5 on your own workload and confirmed the gap. At $10/$50 per MTok, the cost premium is real.
For the broader Anthropic model stack context, see our earlier analysis of Claude Fable 5.1’s September release and our guide on evaluating a new frontier model before switching.
Verdict: Use it. Opus 5.5 is the best cost-per-task model in Anthropic’s stack for production agentic coding. The migration is a half-day project that pays back in the first week of inference spend. Do the API changes, run the effort sweep, and retire Opus 5 from your production loops.
Further Reading
- Anthropic — Introducing Claude Opus 5.5 — the primary source for all benchmark scores and API change documentation.
- Claude Platform Docs — Opus 5.5 Overview — pricing table, context window, supported features, and model IDs by platform.
- CodingFleet — Claude Opus 5.5 Review: Breaking Changes — the most thorough migration checklist and a useful breakdown of where the model falls short.
Your turn: If you’re running Opus 5 or Fable 5.1 in a production coding pipeline today, have you started testing Opus 5.5 — and does the cache read price drop actually show up in your bills? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.
How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

