
Claude Opus 5.5 Just Cut Frontier AI Prices 60% Without Cutting Corners

On September 22, Anthropic released Claude Opus 5.5, and the headline isn't a new benchmark record — it's the price tag. Opus 5.5 lists at $4 per million input tokens and $20 per million output tokens, roughly 60% cheaper than both Anthropic's own Fable 5.1 ($10 / $50) and OpenAI's three-week-old flagship, GPT-6 Astra ($10 / $50, rising to $20 / $75 once a request passes 272,000 input tokens). Cache reads, which quietly make up a large chunk of the bill for long-running agent sessions, dropped to $0.20 per million tokens, versus $0.25 for Fable 5.1 and roughly $1 for Astra.
That alone would be a decent story. What makes it a genuinely interesting one is that Opus 5.5 isn't a discount model — on the benchmarks Anthropic cares most about, it beats the models it undercuts.
The numbers that matter: long-horizon work, not trivia
Frontier labs have spent two years optimizing for one-shot benchmarks like SWE-bench Verified, where scores are now clustered in the low-to-mid 90s and barely distinguish one model from another. Anthropic's pitch for Opus 5.5 leans elsewhere: Terminal-Bench 4.0, which scores agents on multi-step command-line tasks rather than isolated code snippets. There, Opus 5.5 scored 66.4%, against 55.8% for Fable 5.1 and 52.3% for its own predecessor, Opus 5. The gap widens further on Terminal-Bench-Science, a harder subset, where Opus 5.5 hit 58.7% versus 29.0% for Fable 5.1 — roughly double.
Terminal-Bench 4.0: Opus 5.5 66.4% · Fable 5.1 55.8% · Opus 5 52.3%
Terminal-Bench-Science: Opus 5.5 58.7% · Fable 5.1 29.0%
SWE-bench Pro (Anthropic's agentic-coding benchmark): Fable 5 80.3% · Opus 5 79.2%
GPT-6 Astra, for comparison, scored roughly 73–74% on SWE-bench-style agentic coding evals — competitive, but not a clear leader; OpenAI has instead emphasized Astra's computer-use and browser-control scores.
Anthropic also published two anecdotes that say more than the leaderboards: one early tester used Opus 5.5 to audit and fix bugs across a 200,000-line codebase in under three hours, a task that reportedly took Opus 5 over 20 hours. Another ran a 680,000-line code migration in under a day. Whether or not those numbers generalize, they're the kind of claim that matters more to an engineering team than another point of SWE-bench — most real work isn't a single pull request, it's hours of an agent staying coherent across a sprawling repo.
Why this is really a story about market structure
Context matters here. GPT-6 Astra shipped September 3, three weeks before Opus 5.5, and OpenAI called it "the world's most intelligent and aligned model." Its coding scores landed close to Fable 5.1 and Opus 5 — good, but not a decisive win — while its strongest gains showed up in computer-use and desktop-agent tasks. Anthropic's response wasn't to chase Astra's territory; it was to ship a model that's cheaper to run at scale than either competitor while closing the gap on the one workload category — long-horizon agentic coding — where token costs actually add up.
That's the more durable trend underneath this week's headline. Agentic coding tools don't burn tokens on a single reply; they burn them across dozens or hundreds of tool calls, file reads, and self-corrections per task. At that volume, a 60% price cut on input tokens and a 20% cut on cache reads changes the economics of running an autonomous coding agent all day, not just the cost of a single chat message. Pricing, not raw capability, is becoming the lever labs pull once benchmark scores converge.
It also suggests something about Anthropic's model strategy: Opus 5.5 sits alongside Fable 5.1 rather than replacing it outright, and Anthropic explicitly frames it as matching Fable-level performance on most work at a lower operating cost — a mid-cycle release aimed squarely at production agent workloads rather than a full flagship refresh. A Fast mode variant, at $8 / $40 per million tokens, claims up to 2.5x faster throughput for teams that want speed over cost.
What it means if you're building with these models
If you're running long agentic sessions (code review, migrations, large refactors), Opus 5.5's cache-read pricing is the number to watch — it compounds over a multi-hour session far more than the headline input/output rate.
Terminal-Bench and SWE-bench Pro are better proxies for "can this model run unsupervised for hours" than SWE-bench Verified, which most frontier models now cluster around 90%+ on anyway.
GPT-6 Astra's context-length pricing cliff (rates roughly double past 272K input tokens) is worth checking against your actual prompt sizes before assuming its list price applies.
The benchmark race between frontier labs is narrowing at the top; the price and reliability race for agents running unsupervised for hours is just getting started.

