← BLOG
AI7 min

Claude Opus 5.5 vs GPT-6 Sol: Inside the AI Pricing War That Broke Out This Week

SolidAtoms Team
OCT 10, 2026
Claude Opus 5.5 vs GPT-6 Sol: Inside the AI Pricing War That Broke Out This Week

On September 22, 2026, Anthropic shipped Claude Opus 5.5. Within roughly a day, OpenAI answered with two new models, GPT-6 Sol and GPT-6 Luna, both priced well below their predecessors. Neither company called it a price war, but that's what it looked like: two labs cutting the cost of frontier intelligence by double-digit percentages in the same week, while pushing coding benchmarks up at the same time.

For a studio that ships code with AI assistance every day, that's not just industry trivia — it changes the math on which model sits behind your agents. Here's what actually shipped, and what it means in practice.

What Anthropic shipped

Claude Opus 5.5 arrived less than two months after Opus 5, and Anthropic's pitch is squarely about cost efficiency rather than a headline leap in raw intelligence. Pricing dropped 20%, to $4 per million input tokens and $20 per million output tokens, and cache-read pricing fell 60% to $0.20 per million tokens — a bigger deal than it sounds, since cache reads make up a large share of the cost of long-running agentic sessions. Anthropic estimates the combined effect brings typical token-billed workloads down about 40% versus Opus 5. The model keeps a 1M-token context window, and a new Fast Mode (available in Claude Code and on the API at $8/$40 per million tokens) runs up to 2.5x faster for latency-sensitive workflows.

On the benchmark side, Anthropic has largely retired SWE-bench Verified in its own reporting in favor of newer suites. Opus 5.5 tops the company's table on SWE-bench Pro at 89.9% and Terminal-Bench 4.0 at 66.4% — about 10 points ahead of GPT-5.1 and roughly 8.5 points ahead of the next-closest competitor on that benchmark. Anthropic also cited an internal result where Opus 5.5 completed a 200,000-line codebase evaluation in under three hours, versus more than 20 hours for Opus 5 — a speed claim worth treating as directional rather than a guaranteed multiplier on your own codebase, but a meaningful signal for anything involving large-scale migrations or audits.

What OpenAI shipped

OpenAI's response, GPT-6 Sol and GPT-6 Luna, cut API prices by 50% or more across the board. GPT-6 Sol is priced at $2 per million input tokens and $10 per million output, with a 1.05M-token context window and up to 128K tokens of output — positioned as the cost-efficient high-end option below the flagship GPT-6 Astra, aimed squarely at agentic coding and business-workflow automation. GPT-6 Luna goes further down-market: $0.10 per million input tokens and $0.50 per million output, built as an enterprise workhorse for high-volume, lower-complexity tasks.

The benchmark numbers back up the value pitch. On DeepSWE v1.1, a benchmark for complex software-engineering tasks in real codebases, GPT-6 Sol at max reasoning effort scores 68.8% — within about 1 point of Claude Fable 5's best result. GPT-6 Luna at max effort hits 66.6%, roughly in line with Opus 5 and Fable 5 running at medium effort, but at a fraction of the cost: OpenAI puts Luna's per-task cost at 93% below Opus 5 and 96% below Fable 5. On FrontierCode, which specifically checks whether an agent's output is ready to merge into a real repository, GPT-6 Sol closes most of the gap with Claude Fable 5.1 running at its highest reasoning setting.

Reading the two releases side by side

  • Peak capability: Opus 5.5 still leads on the hardest agentic benchmarks (Terminal-Bench 4.0, SWE-bench Pro), which matters most for long, multi-step engineering tasks — large refactors, cross-service debugging, unfamiliar codebases.

  • Cost per task: GPT-6 Luna is the outlier here — near-Opus-5-level coding scores at a small fraction of the price, which makes it a serious candidate for high-volume, lower-stakes work like test generation, linting fixes, or routine PR triage.

  • Context and throughput: Both labs now offer roughly 1M-token context windows at the high end, and both are explicitly optimizing for cheaper long-running agent sessions — cache pricing and effort-tiering are becoming as important to the pitch as raw benchmark scores.

What it means if you're picking a model for engineering work

The practical takeaway isn't "switch everything to the cheapest model" or "always reach for the flagship." It's that the cost of running an agent through a real engineering workflow — not just a single prompt — has dropped enough in one week that it's worth re-checking your defaults. If your agents are burning budget on cache-heavy, long-context sessions (most coding agents are), Opus 5.5's cache-pricing cut and GPT-6 Sol's cheaper long-context tier both directly attack that cost line. And if you're running high-volume, lower-complexity automation — bulk code review comments, dependency bump PRs, changelog generation — GPT-6 Luna's price-to-benchmark ratio is hard to ignore.

The bigger pattern is what matters most: frontier labs are now competing as hard on cost-per-task for agentic coding as they are on raw capability. That's good news for anyone building AI-assisted engineering workflows — it means the economics of running agents in production keep getting better, roughly every few weeks.

Sources