← BLOG
AI7 min

Anthropic Built a Model Too Good at Finding Bugs to Ship It

SolidAtoms Team
SEP 17, 2026
Anthropic Built a Model Too Good at Finding Bugs to Ship It

On September 1, 2026, Anthropic did something most AI labs never do: it shipped a flagship model to the public and, in the same announcement, confirmed that a more capable sibling model would not be released at all. Claude Fable 5.1 became the new generally-available flagship. Claude Mythos 5.1, a model in an entirely new tier Anthropic calls "Mythos-class" — sitting above the familiar Opus/Sonnet/Haiku lineup — stayed behind closed doors. The reason given wasn't a safety refusal in the usual sense. It was that the model is unusually good at one specific, dual-use thing: finding software vulnerabilities on its own.

That's a genuinely unusual move in an industry where the standard playbook is "ship it, add guardrails later." It's worth unpacking both halves of the announcement — what actually shipped, and what didn't.

What Fable 5.1 actually changes

Fable 5.1 is the first point release since Fable 5 launched back in June. On paper it's an incremental update — a 1,000,000-token context window, up to 128,000 output tokens, adaptive extended thinking, and a five-level "effort" setting for trading off speed against depth of reasoning. But the pricing change is the more interesting story for anyone actually running these models in production.

Input tokens are $10/M and output tokens $50/M, with cached-token reads cut 75%, from $1.00 down to $0.25/M. Anthropic says that combination nets out to roughly 25% cheaper than Fable 5 on typical workloads, and up to 45% cheaper on complex agentic or coding tasks, based on an internal analysis of four weeks of August usage. Anthropic also quietly made Sonnet 5's launch pricing ($2/$10 per million tokens) permanent rather than letting it rise to the previously planned $3/$15 — a small thing, but a signal that price competition among coding-focused models isn't easing up.

On benchmarks, Anthropic's own Terminal-Bench-Science 0.1 numbers put Fable 5.1 at 52.6%, more than double Fable 5's 24.7% and well ahead of Opus 5 (29.0%) and GPT-5.6 "Sol" (22.4%). Third-party leaderboards — not Anthropic's own reporting, and worth treating with a bit more caution — have put Fable 5.1 as high as 95.0% on SWE-bench Verified and 80.0% on SWE-bench Pro, ahead of GPT-5.5 (58.6%) and Gemini 3.1 Pro (54.2%) on the Pro variant.

The model that isn't for sale

Mythos 5.1 is where this gets more interesting than a routine model launch. Anthropic isn't making it publicly available at all. Access is restricted to organizations inside Project Glasswing, a cybersecurity initiative Anthropic launched in April 2026 with roughly a dozen partner organizations, reportedly including AWS, Apple, Google, Microsoft, CrowdStrike, and Palo Alto Networks.

According to reporting on the Glasswing program, a preview version of the Mythos model surfaced more than 10,000 high-severity vulnerabilities in widely used software — including bugs that had reportedly survived up to 27 years of human code review and automated testing before an AI found them. Whether or not that exact figure holds up under scrutiny, the underlying claim — that a model can autonomously and systematically hunt for exploitable bugs at a scale no human security team can match — is the actual news here.

A model that's extremely good at finding vulnerabilities is, by construction, extremely good at two very different jobs: defense and offense. Anthropic's decision reads less like a safety refusal and more like an acknowledgment that they haven't figured out how to separate those two uses yet.

That's a meaningfully different kind of gating than what we're used to seeing. Most "we won't release this" moments in AI have been about a model generating harmful content, or refusing certain requests. This is a lab looking at a working, presumably profitable capability and deciding the blast radius of misuse is too large to hand out an API key for — even to paying customers — while still being useful enough to build an entire partner program around.

Why this matters beyond Anthropic

For engineering teams, the practical takeaway isn't about Mythos — you can't buy access to it. It's about what its existence implies for the next 12-18 months of tooling. Autonomous vulnerability discovery at this scale is going to show up in commercial products one way or another, whether through Anthropic's own Glasswing partners, competing labs racing to match the capability, or open alternatives that don't bother with the gating at all. Security teams that have spent years triaging findings from static analyzers and fuzzers should expect the volume and quality of automated findings to jump, and dependency-heavy codebases — the kind most of us ship — are exactly the surface area this class of model is good at.

It's also a useful data point in the broader "which Claude model do I even use now" confusion that followed the Fable/Mythos naming change. Anthropic didn't just increment a version number this time — it created an entirely new tier above Opus, which is disorienting enough that multiple outlets published explainer posts just to map out the new hierarchy. If you're choosing a model for a new project this month, Fable 5.1 is the one you can actually use, and the pricing changes above are the numbers that will actually show up on your bill.

Sources