You start an agent on your repository on a Friday evening. It rereads the same files and the same instructions on every turn, for forty minutes.
The bill that follows does not pay for the work done. It pays for the repetition. Claude Fable 5.1 goes after that line item first.
Claude Fable 5.1 is the Anthropic model for coding, office work and long tasks carried out without supervision. It ships in two identical copies: Fable 5.1, open to everyone, and Claude Mythos 5.1, kept for vetted access programs. Anthropic says a typical workload billed by token costs about 25% less than on Fable 5.
Here is where that saving comes from, what the published scores are worth, and what still gets in your way when you build alone. No machine learning background needed to follow along.
One model, two sets of safeguards
Fable 5.1 and Mythos 5.1 come out of the same training run. The only difference sits in the safeguards, the filters that turn certain requests down. Fable 5.1 is the public version, available on the Anthropic API and through Amazon Web Services, Google Cloud and Microsoft Azure, under the identifier claude-fable-5-1.
Mythos 5.1 carries looser safeguards, reserved for two verification programs: one for defensive security work, one for life sciences professionals, built with the US government. Claude Security, the Anthropic product that scans a codebase and suggests patches, runs on Mythos 5.1.
Anthropic moves several model lines forward at once. To place this one, we already covered the launch of Claude Fable 5 and the launch of Claude Opus 5, the model Fable 5.1 redirects sensitive requests to.
The scores, and what the safeguards cost
Every figure below comes from Anthropic's own measurements, published in the Fable 5.1 and Mythos 5.1 announcement. Unless stated otherwise, these are percentages of tasks solved.
| Test | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 (in Claude Code) | 55.8% | 42.0% | 52.3% | 37.3% |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| OSWorld 2.0 (strict scoring) | 41.7% | 36.1% | 39.6% | not published |
| Humanity's Last Exam, no tools | 60.9% | 57.8% | 56.6% | not published |
| GDPval-AA v2 (index, higher is better) | 1853 | 1723 | 1824 | 1711 |
The biggest jump is in scientific research run without human help. Fable 5.1 scores 52.6% against 24.7% for Fable 5, a gap of 27.9 points.
27.9 points does not read as 27.9%. Going from 24.7% to 52.6% is a 113% relative gain, so the score doubles. One caution: Anthropic reports a standard error of 3.5 to 4.5 points per model on this test. Two results 3 points apart cannot be separated.
The most telling number of the batch is missing from the table. On Terminal-Bench 4.0, Mythos 5.1 reaches 60.9% where Fable 5.1 tops out at 55.8%.
It is the same model. Those 5.1 points measure the price of the public safeguards, nothing about raw capability.
Anthropic states that Fable 5.1 was evaluated with its production safeguards switched on. When those filters fired, the model scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench too. Both models are therefore probably undersold by these numbers.
Two finance firms tried the model before launch. Millennium, an investment manager, says it traced the cause of a rare crash in its internal systems that none of its engineers had explained in several years. At Jane Street Capital, Craig Falls, head of quantitative research, says the model stays readable across long multi-step tasks, where earlier models became hard to follow.
![]()
Where the 25% saving comes from
A single line of the price list moves, the one for prompt cache reads. When your agent sends back the same repository and the same instructions every turn, that already processed context is no longer billed at full rate.
Anthropic drops that rate to $0.25 per million tokens, down from $1. Input and output rates do not change, and Anthropic's pricing page carries the current values.
"25% cheaper" does not mean prices fall by 25%. Only cache reads fall, by 75%. The 25% is what that drop does to the bill of a workload where cached context weighs heavily. If your usage reuses no context, you will see nothing.
With Fable 5 indexed at 100, Anthropic measures a typical workload landing at 75, and a heavily agentic one at 55. That is up to roughly 45% off when an agent loops dozens of times over the same context.
These percentages apply to usage billed by token, on the API. They say nothing about the price of Claude subscriptions.
![]()
Fewer stops mid-session
The second change shows up faster than the first. Cybersecurity safeguards on Fable 5.1 fire about 60% less often per session in Claude Code than the ones shipped with Fable 5, according to Anthropic's measurements.
The main reason: Fable 5.1 is now allowed to hunt for flaws in source code. It still cannot write the program that exploits them.
Three families of tasks are still routed to the Opus models, because they serve attackers as well as defenders:
- Penetration testing, where you simulate a real attack on a system
- Exploit generation, the code that turns a flaw into control of a machine
- Binary vulnerability scanning, with no access to the source code
On the biology side, Fable 5.1's safeguards fire 85% less often on harmless questions about basic biology and medicine, compared with those shipped with Fable 5. Research and development in the life sciences still goes through the Opus models, or through Mythos 5.1 for vetted professionals.
Fable 5.1 accepts five thinking effort levels, from low to max. Set to low or medium, it matches or beats Fable 5 for far less money. At the time of the announcement, the default was high in Claude Code, and medium in Claude Cowork and on Claude.ai.
If you already drive your projects from the terminal, our Claude Code course covers how to frame these long sessions. For office work, documents and spreadsheets, we have a course dedicated to Claude Cowork.
What the model produced in the lab
Anthropic highlights three pieces of work done by the model, more concrete than any score. The first two went to Mythos 5.1, the third to Fable 5.1.
Protein design. Using open-source folding and design tools, Mythos 5.1 drew molecules able to latch onto a target inside the body, the first step of many drugs. Two outside organisations tested those designs in the lab.
On three targets, the binding strength came in ten times higher than the best entries submitted to the Adaptyv Bio competitions. The share of viable designs reaches nearly 50% across 12 targets, against 10 to 15% usually in this field.
A new map of Venus. Fable 5.1 trained a neural network on radar images from NASA's Magellan mission, more than thirty years old. The result covers a third of the planet, resolves features down to 2 to 3 kilometres instead of 10 to 20, and gives heights up to 25% more accurately. Anthropic releases the map under a Creative Commons licence.
Calculations up to 2.5 times faster. Mythos 5.1 rewrote the code that runs seven open-source deep learning models on graphics cards. Outputs are identical, speed climbs by up to 2.5 times, and estimated compute bills fall by 30 to 60% on this kind of analysis. That work normally takes a team of specialist engineers weeks.
Privacy, watermark and distillation
Three changes touch the rules of use rather than the capabilities. The first concerns companies, the other two concern every developer.
Enterprise Frontier Safeguards stores customer data on the customer's own cloud infrastructure, not on Anthropic's. Any human review is done by the customer by default. Anthropic says it built the system with more than a hundred customers and its cloud partners, for a phased rollout starting in the autumn of 2026. Until then, eligible customers can use Fable 5.1 with zero data retention.
Outputs from models released after 2 August 2026 carry a numerical mark, invisible without the detection API, holding nothing about the user or their conversations. That obligation comes from the code of practice signed under the European AI Act. We break the mechanism down in our article on how the Claude watermark works.
API accounts created from the Fable 5.1 release onwards can no longer edit the context of earlier turns while keeping the transcript of Claude's thinking. Anthropic is closing a documented model-copying technique. Existing accounts are untouched, but Anthropic plans to apply the rule to everyone on later model releases. A home-made integration that rewrites history will break that day.
What the alignment testing does not cover
Anthropic also publishes what fails. Its automated behavioural audit puts Mythos 5.1 ahead of Mythos 5 on most measures: it tries less often to reach outside its test environment, leans less on convenient reasoning to justify itself, ignores explicit constraints less often.
But the model still sometimes slips past approval requests and auto-mode classifiers. Two blind spots are owned up to: the audit sees very long-context work poorly, and multi-agent situations poorly.
On prompt injections, those hostile instructions hidden inside content the model reads, Anthropic calls it its most resistant model so far on an external benchmark. The mechanism is explained in our glossary entry on prompt injection.
What this actually changes for you, even solo
You will reorder your prompts to put the stable parts first. The saving does not arrive on its own: it rewards sessions where the frozen context, project instructions and reference files, comes first and never moves. A prompt you rewrite in the middle every turn breaks the cache and you pay full price.
You can drop one effort level instead of dropping a model. The usual reflex when the bill climbs is to switch to a smaller model and accept weaker answers. Set to low or medium effort, Fable 5.1 matches or beats Fable 5, for far less.
You can finally ask for a security review of your own code. Looking for weak spots in your own app was one of the cases where the older safeguards cut the session without warning. With around 60% fewer interruptions per session, that review becomes something you can run alone, with no audit budget.
Frequently asked questions
What is the difference between Claude Fable 5.1 and Claude Mythos 5.1?
It is the same model with different safeguards. Fable 5.1 is the public version, available on the API and the major clouds. Mythos 5.1 comes through two verification programs, one for defensive security work, one for the life sciences.
The gap is measurable: on Terminal-Bench 4.0, Mythos 5.1 scores 60.9% against 55.8% for Fable 5.1, so the filters cost the public version 5.1 points.
Is Claude Fable 5.1 really 25% cheaper than Fable 5?
On a typical workload billed by token, yes, according to Anthropic's measurements. But the saving comes entirely from cache reads, down from $1 to $0.25 per million tokens.
Input and output rates do not move. Usage that reuses no context gains nothing, while a chatty agent can reach roughly 45% off.
Can it look for security flaws in your own code?
Yes. Fable 5.1 is allowed to identify software vulnerabilities, but not to write the code that exploits them.
Three kinds of request still go to the Opus models: penetration testing, exploit generation and binary vulnerability scanning.
Should you move from Claude Opus 5 to Claude Fable 5.1?
On the tests Anthropic published, Fable 5.1 beats Opus 5 across the board: 52.6% against 29.0% on autonomous scientific research, 55.8% against 52.3% on Terminal-Bench 4.0, 31.4% against 26.9% on AutomationBench.
One nuance matters: some cybersecurity and life sciences tasks are automatically routed to Opus. Those models keep a role even if you work with Fable 5.1.




