U.S. artificial intelligence company Anthropic unveiled its new AI models, “Claude Fable 5.1” and “Claude Mythos 5.1.” Anthropic said the models improve performance while lowering usage fees, and also include experimental results from scientific research.
Anthropic said it is “world-class in coding and knowledge work, and has research capabilities that show in advance how AI can contribute to scientific progress.”
First, it is important to understand the relationship between the two models. Fable 5.1 and Mythos 5.1 are, under the hood, the same model. The difference lies in the strength of their safety guardrails.
Fable 5.1 is the general-purpose version available to anyone. Mythos 5.1, by contrast, is open only to verified institutions and individuals. It is a version with adjusted safeguards to support specialized work in fields where risks may arise, such as cybersecurity and the life sciences. It is like using the same engine but dividing it between a version for ordinary roads and one for controlled test tracks.
Anthropic said that, along with performance improvements, it also addressed three issues customers had raised: price, data retention, and safety guardrails.
The most notable change is price. Based on ordinary use, Fable 5.1 is about 25% cheaper than the previous Fable 5. For complex coding or multi-step automated tasks, savings can reach as much as 45%.
The reason is a cut in “cache read” costs. The term may be unfamiliar, but the principle is simple. AI processes user-supplied documents and conversations in chunks called tokens. When handling a long document repeatedly, rereading it from the beginning each time becomes expensive.
So the model stores what it has already processed and retrieves it later. Reading that stored content again is called a cache read. Anthropic lowered the cache read fee from $1 to $0.25 per 1 million tokens, a 75% reduction.
The more repetitive the work, the greater the savings. In automated tasks that repeatedly reference documents, cache reads account for most of the total cost. The remaining fees stay the same: $10 per 1 million input tokens and $50 per 1 million output tokens. Korean developers and startups using Claude will now be able to run the same tasks more cheaply.
Anthropic also released benchmark results across several tests. Benchmarks are performance tests in which models compete on set problems. The figures are Anthropic’s own measurements.
In “TerminalBench-Science,” which measures scientific research ability, Fable 5.1 scored 52.6%. Anthropic said this was more than double the previous model Fable 5’s 24.7%, and higher than the top-tier Opus 5 at 29.0% and competing models as well.
In “TerminalBench 4.0,” which measures coding tasks handled through commands, Fable 5.1 scored 55.8%, while Mythos 5.1 with safeguards relaxed scored 60.9%. In CursorBench, a test for the coding tool Cursor, it scored 73.4%. In “Humanity’s Last Exam,” which asks for knowledge across many fields, it scored 60.9% without tools and 65.0% with tools.
Anthropic emphasized a change in working style more than the raw performance figures. It said Fable 5.1 avoids shortcuts that make results merely look good and instead finds and fixes the root cause of a problem.
The company also cited a real-world example. In a test by investment firm Millennium, Fable 5.1 identified the cause of a rare error in the firm’s internal system. The problem appeared only once in a million cases and had gone unexplained for four to five years by any engineer or AI model. After examining software made by an outside vendor and comparing it with error logs, the model even pinpointed a flaw in that software.
Feedback from companies that had tried the model was also released. Financial firm Jane Street said it solved more coding problems than previous models and was easier to follow on long tasks. Development tools company Cognition said it was moving code review work previously assigned to Opus over to Fable 5.1 because of its lower per-task cost.
MongoDB said it built a complex prototype in three days and kept it running on its own for hours at a time. Fintech company Ramp said that in a 38-hour unattended run, the model found and corrected an error in an existing research result by itself.
Data retention is one of the most sensitive issues when businesses use AI, due to concerns that company secrets might remain on the AI provider’s servers. Anthropic introduced a new approach called “Enterprise Frontier Safeguards” (EFS).
The key point is that data is stored not on Anthropic’s servers, but in cloud infrastructure controlled by the customer. Even when a human needs to review the content, the customer does so directly rather than Anthropic. The company said this preserves the same level of confidentiality as “zero data retention,” while keeping monitoring functions that help prevent misuse.
The system was developed with more than 100 companies in finance, healthcare, manufacturing, and law, as well as with Amazon, Google, and Microsoft cloud services. It will be rolled out gradually starting this fall, and qualified customers can use Fable 5.1 on a zero-retention basis until the preparation is complete.
One of the biggest headaches in AI safety systems is false positives: blocking harmless questions by mistake because they are judged to be dangerous. For example, if medical questions or legitimate defensive work by security staff to protect their own systems are blocked, the tool loses usefulness.
Anthropic said it has significantly reduced false positives in Fable 5.1. In cybersecurity, the new safeguards produced about 60% fewer false positives than before. For Claude Code users, safety interventions per session fell by an average of about 60%.
This is largely because Fable 5.1 can now be used for defensive work that identifies software weaknesses. However, it still blocks efforts to build offensive tools that exploit those weaknesses. Ambiguous tasks such as penetration testing or attack-code creation are still routed to the higher-level Opus model.
The life sciences field was also adjusted. For elementary-level biology or medical questions, the frequency of safety-triggered interventions was reduced by 85%. However, questions tied to real research and development are still routed to the top-tier model instead. In addition, working with the U.S. government, Anthropic has created an access program that opens advanced biology capabilities in Mythos 5.1 to verified life sciences experts.
What Anthropic highlighted most strongly in this announcement was scientific research performance. It presented three examples.
The first was molecular design, which is closely tied to drug development. Many medicines work by binding to targets in the body. Designing binders that attach strongly to those targets is the first step in drug discovery.
Mythos 5.1 used a publicly available protein design tool to create binders, and the results were verified experimentally by two outside institutions. The results were striking. For three targets, the binding strength reached 10 times the best record from a related competition, and across 12 targets the success rate approached 50%. Typical success rates in this field are around 10% to 15%.
The second example was a map of Venus. Fable 5.1 trained a neural network to create a new high-resolution terrain map covering one-third of Venus, based on radar images taken more than 30 years ago by NASA’s Magellan probe.
The new map improved resolution from the previous 10 to 20 kilometers to 2 to 3 kilometers, and altitude information became up to 25% more accurate. Anthropic said it hopes the map will help future Venus missions by NASA and the European Space Agency, and released it under an open license.
The third was computational biology. In biology research, AI models running on graphics processing units are executed thousands of times. Speed directly determines research progress. Mythos 5.1 directly modified programs running on GPUs and increased the speed of seven publicly available deep-learning models by as much as 2.5 times.
The output stayed the same, but the speed improved. In genome-wide analysis tasks, GPU costs were reduced by 30% to 60%. Anthropic said work that would take several performance engineers weeks was completed in just days.
Before launch, Anthropic said it had conducted risk testing. The company and outside reviewers assessed whether the models could be abused to make chemical or biological weapons, and how capable they were in cyberattacks. Although Mythos 5.1’s biological capabilities are higher than before, Anthropic said they did not reach the next tier of its risk threshold, so the model is being deployed with the same safeguards as before.
Anthropic also responded to the European Union’s AI regulations. In July, together with 190 other organizations, it signed the content labeling code under the EU AI Act. Outputs from models released after August 2 will include an invisible marker, or watermark, to indicate the possibility that Claude helped write the text. The marker does not contain information about the user or the conversation, and cannot be verified without a dedicated detection tool.
Fable 5.1 is available starting the day of the announcement on all major platforms, including AWS, Google Cloud, and Microsoft Azure. Developers can access it through the Claude API under “claude-fable-5-1.”
Mythos 5.1 is limited to verified cybersecurity defenders and life sciences experts, and is currently restricted to certain institutions in the United States. Anthropic said it plans to expand access domestically and internationally in consultation with the U.S. government.
The launch, which lowered prices while improving performance and highlighting scientific research achievements, shows that competition among AI models is moving beyond simple performance comparisons toward cost, safety, and real-world usefulness. In particular, progress in scientific fields such as drug design and genome analysis remains a key measure of whether AI can become a true laboratory tool.