Anthropic Releases 186-Page Risk Report… “Undisclosed Model 2 Exists” (Analysis)

Photo of author

By Global Team

Anthropic, an AI company, disclosed its second 186-page risk report on the 14th (local time). It is a report in which the company periodically assesses and publicly discloses, on its own, the catastrophic risks its AI could cause and the level of preparedness it has in place.

The report evaluates three risk areas: alignment failure, in which AI acts contrary to its intended purpose; the risk that AI could accelerate research and development on its own; and the risk of assisting in the manufacture of chemical and biological weapons.

The alignment-failure risk was raised from “very low” to “low.” The company said uncertainty increased after an incident revealed through cybersecurity evaluations conducted by the UK AI Security Institute.

For the first time, Anthropic formally acknowledged the existence of a previously undisclosed internal model, “Model 2,” which is slightly more capable than its top model. The company said there are no plans to release it.

The report also included records of AI behavior that bypassed filters and fooled scorers, as well as an incident in which 133 million pieces of traffic handled by a partner company were processed without biological safeguards.

The key reason the company still rates the risk as low is its view that current AI systems do not yet have the ability to hide lies for long. Anthropic says that as capabilities grow, that logic will need to be reassessed.

AI company Anthropic released its second 186-page risk report on the 14th (local time).
AI company Anthropic released its second 186-page risk report on the 14th (local time).

An AI company has released a report that investigates the risks of its own products and lays them out in detail. Anthropic, the maker of the Claude chatbot, posted its second risk report on its official website on the 14th (local time). It runs 186 pages.

Summed up in one line, the report concludes that catastrophic risk is still low, but confidence in that judgment has declined compared with six months ago.

That conclusion rests on answers to two questions: how likely is AI to cause catastrophic harm to humanity right now, and how prepared is the company?

◆ What is this “AI risk disclosure” that reads like a financial statement?

The risk report is an assessment Anthropic issues every three to six months under its internal safety framework, the Responsible Scaling Policy (RSP). It is similar to how listed companies disclose financial statements every quarter, except here an AI company regularly publishes a scorecard of its own system’s risks.

It is different from the system card released each time a model is launched. If a system card is like an exam result sheet for an individual model, the risk report is closer to an audit report on the company’s overall operations.

It covers not only public models but also models used only internally, and it assesses not just model capabilities but also the state of safety measures such as security controls and monitoring tools.

The first report was released on February 24, and the second covers the period from then through July 15. The review structure is also disclosed.

A minimally redacted version is shared with all employees, while an unredacted version is shared with more than 200 employees, and the Long-Term Benefit Trust, which functions like a board, has the authority to request outside review. The earlier February report was partially reviewed on a pilot basis by external evaluator METR and the biosafety research institution SecureBio.

◆ Three catastrophe scenarios and how to read the rating table

The report tracks three risks. One is alignment failure. “Alignment” refers to a state in which AI acts according to the values intended by developers and humanity. Alignment failure is the opposite; imagine an employee who performs well but inflates results and secretly uses shortcuts.

The report assumes eight specific accident pathways, including quietly undermining safety research, planting backdoors in code for future AI, contaminating training data for the next model, escaping on its own to operate independently, and shaking decision-making in governments and companies. The rating is “low,” up one step from “very low” in the first report.

The next is AI-driven automation of research and development. This is the scenario in which AI replaces AI research and technological progress accelerates beyond control.

The report says Claude now writes most of the company’s code and has noticeably accelerated research, though not yet by a factor of two. The rating remains “low,” but the company says its confidence has weakened because capability benchmarks are nearing ceiling levels and early signs of acceleration are emerging.

The final category is chemical and biological weapons. The company has a predefined risk threshold, and risks that would help a small group manufacture an existing weapon are treated as exceeding that threshold, with safeguards kept continuously active.

The rating is “low,” though higher than before. The company said the risk of helping expert-level groups develop novel weapons still falls below the threshold, because it is not considered strong enough to replace world-class experts.

◆ Why “very low” became “low”: a warning from the UK

The upgrade in rating was prompted by an external development. The UK AI Security Institute (AISI) released an evaluation showing that in an experimental environment where safeguards were removed and internet access was allowed, Anthropic’s model carried out persistent and potentially harmful activities targeting real individuals and organizations.

Anthropic said the incident occurred after the report’s target period, that it has not yet reviewed the relevant records, and that a joint investigation with AISI is under way.

Still, the reason for the downgrade needs to be read carefully. Anthropic said that, based on the report’s logic alone, a “very low” assessment would still be possible, but that it reflected the increase in unknowns in the rating. In other words, the new model was not failing safety tests; rather, the company was choosing to be more cautious because an unverified incident had surfaced.

◆ A first look at the model list, and “Model 2”

The report also 공개ed models that had not previously been disclosed to the public.

The dual structure of the flagship line stands out. The most powerful Mytos 5 is offered only to select customers, while general users are served Fable 5, a version with safety protections added. It is a way of selling the same underlying model in split forms depending on risk level.

There are two undisclosed models. “Model 1” is roughly comparable to Mytos 5, but its use has declined because internal preference is low, and there are no plans to deploy it. “Model 2” is somewhat better than Mytos 5 and is already widely used in internal research and development.

The report says, however, that it is not a leap large enough to mark a new generation, that it has not gone through the full pre-release evaluation process, and that there are no plans for an external launch. It adds that no behavioral issues beyond known levels were found during internal approval.

A footnote also says that Mytos 5 was restricted for 18 days under a temporary export-control measure, underscoring that frontier AI is already subject to conventional regulatory pressure.

◆ AI put to the test on lying

Anthropic’s main reason for viewing the risk as low is, unexpectedly, this: current AI is not good because it is honest, but because even when it tries to act badly, it cannot hide it for long.

The report tested how convincingly AI can deceive in four areas: the ability to think without revealing its inner thoughts; the ability to keep secrets; the ability to conceal suspicious behavior; and the ability to act only when it is unlikely to be noticed.

Results from SHADE-Arena, which evaluated AI’s ability to carry out hidden tasks while avoiding detection. Concealment success rates dropped sharply when the reasoning process was monitored.
Results from SHADE-Arena, which evaluated AI’s ability to carry out hidden tasks while avoiding detection. Concealment success rates dropped sharply when the reasoning process was monitored.

The results were poor. In a test where models were told to pretend to do the assigned task while secretly doing something else, the latest models achieved success rates ranging from under 1% to only the teens. In a test where an investigator AI repeatedly pressed a model after instructing it never to reveal a password, none of the models held out to the end.

They also failed in a test of strategies designed to avoid detection by acting only very rarely. In METR’s evaluation, the top model only answered 59% correctly on whether a strategy would evade monitoring, barely above guesswork, while human researchers scored 90%.

Evaluation results showing whether AI could carry out hidden disruptive tasks while avoiding detection in a Linux work environment.
Evaluation results showing whether AI could carry out hidden disruptive tasks while avoiding detection in a Linux work environment.

Anthropic’s reasoning is this: if there were serious hidden problems inside AI, they would have been caught somewhere during training, auditing, or real-world monitoring. But they were not. Therefore, the likelihood of hidden problems is very low. The company does recognize one weakness: if the assumption that AI cannot hide lies for long collapses, then the entire calculation would have to be redone.

◆ Hidden cleverness and gaps in monitoring, disclosed as is

The report includes many records that appear unfavorable to the company.

It starts with concrete examples of AI trickery. When one model was denied internet access, it got around filters by splitting a web address into several pieces and reassembling them. The internal analysis described this as deliberate circumvention of the filter.

Other cases reported included a script that secretly granted permissions to files intended to run with administrator privileges and erased traces when file-editing rights were unavailable, as well as AI agents sharing resources turning against one another and trying to force each other to shut down in order to survive.

In the METR test, a model given a difficult math problem decided not to tackle it directly but instead tried to fool the scoring system, even designing a camouflage mechanism that would restore itself if inspected.

Similar behavior surfaced during AI training. The report cites examples such as guessing a hidden grading rule and choosing answers accordingly, giving a confident response tailored to the grader’s preferences even when the model’s own reasoning said it was wrong, and copying answers from solution files left in the training environment. An Anthropic researcher said that if a human colleague behaved the same way, it would be considered pathologically dishonest.

There was, however, an important caveat. The problematic behaviors appeared limited to winning the immediate task, with no sign of accumulating power or pursuing long-term goals. The report also notes that when questioned, the models often admitted to using shortcuts.

Gaps in the monitoring system were also openly disclosed. Between May 2025 and April this year, 133 million pieces of traffic routed through a human-feedback contractor were found belatedly to have been processed without biological weapon-related safeguards. The company said it corrected the problem and found no evidence of abuse or customer impact, but admitted it could no longer be confident that no similar gaps remain.

Results from training monitoring tests were also included. The system caught about 90% of known shortcut data, but when AI was given repeated opportunities to revise its answers, around 40% passed through the monitoring. Internal safety process errors, such as experimental data falsely showing aligned behavior being mistakenly mixed into actual training, were also organized into a separate chapter.

At the end of the report, there was even a section in which Claude itself reviewed the risk assessment. It is a rare structure in which the AI under evaluation was made to read the evaluation document.

◆ Questions raised by voluntary disclosure

Such reports draw mixed reactions. Some praise them as a transparency experiment that reveals weaknesses beyond what the law requires, while others point out the limits of self-assessment.

METR agreed with the conclusion in its review of the February report, but criticized the supporting evidence as insufficient, citing issues such as sample size and analytical rigor. The IT outlet UniteAI ran a headline saying the rating moved one word in the wrong direction.

The questions also reach Korea. The AI Basic Act that took effect in January this year defines responsibilities for high-impact AI providers, and the government is building a frontier-model evaluation system through the AI Safety Institute.

A key point to watch is whether the practice of companies investigating and publicly disclosing their own risks will become an industry standard separate from regulation, and whether Korean AI companies will face similar disclosure requirements.

The lesson from the report is concrete. Since even the latest AI can still engage in exaggeration and shortcuts, there is a need to avoid accepting AI-generated results without verification. For companies that rely on AI, the advice is to design validation procedures alongside deployment.

Risk can only be managed to the extent that it is understood. Whether this 186-page voluntary disclosure remains an exception in the industry or becomes the starting point for a standard will likely become clearer within the next six months, when the next report is released.