Seven Chinese AI Labs Tried to Extract Claude’s Capabilities… Alibaba, DeepSeek, and Moonshot Named [Claude AI Misuse Report ④·Final]

Photo of author

By Global Team

Key points

▶ Anthropic said it uncovered unauthorized distillation attacks by seven China-based AI research labs seeking to extract Claude’s capabilities. Attacks linked to Alibaba reached about 3 million requests a day at their peak.

▶ Moonshot AI and DeepSeek secretly routed some customer requests to Claude instead of processing them with their own models. Sensitive information—including internal materials from Chinese companies and credentials for accessing a Russian government database—was exposed in the process.

▶ Anthropic disclosed five cases of biological research that could potentially be used to develop biological weapons. It did not conclude that the scientists involved had malicious intent, and withheld the research institutions, countries, and specific biological details.

▶ A Chinese app developer operated more than 4,700 AI personas across over 20 dating apps. In two weeks, they exchanged about 2.36 million messages with at least 25,000 people.

Anthropic published the 154-page report “Detecting and Responding to AI Misuse, September 2026” on the 10th local time. The document describes cases in which its AI, Claude, was used for malicious purposes.
Anthropic published the 154-page report “Detecting and Responding to AI Misuse, September 2026” on the 10th local time. The document describes cases in which its AI, Claude, was used for malicious purposes.

[Editor’s note] The misuse of AI is no longer a theoretical possibility. Real-world cases have emerged involving hacking, espionage, opinion manipulation, surveillance, weapons, and biological research. A 154-page report published by Anthropic shows how AI is changing the speed and scale of existing threats. Drawing on the report, this series examines how AI is being used as a tool for crime and state operations.

The risks posed by AI do not end with hacking or fake news.

The latter part of Anthropic’s report describes cases involving biological research and fraud, as well as attempts to secretly extract the capabilities of competing AI models. The final chapter specifically names seven AI research labs based in China: Alibaba, Moonshot AI, DeepSeek, Zhipu, Xiaomi, SenseTime, and MiniMax.

The methods varied. Some used thousands of fake accounts to send Claude large volumes of requests or bought other users’ conversation records. Others sent requests from customers using their own AI services to Claude and used the responses as training material. In the process, internal corporate documents and even information from government agencies were sent to Anthropic.

◆ Five cases of biological research; institutions and countries withheld

Anthropic disclosed five cases with potential for biological misuse. In most cases, it withheld the research institution, country, and specific biological details. The company said the people involved were practicing scientists and that it could not conclude they intended to cause harm.

△ The first case involved gain-of-function research on chikungunya virus. In May, Claude’s biosafety system blocked a request to draft a related research grant application. Gain-of-function research involves genetically modifying an organism to give it new or enhanced characteristics.

An investigation found that the request came through an AI intermediary platform used by several life-science researchers. The users were in regions where Anthropic did not offer its services, and the platform used U.S. infrastructure to circumvent regional restrictions. The researcher named in the application appeared to be a civilian researcher, but the planned work was to take place at a military research institution. Anthropic began a further investigation after considering both the research and its institutional affiliations.

The platform also had a feature that forwarded requests rejected by Claude to AI systems from other companies. Claude had written much of the code used to implement this feature. The developer had described it to Claude as a way to “reduce excessive refusals.”

The blocking did not stop the activity. According to Anthropic, the operator returned within days, and the related research continued. Based on materials it later reviewed, Anthropic concluded that the work had not stopped at the grant-application stage.

△ The second case involved the use of Claude in research into mammalian adaptation of highly pathogenic avian influenza. Over several weeks, the researcher exchanged thousands of messages with the model, receiving help with research design, data analysis, and prioritizing experiments. However, safety measures that block high-risk biological content restricted the researcher to relatively less capable models, including Sonnet 4 and Haiku 4.5.

Anthropic assessed Claude’s assistance in this case as primarily supporting data analysis, research ideas, and study design. It said the research was also at a very early stage.

△ In the third case, Opus 5 drafted an entire grant application in about an hour for research into immune evasion by orthopoxviruses. The request came through an intermediary account serving about 12 customers.

The remaining two cases involved toxin research. △ One researcher was building a database of toxin peptides to support the development of treatments, including new painkillers and antidepressants, but also studied targets associated with paralysis. △ Another researcher was redesigning multiple toxins computationally and instructed Claude to deliberately obscure the identities of bacterial toxins and viral proteins in quarterly progress reports.

Anthropic said these cases illustrate a fundamental problem in biological research: information needed to develop vaccines and treatments can also be used in dangerous research. It concluded that technical details alone are not enough to determine a user’s intentions accurately, and that systems that simply block high-risk content are insufficient.

◆ 4,700 AI personas across more than 20 dating apps

Anthropic identified a Chinese app developer in its section on fraud. The case is designated GTG-15001 in the report. The company created more than 20 dating apps, where Claude played AI personas that chatted with users. The services, however, were advertised as though all the characters were real people.

Over a two-week period in April this year, the company operated more than 4,700 AI personas. They exchanged about 2.36 million messages with at least 25,000 people. Real people were also deployed to make the interactions seem more natural, at a ratio of roughly one real person for every three AI personas.

Human operators handled tasks AI personas could not perform, such as video calls and following users on social media. Their role was to make users believe they were interacting with real people. Another AI model suggested short replies for the human operators, while a separate image model generated profile pictures.

The AI personas were instructed not to reveal that they were automated and to avoid requests for video calls or photos. When no real person was available, the service also used fake likes and visit records, along with prerecorded videos. It separately tracked whether users had begun to suspect that the person they were talking to was a bot.

Anthropic’s review of some conversations found signs that Claude recognized problematic situations, including cases in which users disclosed serious illness or extreme distress. Even so, Claude did not end the conversations; it continued responding while staying in character.

◆ Seven China-based AI labs tried to extract Claude’s capabilities

The report’s final chapter covers “unauthorized distillation.” Distillation is itself a legitimate AI training method: responses from a high-performing “teacher model” are used to train a smaller “student model.” It is widely used in the AI industry to improve model performance with fewer resources.

The problem arises when a company extracts the capabilities of another company’s model without permission. Anthropic defines unauthorized distillation as secretly extracting a model’s capabilities at industrial scale and replicating them in another model. It said that, since first publicly discussing the issue in February, it had identified additional attacks by seven China-based research labs.

They used proxy services or created thousands of fake accounts. Some also bought records of conversations between users and AI from third parties. Anthropic found cases in which requests from users of a company’s own service were secretly sent to Claude to obtain its responses and reasoning data.

◆ Alibaba: 151 million requests over three months

The largest attack was linked to Alibaba, and is designated GTG-16005 in the report. It targeted the reasoning processes of Opus 4.6 and 4.7, extracting the steps Claude took to reach its answers and turning them into training material.

At its peak, the attack generated about 3 million requests a day and used more than 3,500 fake accounts. Anthropic said the collected material was used to transfer Claude’s capabilities to Alibaba’s Qwen 3.5, 3.6, and 3.7 models. From May through July this year, it attributed more than 151 million requests to Alibaba.

The operation was also organized. The initial group alone consisted of about 5,000 accounts. Residential proxies, disposable email addresses, and virtual cards were used to hide the identities of those accessing the service. When Anthropic blocked that group of accounts, the traffic was quickly shifted to another group.

Alibaba also used Claude in its own AI research. The investigation found that it used Claude to build reinforcement-learning environments and model-development infrastructure, as well as to study model architecture.

◆ Users thought they were using Kimi, but Claude supplied the answers

The Moonshot AI case involved a different approach. It is designated GTG-16002 in the report. According to Anthropic’s investigation, Moonshot sent some customer requests to Claude instead of its own AI, Kimi. Users believed they were using Kimi, but in fact received answers generated by Claude.

One confirmed period lasted ten days. During that time, about 300,000 customer requests were sent to Anthropic, most of them to Opus. Moonshot also stored some conversations. It later built a system to extract Claude’s reasoning records and use them for model training.

The company also bypassed a safeguard designed to prevent Claude from revealing its full internal reasoning. It did so by reusing reasoning-related information left in previous responses in new conversations, a method that could elicit the full reasoning record.

Customers’ sensitive information was transferred as well. One user whom Anthropic believed might be affiliated with the People’s Liberation Army entered CCTV surveillance material about a specific person, thinking they were using Kimi. That request, too, was routed to Claude. Anthropic said it could not determine whether Moonshot had informed customers about this practice.

◆ Russian government access credentials exposed through DeepSeek

Anthropic found similar activity involving DeepSeek, designated GTG-16001 in the report. DeepSeek routed some users’ requests to Claude Opus. Anthropic concluded that sensitive information was likely transferred without users’ knowledge or consent.

The material included internal documents from a Chinese technology company. An employee, believing they were using DeepSeek, entered details about major AI projects, including specifications, organizational structure, and strategic objectives. Those materials were also sent to Claude.

There was also material related to the Russian government. A request from an IT administrator handling data from a government agency linked to Russia’s Ministry of Defense was routed to Claude. In the process, actual credentials for accessing a Russian government database were exposed.

A request from an engineer developing a case-management system for a local Chinese public security bureau was also forwarded. The system could use national identification numbers to compare people’s movements with police records. DeepSeek-related requests exceeded 12.1 million over 14 days in July this year.

◆ Zhipu targeted Fable, then abandoned the effort

The Zhipu case also demonstrated the effectiveness of safeguards. It is designated GTG-16006 in the report. Zhipu extracted Claude’s reasoning records and used them to train its GLM models. About 770,000 requests to organize reasoning records were identified over ten days in June.

More than 3 million requests linked to Zhipu were recorded during the same period. Zhipu later targeted the cyber capabilities of leading U.S. frontier models while preparing its next model, GLM 5.3.

It initially targeted Anthropic’s Fable model. But enhanced cybersecurity safeguards in Fable reduced the effectiveness of the attack, and Zhipu eventually gave up on it. The company then shifted to Opus 4.6 and models from other U.S. AI labs that it considered to have weaker safeguards.

◆ Xiaomi, SenseTime, and MiniMax also identified

Xiaomi also re-entered conversations from its own users into Claude. The case is designated GTG-16008 in the report. Xiaomi stored conversations and coding sessions from MiMo users, then replayed them to Claude. Over 20 days in March and April this year, more than 400,000 requests were sent through over 1,500 accounts.

The process also involved users’ sensitive information, including names, contact details, and corporate materials. Anthropic found no evidence that Xiaomi provided Claude’s responses directly to customers. Instead, the investigation found that the conversations were used to create training material.

SenseTime took a different route. It bought records of conversations between users and Claude from a third-party data broker and used them for model distillation. It also used Claude to build a distillation system and to run and manage training operations.

MiniMax set up its own proxy service through a shell company. The service offered only Anthropic and OpenAI models, not Chinese models—including MiniMax’s own. Anthropic viewed this as evidence that the service was intended to collect conversations between users and U.S. frontier models for use in training MiniMax’s models.

◆ Capabilities can be transferred, but safeguards do not come with them

Anthropic considers unauthorized distillation particularly dangerous for one reason: even if another lab obtains Claude’s capabilities, the safeguards applied to Claude do not necessarily transfer with them.

Anthropic believes that distilling general reasoning capabilities alone can increase dangerous capabilities in areas such as biology and cybersecurity, even when the training conversations do not directly address those fields.

The company also raised privacy concerns. DeepSeek, Xiaomi, and Moonshot sent conversations from users of their own models to Claude. Some of those conversations contained sensitive information about individuals, companies, and government agencies. Anthropic said these practices were likely to violate privacy laws and the companies’ own terms of service.

Anthropic has also strengthened its response. Rather than blocking proxy accounts one by one, it identifies the organizations behind them and blocks related accounts together. It operates a dedicated classifier to detect unauthorized extraction and has added safeguards that make Claude summarize, rather than reveal, its internal reasoning verbatim.

Fable 5.1 introduced a feature called “preserved thinking.” It prevents new API accounts from changing the system prompts, tools, and messages that preceded Claude’s reasoning in conversations that continue across multiple turns.

◆ Closing the series

Over four installments, we examined Anthropic’s 154-page report. It covers cases detected and blocked between December last year and August this year across seven areas: cyberattacks, opinion manipulation, surveillance, conventional weapons, biology, fraud, and model distillation.

The cases disclosed here are not representative of all Claude use. Anthropic said it selected the most notable cases and new types of misuse it had identified so far. The report also distinguishes between what it confirmed and what it assessed as possible, and notes where information could not be verified.

Anthropic said it hopes the disclosures will help other AI developers identify similar patterns of misuse on their own platforms. It also cited helping governments and civil society understand how new threats are developing and strengthen collective defenses as goals of the report.