Wain AI/Tech Blog

AI news and trends worldwide, updated nearly every day

Anthropic Publishes Its September Threat Report - Distillation by Seven China-Based Labs, and Traffic Relayed to Claude Without Users Being Told

Image: Anthropic official website

Anthropic Publishes Its September Threat Report - Distillation by Seven China-Based Labs, and Traffic Relayed to Claude Without Users Being Told

Anthropic published a threat intelligence report on September 10, 2026. It says that since the first disclosure in February it has identified distillation attacks from seven China-based labs, and that DeepSeek, Moonshot and Xiaomi relayed their own users' conversations to Claude without telling them.

On September 10, 2026, local time, Anthropic published a threat intelligence report titled “Detecting and countering misuse of AI: September 2026.” It covers activity the company disrupted between December 2025 and August 2026, spread across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation1.

In the illicit distillation section, Anthropic says that since its first disclosure in February 2026 it has identified and disrupted further distillation attacks from seven labs based in China. The report also raises a second issue, separate from the harm distillation does to Anthropic itself: the claim that DeepSeek, Moonshot AI and Xiaomi took conversations received from users of their own models and relayed them to Claude without telling those users.

151 Million Exchanges From Alibaba Alone

Distillation uses the outputs of a capable “teacher” model to train a smaller “student” model, and as a technique it is both widespread and legitimate (for the mechanics, see our explainer on model distillation and quantization). Anthropic separates that from what it calls illicit distillation, which its report defines as “an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization.” What makes it possible in practice, the company says, is fraud: networks of fake accounts built on stolen cards, credentials and API keys.

The change in scale is visible in the numbers. Anthropic’s first report, published in February 2026, said it had identified industrial-scale campaigns by three labs — DeepSeek, Moonshot and MiniMax — accounting for more than 16 million exchanges through roughly 24,000 fraudulent accounts3. The campaign attributed to Alibaba (Qwen / Tongyi Lab) in the new report exceeds 151 million exchanges for the May-to-July window alone. At its peak it reached close to 3 million exchanges in a single day, launched from more than 3,500 fraudulent accounts1. Anthropic describes it as the largest distillation attack it has ever measured. TechCrunch puts the report’s total at close to 200 million exchanges attributed to five separate campaigns2.

Alibaba’s pipeline, according to the report, injected a fixed prompt into every request so that Claude would write out its reasoning inside inline tags, then converted those transcripts into data for supervised fine-tuning (SFT). Anthropic says the resulting transcripts were used to distill Claude’s capabilities into Qwen 3.5, 3.6 and 3.7. Beyond distillation, it says Alibaba also used Claude to build internal infrastructure for model development, to construct reinforcement learning environments, and to advance model architecture research1. In July, an internal notice barring employees from using Claude Code at the same company was reported.

The Target Was the Reasoning Trace

Anthropic says the campaigns went after some of Claude’s most valuable capabilities, listing “agentic capabilities and tool use, coding and data analysis, and logical reasoning.” The technical target is narrower than that: the record of how the model gets to an answer, its chain of thought. What normally reaches the user is summarized thinking, not the raw trace. To pull the raw version out, the attackers used techniques close to prompt injection.

The report reproduces several of the prompts. One casts Claude as an expert translator and asks it to render its previous working memory into katakana-only Japanese — a translation task on its face, and a way to get the reasoning trace out in another shape. In another case, an unauthorized lab ran an experiment of more than twelve thousand requests to find out which techniques would work; most were rejected, but the ones that succeeded were then used to launch a larger attack.

Moonshot AI and DeepSeek are described as having circumvented a control Anthropic had put in place, the “thinking signature.” Rather than returning raw thinking, Claude returns a reference to it, which the API uses to look the trace up on later calls. Anthropic’s account is that both companies saved that signature, opened a fresh session, and then coaxed Claude into converting it back into the full trace.

Relayed to Claude Without the User Being Told

The heavier part of the report is what came next. Anthropic says Moonshot AI forwarded customer requests to Claude while presenting them as processed by Kimi, and showed Claude’s responses to users as they were. Over one ten-day span, roughly 300,000 customer requests were relayed to Anthropic, most of them routed to Opus. The proxy network behind it ran on 5,380 fraudulent accounts, and most of those appeared to sit in Singapore or Japan1.

DeepSeek is reported to have relayed traffic to Claude without informing its customers as well. What stands out in Anthropic’s account is how the targets were picked: requests arriving through coding harnesses such as Claude Code, the Claude Agent SDK and OpenCode were identified by strings they contained, tagged, and — for selected users — relayed to Claude Opus1. The users being singled out, in other words, were the ones running agentic development workflows.

What got relayed included material the sender would not have expected a third party to see. Anthropic cites a request from an IT operator working with data from a Russian government agency tied to its Ministry of Defense, which exposed live credentials for a Russian government database, and relayed exchanges from engineers building a case management system for a municipal Public Security Bureau in China. On the Moonshot side, an engineer building an internal system for a large Chinese state-owned enterprise revealed internal code and live credentials from several companies. The user “had no way of knowing that their use of Kimi was being forwarded to Claude,” the report says1. One user handling surveillance data is described as likely affiliated with the People’s Liberation Army — written as Anthropic’s assessment, not as an established fact.

Xiaomi’s case has a different shape. Anthropic’s investigation did not find that Xiaomi served Claude’s responses to its own users; instead, it says, Xiaomi saved exchanges between its MiMo models and its users and replayed them through Claude to generate training data. The relayed requests contained users’ names, contact information and corporate data, and most of them arrived by way of third-party model routing services that people in the United States and Europe use every day1. Anthropic’s view is that this handling of data by DeepSeek, Xiaomi and Moonshot is likely inconsistent with privacy laws and with the labs’ own terms of service.

The implication, then, is that even someone who never signed a contract with a Chinese lab — who only tried its model through a router — may have had their exchanges saved or forwarded somewhere along the path. Anthropic also says SenseTime’s distillation pipeline included transcripts of Claude conversations purchased from third-party data vendors, pointing to a secondary market in which intermediaries log sessions and resell them. On MiniMax, it says the company operated a proxy service through a shell company that does not disclose its relationship to its parent, and notes that the service offers only Anthropic and OpenAI models and no Chinese ones — evidence, in Anthropic’s reading, that the point was to harvest exchanges with US frontier models.

Anthropic’s Defenses, and What Users Can Review

The countermeasures described are layered. Classifiers built to detect adversarial extraction, strengthened alongside the launch of Fable 5. Summarizing internal reasoning before responding, which lowers the training value of any transcript that is stolen. And with Fable 5.1, a feature called “preserved thinking”: for API accounts created since then, the parts of a multi-turn conversation that sit ahead of Claude’s reasoning — the system prompt, the tools, the earlier messages — can no longer be rewritten1. That change was already described as an anti-distillation measure when Fable 5.1 and Mythos 5.1 shipped; this report fills in the attacks behind it.

One detail worth noting is that the strength of a defense appears to have shifted attacker behavior. Ahead of releasing GLM 5.3, Zhipu went after the cyber capabilities of frontier models, but Anthropic observed that it gave up on Fable once the company’s cyber safeguards degraded the attacks, and moved to Opus 4.6 and the leading model of another US lab, which it assessed as having weaker safeguards. The report does not name that other lab.

There is a limit to what a user can check from the outside, but taking inventory of the path is available today. Which model router are your requests going through; what its terms say about storing conversations and passing them to third parties; whether the environment variables and config files you hand a coding agent contain live credentials. The harm described in this report comes less from which model was chosen than from which route the inference travelled and whose hands it passed through.

Note that neither the report nor the coverage of it includes a response from the companies named. The facts rest on Anthropic’s own investigation and attribution — the same situation as in July, when the US administration accused Moonshot over Kimi K3 and raised the possibility of sanctions without showing how it had verified its claims. Anthropic states that the safeguards preventing misuse of Claude do not carry over when its models are distilled by an unauthorized lab1, an argument continuous with the company’s stated position on open-weight models.

Sources

  1. Detecting and countering misuse of AI: September 2026 - Anthropic’s official threat intelligence report (published September 10, 2026)
  2. Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek - TechCrunch
  3. Detecting and preventing distillation attacks - Anthropic official (the first disclosure, February 2026)

We publish the latest AI news nearly every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →