-
3 minutes, 21 seconds
On Thursday, Anthropic published a report alleging that China-based AI companies are conducting persistent distillation attacks against its Claude models. According to the company, these efforts are not isolated incidents but a sustained campaign aimed at extracting the capabilities of its systems.
The report claims that the attackers use thousands of fraudulent accounts to generate large volumes of outputs from Claude, then use that data to train their own competing models. Anthropic says the activity has continued despite its efforts to detect and block it.
In the report, Anthropic frames the allegations as a direct threat to its business and to the broader AI ecosystem. It warns that distillation attacks allow competitors to replicate advanced capabilities without bearing the cost of original research and development. The company also suggests that such activity undermines incentives for continued investment in frontier AI.
Anthropic has not publicly named all of the companies involved, but it describes the actors as China-based and characterizes the campaign as persistent. The report marks one of the most detailed public accusations of its kind from a major AI developer.
Distillation attacks are the practice at the center of Anthropic's report. In this context, distillation refers to using the outputs of a powerful AI model to train a competing model, effectively transferring capabilities from the stronger system to a weaker one without the original developer's involvement. The technique lets a rival capture much of the value of a frontier model while spending far less on the research and compute that produced it.
Anthropic describes the activity as an attack because it circumvents the intended terms of access. Rather than building capabilities independently, the distilling party relies on queries to a target model and uses the responses as training signal. The concern is not merely competitive: it also raises questions about safety guardrails, since a distilled model may not inherit the same protections as the model it was derived from.
The report frames distillation attacks as a growing problem for developers of frontier AI systems, and it is this framing that shapes the allegations that follow.
The alleged attacks have escalated in recent months. According to Anthropic, the campaign has grown in both scale and sophistication, with the company reporting that the activity intensified as competition in the AI industry heated up. Rather than a single isolated incident, the alleged distillation attempts appear to represent a sustained and expanding effort.
Anthropic's disclosure points to a pattern in which the frequency and ambition of these operations have increased over time. The company frames the escalation as a direct challenge to the safeguards it has built into its models, and it has publicly named the firms it believes are responsible.
This timing is significant. The reported surge comes amid a period of intense rivalry among AI developers, where access to frontier capabilities can translate into a competitive edge. For Anthropic, the recent escalation underscores the difficulty of protecting proprietary technology once a model is exposed to external queries.
The timing of Anthropic's report is no accident. The escalation comes as competition in the AI space has intensified, with rivals racing to close the gap on leading models. Distillation attacks offer a shortcut: instead of paying the full cost of training a frontier system, a competitor can extract much of that capability by querying a deployed model at scale.
That dynamic makes the alleged activity more than a technical curiosity. If distillation can compress years of expensive research into a fraction of the effort, it threatens the advantage that labs like Anthropic have built through massive investment in training and safety work.
The pressure runs in both directions. As more capable models reach the market and the rewards for matching them grow, the incentive to harvest outputs from the best systems rises accordingly. Anthropic's disclosure should be read in that context: a warning issued precisely when the temptation to cut corners is greatest.
Comment