Introduction
Anthropic has published a new report detailing a significant escalation in distillation attacks targeting its Claude family of models. The campaigns, traced to China-based AI firms, represent a growing threat to frontier model security as competition intensifies across the global AI landscape.
What Happened
Over the past several months, unauthorized labs have developed increasingly sophisticated methods to bypass Anthropic's defenses and extract the models' chain of thought. The company identified five distinct campaigns, the largest attributed to Alibaba, which generated 151 million exchanges between May and July 2026 alone, peaking at nearly three million interactions per day. Another campaign linked to Moonshot AI reportedly funneled requests through a network of 5,000 accounts, many tied to Chinese military interests, targeting Anthropic's Opus model with queries including surveillance footage analysis. Additional efforts from DeepSeek and other labs were also documented, with OpenAI previously flagging similar activity. Attackers employed tactics such as framing queries as translation requests to coax the model into revealing its internal reasoning traces, which can then be used to train smaller models via supervised fine-tuning.
Why This Matters
Distillation attacks undermine the competitive advantage of US-developed frontier models by enabling cheaper, smaller alternatives to replicate advanced capabilities. The scale and aggression of the campaigns highlighted in Anthropic's report signal a shift toward more organized and persistent efforts to harvest model behavior at scale. For developers and policymakers, the findings underscore the need for stronger safeguards, transparent training practices, and international cooperation to protect AI intellectual property as the technology becomes increasingly strategic.
Key Takeaways
- Anthropic identified nearly 200 million exchanges linked to distillation attacks across five campaigns.
- Alibaba's campaign was the largest, producing 151 million exchanges in under three months.
- Moonshot AI routes were allegedly tied to Chinese military interests and targeted Opus via surveillance-related queries.
- Attackers used translation framing and other prompts to extract chain-of-thought data.
- The attacks highlight escalating risks to AI model security in an increasingly competitive global market.
Conclusion
As AI competition heats up, the threat of distillation attacks is moving from isolated incidents to coordinated, high-volume campaigns. Anthropic's report serves as a warning that current defenses may not be sufficient against adversaries willing to invest significant resources in extracting model capabilities. The industry will likely need to adopt new security standards, such as obfuscated reasoning outputs and stricter access controls, to safeguard the next generation of frontier models.




Discussion
Join the conversation
Thoughtful reactions, questions, and follow-up ideas help shape the next story.