Anthropic has gone public with a major data theft by Chinese conglomerate Alibaba. According to a Forbes report, Alibaba allegedly used 25,000 fraudulent accounts to conduct 28.8 million prompt-response exchanges over three months—extracting an estimated 57.6 billion tokens from Claude. The goal: to harvest Anthropic's knowledge for training its own language model, particularly in advanced areas like agentic reasoning.
The essentials
- Alibaba allegedly accessed Claude via 25,000 fake accounts
- 28.8 million prompt-answer pairs extracted in three months
- Estimated data volume: 57.6 billion tokens
- Anthropic has called on Congress for countermeasures and information sharing
The technique: Distillation as a weapon
The attack exploits a method called distillation—a technique originally designed for legitimate purposes. Normally, AI makers use distillation to transfer knowledge from large models to smaller ones, like a teacher-student relationship. Alibaba appears to have weaponized this approach: instead of legitimate use, Claude was systematically "squeezed" to extract advanced capabilities and integrate them into its own model.
Disguised as normal usage
The insidious part: the operation was designed to look normal. The 25,000 accounts mimicked genuine user behavior to avoid detection. Over three months, millions of prompts were submitted—an industrial-scale data heist only possible through automated systems. Anthropic discovered the pattern and reported it.
What Anthropic demands from Congress
In a letter to Congress, Anthropic calls for concrete measures:
| Demand | Purpose |
|---|---|
| Information sharing between agencies | Early detection of similar attacks |
| Penalties for industrial siphoning operations | Deterring foreign actors |
| Access pattern monitoring | Prevention of mass extraction |
The company warns this won't be the last such attack—a signal to the US government that AI competition increasingly involves security concerns.
What this means for German companies
The Anthropic breach is a wake-up call for every German AI provider and company running proprietary models. If even a well-guarded US firm can be exploited through millions of prompts, German companies aren't automatically safer—likely the opposite, given fewer resources for security monitoring. The question becomes: how do you detect systematic model extraction? What technical and organizational safeguards do you need? At the same time, the case shows that terms of service alone aren't enough—you need active anomaly detection and possibly regulatory support.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




