A security bulletin has accused several Chinese artificial intelligence companies of harvesting billions of tokens from prominent American models. The report suggests that entities such as DeepSeek and Alibaba have been targeting systems like Claude and GPT since late 2024.

Advertisement

The DeepSeek and Alibaba token harvest

The security bulletin alleges that a group of Chinese firms—including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z. AI—have been aggressively extracting data from leading U.S. models. According to the report, these companies have targeted high-profile systems such as Anthropic’s Claude, OpenAI’s GPT, Google’s Gemini, and xAI’s Grok.

This massive data collection,which reportedly involves billions of tokens, has been described as a targeted and malicious effort.. The scale of these exchanges suggests a systematic attempt to replicate the capabilities of the world's most advanced AI through a massive volume of requests, rather than through original research.

Anthropic’s warning on the risks of distilled models

Model distillation, a process where a "student" model learns from a more powerful "teacher" model, is at the heart of this controversy. While the technique is a standard industry practice, Anthropic has warned that illicitly distilled models often lack the essential safety safeguards found in the original versions. The report notes that such uncontrolled distillation could pose significant risks to national security by creating less stable or more easily manipulated AI.

The industry's response to these risks appears to be shifting. While OpenAI CEO Sam Altman previously condemned DeepSeek for unauthorized distillation, he has recently suggested that such activities are no longer among the top ten concerns for his company. This pivot comes even as Anthropic continues to research how to prevent its own models from being rescripted through these methods, highlighting a growing divide in how top-tier labs prioritize the threat of data theft.

G20 discussions and the push for a coordinated slowdown

The tension between rapid innovation and security has reached the global stage , specifically within G20 sessions. Executives from major players like Nvidia, Palantir, Anthropic, and OpenAI have recently advocated for a coordinated slowdown in the release of new AI models. They argue that such a pause is necessary to ensure legal compliance and to maintain oversight as AI capabilities expand .

However, observers have pointed out a contradiction in this stance.. While industry leaders call for increased regulation and caution, they simultaneously continue to deploy increasingly capable models at a rapid pace. This creates a friction point between the public desire for safety and the commercial drive for dominance, leading some to question if the call for regulation is a way to manage the fallout of their own rapid deployment cycles.

The unverified role of the Chinese government

Several critical questions remain regarding the scale and intent of these data extractions. While the security bulletin suggests these activities may have the "tacit approval" of the Chinese government, the report does not provide direct evidence of state-level direction. it remains unclear how much of this harvesting is driven by independent commercial interests versus state-backed R&D mandates.

Furthermore, the effectiveness of proposed technical defenses remains unproven. While experts are evaluating tools like federated learning, differential privacy, and the "Nexus Group" approach to lock models to their owners, it is yet to be seen if these can truly stop a determined actor from extracting high-value knowledge. The industry is also looking toward solutions like AI StackGuard, but the rapid pace of Chinese model development may outstrip these defensive measures.