AI distillation - the technique of training a smaller "student" model on a powerful "teacher" model's outputs - has become one of the most contested issues in global AI competition. Anthropic and OpenAI allege that Chinese firms including Moonshot AI, DeepSeek, MiniMax, and Alibaba have run industrial-scale campaigns to extract reasoning, coding, and agentic capabilities from U.S. frontier models like Claude - using tens of thousands of fraudulent accounts and tens of millions of exchanges. At the same time, PLA-linked research institutions have published work applying distillation to military and security applications.
EdgeTheory's latest assessment cuts through the noise to separate legitimate model distillation from suspected unauthorized extraction, and explains why this fight matters for the future of U.S. AI leadership.
What You'll Learn
- How distillation attacks actually work - the technical mechanics behind how a "student" model learns from a "teacher" model's outputs, and why reasoning traces and agentic trajectories are far more valuable to extract than simple answers.
- Who's behind the alleged campaigns - a breakdown of the key suspected actors, including Moonshot AI, DeepSeek, MiniMax, Alibaba/Qwen, and PLA-linked research institutions, and what's actually been proven versus alleged.
- Why this threatens the U.S. AI advantage - how distillation shifts the competition from controlling advanced chips to protecting model capabilities, and what that means for the future of export controls.
- The geopolitical fallout for the global AI race - how distillation could accelerate China's AI catch-up, expand its global influence, and drive deeper U.S.-China technology fragmentation.
- The real national-security stakes - how commercially developed AI capabilities could diffuse into military, cyber, surveillance, and intelligence applications, and what policymakers are already proposing in response.
Understanding these dynamics is no longer optional for anyone tracking U.S.-China technology competition.
As the fight over AI shifts from silicon to software, the actors, methods, and policy responses covered in this report will shape how the next phase of the AI race is won - or lost.
Download the report
AI distillation has emerged as a growing point of friction in the global AI race as U.S. officials and frontier AI companies raise concerns that foreign actors are using American model outputs to accelerate their own development. While distillation is a legitimate technique for transferring capabilities from a powerful “teacher” model to a smaller “student,” alleged industrial-scale extraction creates broader economic and national-security concerns when proprietary capabilities are obtained without authorization. This assessment seeks to answer four key questions:
Distillation could alter the balance of AI capabilities between the United States and China. Understanding these activities is increasingly important as competition expands beyond control of advanced semiconductors.
By distinguishing legitimate model distillation from suspected unauthorized extraction, this assessment identifies how distillation could alter the balance of AI capabilities between the United States and China. Understanding these activities is increasingly important as competition expands beyond control of advanced semiconductors to the protection of model outputs, intellectual property, and commercially developed AI capabilities.
This assessment used WatchBuilder to refine and customize data collection around emerging trends by automatically evaluating the coverage provided by existing Watches against the articles and content associated with each trend.

Analysts used WatchBuilder to configure a new Watch Group by selecting the relevant module and trend, defining the target Watch Group, and setting parameters.
WatchBuilder’s LLM agent assessed the available content and incumbent Watches to identify gaps or redundancies in coverage and determine which Watches should be added, retained, or pruned. This supported the assessment by ensuring that collection remained focused on the most relevant content while reducing the need for analysts to manually review and reconfigure Watches as the information environment changed.

WatchBuilder generated a Watch Group Agent Job to evaluate the selected trend and automatically determine the Watches needed to provide targeted coverage
WatchBuilder also supported the rapid expansion and optimization of collection by automatically creating new Watches when additional coverage was required and queuing them for caching. Analysts were able to monitor the workflow and review resulting Watch changes through the job detail page. By automating the process of evaluating and adjusting Watch coverage, WatchBuilder supported a more responsive analytic workflow, allowing analysts to spend less time managing collection and more time assessing the narratives, actors, and trends identified within the resulting data.
Distillation could undermine a central pillar of the US’ AI advantage by shifting competition from controlling compute to protecting capabilities.
Moonshot AI, DeepSeek, MiniMax, and Alibaba-linked operators have faced direct allegations of large-scale extraction from U.S. frontier models, while PLA-linked and Chinese security institutions have publicly researched distillation for military and security applications. These activities demonstrate both commercial and strategic interest, but available evidence does not establish that Chinese frontier models were primarily developed through unauthorized distillation.
U.S. semiconductor restrictions constrain China's access to advanced hardware, but distillation provides another pathway for acquiring selected reasoning, coding, cyber, and agentic capabilities through model outputs. At scale, this could accelerate Chinese catch-up and increase pressure for controls extending beyond chips to APIs, model access, and overseas compute.
Unauthorized extraction could allow competitors to benefit from costly U.S. research while offering capable models at lower prices, reducing the durability of U.S. technological advantages. More strategically, those capabilities could diffuse through lower-cost Chinese models and into military, cyber, surveillance, and global commercial applications. The resulting economic exposure may be substantial, but available evidence does not support a defensible aggregate dollar estimate.
AI distillation is a training technique in which a smaller “student” model learns from the outputs of a more powerful “teacher” model, allowing it to replicate selected capabilities without accessing the teacher’s underlying weights or original training data. Distillation is a legitimate AI-development technique; it becomes a security concern when proprietary model capabilities are systematically extracted without authorization.
Actors query frontier models at scale and collect outputs that can improve multiple stages of model development. Adversarial distillation can support synthetic data generation, chain-of-thought extraction, data cleaning, and reward modeling, allowing a target model to serve not only as a source of training examples but also as a grader for reinforcement learning. These outputs can then support pre-training, supervised fine-tuning, or reinforcement-learning pipelines. Rather than eliminating the need for domestic model development, distillation can allow developers to make larger capability gains at individual stages of training by leveraging work already performed by frontier laboratories.
Some proxy networks operated more than 20,000 accounts simultaneously, replacing disabled accounts and distributing requests across different access pathways. The campaigns concentrated on high-value capabilities including reasoning, coding, tool use, computer use, and agentic behavior.
At industrial scale, campaigns may rely on fraudulent accounts, proxy services, and distributed access infrastructure to evade provider restrictions. Anthropic identified approximately 24,000 fraudulent accounts generating more than 16 million Claude exchanges across campaigns attributed to DeepSeek, Moonshot AI, and MiniMax. Some proxy networks operated more than 20,000 accounts simultaneously, replacing disabled accounts and distributing requests across different access pathways. The campaigns concentrated on high-value capabilities including reasoning, coding, tool use, computer use, and agentic behavior.
Importantly, distillation does not allow an actor to simply reproduce an entire frontier model. Closed-model APIs generally provide outputs rather than access to model weights, internal probability distributions, or the complete training process. The value of extraction therefore varies by technique. Simple imitation of final responses appears less effective at transferring underlying capabilities, while reasoning traces, reward-model outputs, and agentic trajectories may provide substantially more valuable training data. Distillation can consequently accelerate development in targeted areas without eliminating the need for pre-training, compute, domestic engineering, reinforcement learning, or original research. Distilled capabilities may also generalize less effectively beyond the types of tasks represented in the extracted training data.
This suggests that Moonshot has become not only a prominent suspected actor, but also a focal point for wider concerns over U.S.-China technological competition.
China is the principal foreign ecosystem associated with suspected AI distillation activity. Moonshot AI, DeepSeek, and MiniMax represent the strongest commercial suspects, while PLA-linked and security research institutions provide the clearest indication that distillation has potential applications beyond commercial competition, including military, cyber, surveillance, and national-security objectives.
Moonshot AI is the most prominent current Chinese actor accused of unauthorized distillation. U.S. officials allege that Moonshot conducted large-scale extraction from Anthropic models and used multiple access methods to evade detection. Its Kimi K3 model has drawn particular scrutiny because of its rapid development and frontier-level capabilities, although publicly available evidence does not conclusively establish that K3 was trained using illicitly obtained outputs.


EdgeTheory Emotion Profile Classifier
The emotion profile indicates that reporting around Moonshot’s alleged distillation activity was driven primarily by anger and fear, with both emotions spiking around key allegations in late July. High-anger narratives focused on claims that Chinese firms covertly extracted U.S. model capabilities, while fear reflected broader concerns over intellectual-property loss and China’s ability to narrow the U.S. AI advantage. This suggests that Moonshot has become not only a prominent suspected actor, but also a focal point for wider concerns over U.S.-China technological competition.
Anthropic has accused DeepSeek of participating in industrial-scale distillation of Claude, while OpenAI has separately reported detecting attempts by DeepSeek to distill its models. DeepSeek's ability to produce highly capable models despite U.S. restrictions on advanced computing has increased scrutiny of its development methods. However, available evidence does not prove that its flagship models were primarily developed through unauthorized distillation.
Anthropic accused operators affiliated with Alibaba and its Qwen AI lab of conducting its largest known distillation attack, allegedly generating more than 28.8 million Claude exchanges through nearly 25,000 fraudulent accounts between April and June 2026.
Anthropic identified MiniMax alongside Moonshot and DeepSeek in alleged industrial-scale distillation activity. Less public information is available regarding MiniMax's specific methods or targeted capabilities, making the allegations less developed than those concerning Moonshot.
Anthropic accused operators affiliated with Alibaba and its Qwen AI lab of conducting its largest known distillation attack, allegedly generating more than 28.8 million Claude exchanges through nearly 25,000 fraudulent accounts between April and June 2026.

EdgeTheory Narrative Classifier on Alibaba’s extraction of data from Claude AI
The campaign reportedly targeted advanced capabilities including agentic reasoning and software engineering. However, these claims remain allegations and do not establish that Alibaba’s Qwen models were developed through illicit distillation.
Research has included military code processing, UAV applications, attack knowledge, surveillance, and methods for making distillation attacks more difficult to detect.
These publications demonstrate capability and strategic interest but do not establish that these institutions participated in the commercial extraction campaigns alleged by U.S. companies.
Chinese military, academic, and security-linked institutions have publicly researched or applied distillation techniques using U.S. models. Reported actors include PLA Unit 96941, Army Engineering University, the National University of Defense Technology, Xidian University and a People's Armed Police counterterrorism laboratory, and the University of Science and Technology of China. Research has included military code processing, UAV applications, attack knowledge, surveillance, and methods for making distillation attacks more difficult to detect.
These publications demonstrate capability and strategic interest but do not establish that these institutions participated in the commercial extraction campaigns alleged by U.S. companies.

Edge Theory Narrative Attack Classifier
The reporting suggests that distillation extends beyond commercial competition into Chinese military and national-security applications. The use of U.S. model outputs for defense-related systems indicates that distillation could provide Chinese institutions with advanced AI capabilities while reducing dependence on restricted U.S. hardware. This makes distillation particularly significant to U.S.-China competition because commercially developed American AI capabilities could potentially be adapted for military, surveillance, and cyber purposes, partially undermining the strategic objectives of U.S. technology controls.
Distillation therefore complicates U.S. efforts to maintain an AI advantage because protecting chips and physical infrastructure alone may be insufficient if valuable capabilities can also be extracted through access to U.S. models.
Taken together, the actors represent two distinct dimensions of China's distillation ecosystem. Commercial firms such as Moonshot AI, DeepSeek, MiniMax, and Alibaba are primarily associated with efforts to accelerate model development and commercial competitiveness, while PLA-linked and security institutions demonstrate potential military, cyber, and surveillance applications. The overlap suggests that the strategic significance of distillation extends beyond whether individual Chinese models copied U.S. capabilities: U.S.-developed model outputs may contribute to a wider Chinese AI ecosystem spanning both commercial and national-security objectives.
Distillation attacks could compress the U.S.-China AI capability gap, weaken the strategic value of semiconductor controls, and accelerate the global diffusion of Chinese AI systems. The resulting competition is likely to shift from simply developing the most advanced model toward controlling how rapidly capabilities can be protected, replicated, deployed, and integrated into global technology ecosystems.
Distillation can narrow the U.S.-China capability gap by allowing Chinese developers to transfer selected reasoning, coding, cyber, and agentic capabilities from frontier models without reproducing the full cost of their development. This is particularly significant under U.S. semiconductor restrictions because it provides another pathway for acquiring advanced AI capabilities despite constraints on cutting-edge compute. However, distillation supplements rather than replaces domestic research, engineering, data, and infrastructure.

Edge Theory Narrative Attack Classifier
This narrative item strengthens the assessment that distillation is part of a broader U.S.-China competition over AI and intellectual property. Its emphasis on protecting U.S. AI research from foreign espionage suggests that policymakers increasingly view advanced model capabilities as strategic assets comparable to other sensitive technologies. Distillation therefore complicates U.S. efforts to maintain an AI advantage because protecting chips and physical infrastructure alone may be insufficient if valuable capabilities can also be extracted through access to U.S. models.
Distillation shifts part of the AI competition from control over chips to control over model outputs. Semiconductor restrictions can limit access to advanced computing hardware, but they cannot fully prevent capabilities developed on U.S. infrastructure from being transferred through APIs or other model interactions. This could reduce the ability of hardware controls alone to preserve a durable U.S. technological lead and encourage additional restrictions on APIs, overseas compute, and model access.
Lower-cost and open-weight Chinese models can spread extracted or independently developed capabilities across foreign companies, governments, cloud platforms, and applications. China's competitive advantage may therefore come less from consistently possessing the world's most capable model and more from rapidly diffusing sufficiently advanced models at lower cost, particularly in developing markets. This could increase Chinese influence over global AI ecosystems, technical standards, and infrastructure dependencies.
Suspected distillation attacks are likely to intensify U.S. scrutiny of Chinese AI firms and could contribute to sanctions, Entity List designations, tighter identity verification, restrictions on frontier-model access, and greater monitoring of overseas computing infrastructure. The result could be increasingly separate U.S.- and China-centered AI ecosystems, with countries facing greater pressure over which models, infrastructure, and standards to adopt.
IAPS has recommended considering Entity List restrictions against firms conducting unauthorized distillation, evaluating sanctions under the Protecting American Intellectual Property Act
These responses are increasingly moving from hypothetical to active policy debate. IAPS has recommended considering Entity List restrictions against firms conducting unauthorized distillation, evaluating sanctions under the Protecting American Intellectual Property Act, and developing a NIST-led AI Distillation Defense Framework establishing common standards for access controls, detection, monitoring, and response. Such measures would expand U.S.-China AI competition beyond semiconductor controls toward regulation of the broader infrastructure through which frontier-model capabilities can be accessed and transferred.
Distillation can transfer commercially developed capabilities into military, cyber, surveillance, or intelligence applications at comparatively low cost. Chinese military-linked research involving U.S. model outputs for UAVs, military code processing, and monitoring illustrates how private-sector AI innovation could diffuse into state-security systems. A distilled model may also reproduce useful capabilities without retaining the original provider's safeguards or monitoring controls.
Chinese military-linked research involving U.S. model outputs for UAVs, military code processing, and monitoring illustrates how private-sector AI innovation could diffuse into state-security systems. A distilled model may also reproduce useful capabilities without retaining the original provider's safeguards or monitoring controls.
AI distillation is emerging as a significant dimension of U.S.-China technological competition. While distillation itself remains a legitimate AI-development technique, suspected industrial-scale extraction could allow competitors to acquire advanced capabilities without reproducing the full cost of their development. China represents the principal foreign ecosystem associated with suspected activity, spanning commercial AI firms and military-linked research institutions, although available evidence does not establish that Chinese frontier models were primarily developed through illicit distillation.
The broader strategic risk is the accelerated diffusion of U.S.-developed AI capabilities. Distillation could narrow the U.S.-China capability gap, reduce the effectiveness of hardware-focused technology controls, and enable commercial AI advances to spread into competing global platforms and national-security applications. Maintaining U.S. AI leadership will therefore increasingly depend not only on controlling access to advanced compute, but also on protecting the capabilities produced by it while preserving legitimate and responsible AI development.