Review Finds U.S. Model Outputs in Chinese Military-Linked Research

A July 30 Jamestown Foundation review of Chinese research found military- and security-linked institutions using distillation, including outputs from OpenAI and Anthropic models, to train smaller systems for defense, surveillance and cyber uses. SiliconANGLE, summarizing a Reuters review of more than 80 papers and patents, reported that such models can run locally with less compute than frontier training requires.
A Jamestown Foundation analysis published July 30 found that Chinese academic and industry papers from 2024 through 2026 document distillation work involving institutions linked to the People's Liberation Army, the defense industry and state research bodies. Some of the reviewed work used outputs from OpenAI or Anthropic models; other papers examined ways to conceal or attack distillation processes.
SiliconANGLE reported on August 2 that a Reuters review of more than 80 Chinese papers and patent applications found a similar pattern: researchers used outputs from powerful U.S. models to train smaller systems for defense-related applications. Both accounts describe evidence in published research, not a benchmark proving that a specific deployed Chinese defense model matches a U.S. frontier system.
What the documents show
Distillation trains a smaller student model using outputs from a stronger teacher model. It is a common efficiency technique, but the Jamestown analysis distinguishes ordinary use from adversarial distillation that violates provider terms, targets protected capabilities or attempts to hide a model's provenance.
The review cites a PLA-linked research team that used GPT-3.5 outputs to train smaller code-summarization models. It also describes researchers at North University of China using Claude outputs for a classifier proposed for social-media monitoring, and military-university work on black-box attacks and methods for weakening model-watermark signals.
Those examples have different stages and purposes. Some describe completed experiments; others propose applications or attack methods. They support the narrower conclusion that U.S. model outputs and distillation techniques appear in Chinese military- and security-linked research. They do not establish how widely any resulting systems have been deployed or how they perform in operations.
Why the distinction matters
A distilled model can be smaller and cheaper to run locally than its teacher. That can help organizations move selected capabilities onto constrained or isolated hardware, but it does not replace the cost and research required to build a frontier model from scratch.
LDS interpretation: hardware export controls and model-access controls address different parts of the risk. Chip restrictions can raise the cost of large-scale training, while providers still need query monitoring, account-abuse detection and provenance research to identify attempts to extract model behavior through an API. The evidence also argues for separating legitimate compression from covert extraction instead of treating every use of distillation as equivalent.
Key Points
- 1Jamestown's review documents OpenAI and Anthropic outputs in Chinese military- and security-linked distillation research, including code, monitoring and attack-related work.
- 2The evidence comes from published experiments and proposed uses; it does not prove broad operational deployment or frontier-level performance of a specific defense model.
- 3Smaller distilled models can reduce local deployment compute, making model-access monitoring a separate control from restrictions on advanced training chips.
Scoring Rationale
The documentary review spans dozens of papers and military- or security-linked institutions, making the model-access and national-security implications significant. The score remains below the highest tier because published research and proposed uses do not by themselves establish broad operational deployment or performance.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

