China-Based AI Companies Conducting Industrial-Scale Knowledge Distillation Campaigns Against U.S. Frontier AI Models
China-Based AI Companies Conducting Industrial-Scale (TL-2026-2405), also tracked as AA26-251A, is a high-severity advanced persistent threat campaign, first published 2026-09-08. It is attributed to Moonshot AI (China) with high confidence, affects Anthropic Claude frontier models (3.7, Sonnet 4/4.5, Opus 4.1/4.5/4.8, maps to 11 MITRE ATT&CK / ATLAS techniques (AML.T0008, AML.T0024.002, AML.T0040), and is covered by 9 detection rules and 2 indicators of compromise.
Key facts for TL-2026-2405
- Threat ID
- TL-2026-2405
- Also known as
- AA26-251A, Adversarial Distillation of American AI Models, The $100M AI Heist
- Severity
- HIGH
- Status
- ACTIVE
- Category
- APT
- First published
- 2026-09-08
- Last reviewed
- 2026-09-08
- Attribution
- Moonshot AI
- Attribution confidence
- HIGH
- Nation-state nexus
- China
- Motivation
- ESPIONAGE
- Target sectors
- technology, artificial-intelligence, cloud-computing, software
- Target regions
- North America, Europe
- Detection rules
- 9
- Indicators of compromise
- 2
Malware and tooling in China-Based AI Companies Conducting Industrial-Scale
Malware and tooling: GROK
CISA, NSA, and FBI jointly warn that six China-based AI companies - DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI - are conducting systematic industrial-scale knowledge distillation of proprietary capabilities from U.S. frontier AI models, extracting billions of tokens across millions of requests since at least late 2024, likely with the knowledge of the Chinese government.
How China-Based AI Companies Conducting Industrial-Scale works
The joint advisory AA26-251A, released September 8, 2026 by CISA, NSA, and FBI, describes deliberate, industrial-scale campaigns in which six China-based AI companies systematically distill proprietary reasoning, coding, and agentic capabilities from U.S. frontier AI models. The advisory characterizes this distillation as forming the core - not merely a supplement - of the Chinese AI sector's development strategy, which has 'turned to a comprehensive distillation strategy' to bridge the capability gap with U.S. frontier models. The companies assert these activities are 'likely with the knowledge of the Chinese government,' a finding echoed in the White House NSTM-4 memorandum and State Department messaging to Beijing.
The campaigns target Anthropic Claude (3.7 through Opus 4.8 and Fable 5), OpenAI GPT models (GPT-4o through GPT-5.5), Google Gemini (2.5 through 3 Pro), and xAI Grok. DeepSeek has conducted organized distillation since at least late 2024 to generate synthetic training data for its R1 and V3 models, targeting reasoning capabilities, chain-of-thought (CoT) drafting, and agentic functions. Moonshot AI, active since at least mid-2025, extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data for Kimi-K2, using millions of exchanges targeting agentic reasoning, tool use, coding/data analysis, and computer vision. Alibaba executed what Anthropic called 'the largest known distillation attack to date' - over 28.8 million exchanges through nearly 25,000 fraudulent accounts between April 22 and June 5, 2026, targeting agentic reasoning and software engineering capabilities. MiniMax, distilling since late 2025 to improve its M2 model, used prompt injections against Claude Code and redirected exchanges to newly released Claude models within 24 hours. StepFun was active between late 2025 and early 2026 distilling code and agentic functions for its Step 4 model. Z.AI had distilled billions of tokens of GPT-5.5 and Claude Opus 4.8 data by mid-2026, focused on CoT reasoning.
The adversarial infrastructure spans three access pathways (native APIs, remote cloud providers, and third-party aggregators that obfuscate user metadata) supported by a gray market of API proxies, or 'transfer stations,' that resell frontier-model access at roughly ten percent of official pricing. Enablers include bulk premium subscription exploitation across shared account pools, centralized request-routing infrastructure, automated metadata sanitization, systematic quota and cost optimization, CoT elicitation prompts that instruct models to reconstruct internal reasoning, prompt injection and jailbreaking to reveal hidden reasoning, and automated failover between pathways during blocking attempts. Anthropic's February 2026 investigation measured over 16 million exchanges across approximately 24,000 fraudulent accounts operated by DeepSeek, Moonshot AI, and MiniMax in a single detected wave, using 'hydra cluster' proxy architectures managing 20,000+ accounts concurrently.
The U.S. government response includes White House NSTM-4 (issued April 23, 2026), which found the distillation campaigns 'unacceptable' and commits to information sharing with AI companies, defense-coordination best practices, and accountability measures; the Deterring American AI Model Theft Act of 2026 (H.R. 8283), advanced by the House Foreign Affairs Committee on April 22, 2026, which would mandate extraction-attack assessments, an 'AI Model Extraction Attackers List,' and Entity List/IEEPA sanctions; and a State Department demarche to Beijing. Recommended mitigations for U.S. AI providers include detecting anomalous prompts, accounts, networks, and behaviors; subtly altering responses to suspected distillation attempts to attenuate payoffs; and establishing cross-organization intelligence sharing to correlate distributed campaigns.
MITRE ATT&CK / ATLAS techniques used in TL-2026-2405
resource-development
AML.T0008 Acquire Infrastructure
Exfiltration
AML.T0024.002 Exfiltration via AI Inference API: Extract AI Model; T1048 Exfiltration Over Alternative Protocol
ai-model-access
AML.T0040 AI Model Inference API Access
ai-attack-staging
AML.T0042 Verify Attack
Impact
AML.T0048.004 External Harms: AI Intellectual Property Theft
Execution
AML.T0051 LLM Prompt Injection
defense-evasion
AML.T0054 LLM Jailbreak
Defense Evasion
T1027 Obfuscated Files or Information; T1078 Valid Accounts
Resource Development
Affected products and versions in China-Based AI Companies Conducting Industrial-Scale
- Anthropic — Claude frontier models (3.7, Sonnet 4/4.5, Opus 4.1/4.5/4.8, Fable 5, Claude Code)
Vulnerable versions: Claude 3.7; Sonnet 4; Sonnet 4.5; Opus 4.1; Opus 4.5; Opus 4.8; Fable 5; Claude Code - OpenAI — GPT model families (GPT-4o, GPT-5, GPT-5.1, GPT-5.2, GPT-5.5, GPT-oss-20b)
Vulnerable versions: GPT-4o; GPT-5; GPT-5.1; GPT-5.2; GPT-5.5; GPT-oss-20b - Google DeepMind — Gemini models (2.5 Flash/Pro, 3 Pro)
Vulnerable versions: Gemini 1; Gemini 2.5 Flash; Gemini 2.5 Pro; Gemini 3 Pro - xAI — Grok models (Grok 3 Mini, Grok 4, Grok Code Fast-1)
Vulnerable versions: Grok 3 Mini; Grok 4; Grok Code Fast-1
Remediation for China-Based AI Companies Conducting Industrial-Scale
Patches
- Apply response-alteration countermeasures: varying reasoning depth and stylistic consistency across requests to complicate quality evaluation
Immediate actions
- Implement comprehensive detection of anomalous prompts, accounts, networks, and behaviors
- Monitor subscription-to-usage ratios and immediately-flagged maximum-usage new accounts
- Deploy behavioral fingerprinting for chain-of-thought elicitation and coordinated multi-account activity
- Share infrastructure and behavioral indicators with other model providers, cloud platforms, and API aggregators
Workarounds
- Limit and rate-limit AI service query volume per account and pathway (AML.M0004)
- Never inform suspected distillers of model downgrades, as it would improve their defense evasions
Longer-term hardening
- Establish cross-organization intelligence sharing to correlate distributed distillation campaigns
- Strengthen access controls and verification for educational accounts, research programs, and startups
- Apply NIST AI 100-2e2025 mitigations: differential privacy, pre/post-training safety interventions, prompt instruction hardening
- Deploy AI telemetry logging (AML.M0024) and adversarial input detection (AML.M0015)
Timeline of China-Based AI Companies Conducting Industrial-Scale
- DeepSeek begins organized distillation campaigns against U.S. frontier AI models, generating synthetic training data for its R1 and V3 models; targets reasoning capabilities, CoT drafting, and agentic functions.
- DeepSeek releases R1, a reasoning model whose accelerations were substantially enabled by distilled data acquired through extensive malicious distillation.
- Moonshot AI begins its widespread distillation campaign, ultimately extracting Claude Fable 5 data for Kimi-K3 and GPT-4o data for Kimi-K2 across millions of agentic-reasoning, coding, and computer-vision exchanges.
- Alibaba and MiniMax begin distillation activity: Alibaba targeting software engineering and agentic workflows from Claude 4/Opus and GPT-5; MiniMax targeting CoT reasoning and software engineering for its M2 model, including prompt injections against Claude Code.
- Anthropic publishes 'Detecting and preventing distillation attacks,' attributing over 16 million exchanges across approximately 24,000 fraudulent accounts to DeepSeek, Moonshot AI, and MiniMax, using hydra-cluster proxy architectures managing 20,000+ accounts concurrently.
- House Foreign Affairs Committee advances the Deterring American AI Model Theft Act of 2026 (H.R. 8283, ordered reported 43-0), authorizing Entity List designations, IEEPA sanctions, and an AI Model Extraction Attackers List.
- White House OSTP issues NSTM-4 'Adversarial Distillation of American AI Models,' finding the campaigns unacceptable and committing to information sharing, public-private defense coordination, and accountability measures.
- State Department instructs diplomats to warn foreign counterparts about extraction by DeepSeek, Moonshot AI, and MiniMax and delivers a formal message to Beijing.
- Anthropic writes to the Senate Banking Committee describing the largest known distillation attack to date: over 28.8 million exchanges through nearly 25,000 fraudulent accounts affiliated with Alibaba and Alibaba Qwen between April 22 and June 5, 2026.
- Z.AI had distilled billions of tokens of GPT-5.5 and Claude Opus 4.8 data by mid-2026, focused on developing chain-of-thought reasoning capabilities.
- CISA, NSA, and FBI release joint advisory AA26-251A documenting industrial-scale distillation campaigns by six named China-based AI companies, assessed as likely conducted with the knowledge of the Chinese government.
Sources cited for China-Based AI Companies Conducting Industrial-Scale
- AA26-251A: China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies (CISA/NSA/FBI)
- Anthropic: Detecting and preventing distillation attacks
- White House NSTM-4: Adversarial Distillation of American AI Models
- H.R. 8283: Deterring American AI Model Theft Act of 2026
- Anthropic letter to Senate Banking Committee on Alibaba distillation attack (via CNBC)
- Just Security: The Emerging U.S. Response to Adversarial Distillation
- The Decoder: How China's gray market sells Claude tokens
- Digital Applied: Anthropic Distillation Attacks - DeepSeek, Moonshot, MiniMax
- NIST AI 100-2e2025: Adversarial Machine Learning - A Taxonomy and Terminology of Attacks and Mitigations
- DeepSeek-V3 Technical Report (arXiv)
More in apt
- Nation-State and Financially Motivated Actors Weaponize Claude AI Multi-Agent Frameworks for Automated Cyberattacks and Data Theft
- Midnight Blizzard (GTG-20006) Used Claude AI Agents to Automate Malware Evasion, Hijack Hotel Wi-Fi (CaptiveCrunch), and Take Over WhatsApp Accounts Against Ukrainian/European Government and Drone-Supply-Chain Targets
- Iran Exploits SS7 Roaming Infrastructure and Commercial Ad-Tech to Track US Military Smartphones During Operation Epic Fury
- China-Nexus and India-Nexus Espionage Groups Converge on Pakistani Law Enforcement Digitalization Platforms ("One Target, Two Flags")
- Chinese-Speaking Operator "Nie" Uses SecFlow AI Orchestration Framework (Claude, Qwen, DeepSeek) and GLUTTON Steganographic Webshell in Multi-Country Espionage Campaign
Detection coverage for TL-2026-2405
As of 2026-09-08, Threadlinqs Intelligence publishes 9 detection rule(s) for TL-2026-2405 across Splunk SPL, Microsoft KQL and Sigma, covering 2 indicator(s) of compromise. The whole corpus is readable without an account; a free account unlocks full detection query text in Splunk SPL, Microsoft KQL and Sigma; paid tiers add raw indicator values, correlation and the MCP server. Threadlinqs MCP server · View plans.