Threat reportThreat IntelligenceTL-2026-2915
Coordinated model-distillation campaign against OpenAI: 15,000+ accounts attempt to extract protected model reasoning, linked to Moonshot AI-associated individuals
Coordinated model-distillation campaign against OpenAI (TL-2026-2915), also tracked as OpenAI reasoning-extraction campaign, is a medium-severity tracked intrusion set, first published 2026-10-04. It is attributed to Individuals associated with Moonshot AI (China) with medium confidence, affects OpenAI OpenAI model APIs (protected/encrypted reasoning), maps to 6 MITRE ATT&CK / ATLAS techniques (AML.T0005, AML.T0024, AML.T0024.002), and is covered by 9 detection rules and 4 indicators of compromise.
- Severity
- MEDIUMAssessed severity
- CVEs
- 0None referenced
- Techniques
- 6MITRE ATT&CK / ATLAS
- Actors
- 1Individuals associated with Moonshot AI
- Detection rules
- 9SPL · KQL · Sigma
- IOCs
- 4Indicators of compromise
Key facts for TL-2026-2915
- Threat ID
- TL-2026-2915
- Also known as
- OpenAI reasoning-extraction campaign, Encrypted reasoning replay distillation
- Severity
- MEDIUM
- Status
- MONITORING
- Category
- THREAT_INTEL
- First published
- Last reviewed
- Attribution
- Individuals associated with Moonshot AI
- Attribution confidence
- MEDIUM
- Nation-state nexus
- China
- Motivation
- FINANCIAL
- Target sectors
- technology, artificial-intelligence
- Target regions
- Global
- Detection rules
- 9
- Indicators of compromise
- 4
How Coordinated model-distillation campaign against OpenAI works
OpenAI disrupted a coordinated adversarial model-distillation effort aimed at extracting protected internal reasoning from its models. Activity began at low volume on July 1, 2026, spiked July 24-25 (~16,000 extraction-pattern requests from 4,000+ users), and a wider cluster of 15,000+ users was fully disrupted by July 28, 2026. OpenAI attributed a core cluster to individuals associated with Moonshot AI (developer of Kimi) but said it is unclear whether all operators were a single actor.
OpenAI publicly disclosed on September 30, 2026 that it had disrupted a coordinated model-distillation campaign targeting the protected (encrypted) reasoning of its models. Per OpenAI as relayed by secondary reporting, the operators did not break the encryption or access any database or customer conversation. Instead they manipulated model interactions: encrypted reasoning produced in one conversation was copied into a second conversation, where the model was instructed to decrypt and transcribe it, using attacker-controlled prompts and reusable context. OpenAI closed a pathway that let someone who already held another user's encrypted reasoning replay it and recover its contents.
Timeline and scale: activity began around July 1, 2026 at low volume, spiked on July 24-25 with roughly 16,000 requests matching the extraction pattern from more than 4,000 users, and a broader sweep for related prompt-pattern activity identified a cluster of more than 15,000 users. The whole cluster was fully disrupted by July 28, 2026.
Attribution: OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi, while stating it is unclear whether all operators observed in the period originated from a single actor. Reporting notes OpenAI did not establish involvement of Moonshot founder Yang Zhilin or company leadership. No Moonshot response was documented in the reviewed articles.
Context: the technique class was independently described on August 10, 2026 in the arXiv paper 'Stealing Reasoning Traces from Proprietary LLM APIs', which reported that encrypted reasoning blocks from OpenAI, Anthropic and Google could be replayed into weaker sibling models to recover plaintext traces. Earlier, in February 2026, Anthropic reported industrial-scale distillation of Claude by DeepSeek, Moonshot and MiniMax via roughly 24,000 fraudulent accounts and over 16 million exchanges.
Response: OpenAI banned or restricted fraudulent accounts, tightened signup and infrastructure controls, expanded monitoring for related accounts, coordinated with third-party services the activity ran through, closed the replay pathway, added checks that detect and hold streamed output that might expose reasoning, strengthened cross-user and cross-organization protections, and shared intelligence via the Frontier Model Forum and government channels. No CVE, CVSS score, or network indicators (IPs, domains, hashes) have been published. The OpenAI primary report could not be fetched directly (HTTP 403) and facts here come from consistent secondary reporting.
MITRE ATT&CK / ATLAS techniques used in TL-2026-2915
AI Attack Staging
AML.T0005 Create Proxy AI Model
Exfiltration
AML.T0024 Exfiltration via AI Inference API; AML.T0024.002 Extract AI Model
AI Model Access
AML.T0040 AI Model Inference API Access
Execution
AML.T0051 LLM Prompt Injection
Resource Development
Affected products and versions in Coordinated model-distillation campaign against OpenAI
- OpenAI — OpenAI model APIs (protected/encrypted reasoning)
Vulnerable versions: Prior to replay-pathway fix, July 2026
Fixed in: Replay pathway closed by OpenAI (server-side)
Remediation for Coordinated model-distillation campaign against OpenAI
Patches
- OpenAI closed the encrypted-reasoning replay pathway server-side; no customer action required
Immediate actions
- Block or restrict the fraudulent accounts identified in the cluster
- Hold or inspect streamed model output that may expose protected reasoning
- Monitor for related accounts and shared infrastructure
Workarounds
- Do not publish raw agent transcripts that contain encrypted reasoning blocks
Longer-term hardening
- Bind encrypted reasoning artifacts to the originating user, organization and conversation so they cannot be replayed cross-account
- Strengthen signup and infrastructure controls against bulk fraudulent account creation
- Share distillation-abuse intelligence through industry bodies such as the Frontier Model Forum
Timeline of Coordinated model-distillation campaign against OpenAI
- Anthropic publicly reports industrial-scale distillation of Claude by DeepSeek, Moonshot and MiniMax via roughly 24,000 fraudulent accounts and over 16 million exchanges (precedent for the same class of abuse).
- Low-volume reasoning-extraction activity against OpenAI models begins.
- Spike begins: ~16,000 requests matching the extraction pattern over July 24-25 from more than 4,000 users.
- Peak ends; a broader sweep links related prompt-pattern activity to a cluster of more than 15,000 users.
- OpenAI fully disrupts the cluster, banning or restricting fraudulent accounts.
- arXiv paper 'Stealing Reasoning Traces from Proprietary LLM APIs' shows encrypted reasoning blocks can be replayed into weaker sibling models to recover plaintext across OpenAI, Anthropic and Google.
- OpenAI publishes its report attributing a core cluster to individuals associated with Moonshot AI, with caveats on single-actor attribution.
- GBHackers reports the disruption and mitigations.
Sources cited for Coordinated model-distillation campaign against OpenAI
- OpenAI Blocks 15,000 Requests Trying to Extract Protected Model Reasoning
- Disrupting a coordinated model-distillation campaign (OpenAI)
- OpenAI says people linked to Moonshot AI ran a campaign to extract its reasoning
- OpenAI says Moonshot-linked operators tried to extract hidden model reasoning
- OpenAI Says Moonshot AI-Linked Users Tried to Extract its Models' Secret Reasoning
- Detecting and preventing distillation attacks (Anthropic)
- Encrypted Reasoning Traces Let Attackers Steal Hidden Chain-of-Thought (CSA research note)
- Stealing reasoning traces (Simon Willison's Weblog)
Detection coverage for TL-2026-2915
As of 2026-10-04, Threadlinqs Intelligence publishes 9 detection rule(s) for TL-2026-2915 across Splunk SPL, Microsoft KQL and Sigma, covering 4 indicator(s) of compromise. The whole corpus is readable without an account; a free account unlocks full detection query text in Splunk SPL, Microsoft KQL and Sigma; paid tiers add raw indicator values, correlation and the MCP server. Threadlinqs MCP server · View plans.