ZeroHour
GBHackerspublished ()ingested Mayura Kathir1· 1 read
Part of a story covered by 13 sources: “US Agencies Accuse Six Chinese AI Firms of Industrial-Scale Model Distillation; Anthropic Details 200 Million Claude Exchanges” — merged summary and timeline →

DeepSeek, Alibaba and Chinese AI Firms Extract Billions of Tokens From U.S. AI Models

mediumAI safety & security exploited in the wildimportance 72
AI summary · glm-5.3-flash

NSA, CISA and FBI advisory AA26-251A accuses DeepSeek, Alibaba and four other Chinese AI firms of industrial-scale distillation of US frontier models.

Joint advisory AA26-251A from NSA, CISA and FBI accuses DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI of extracting billions of tokens from Claude, GPT, Gemini and Grok variants since at least late 2024, likely with Chinese government awareness. Campaigns allegedly used API proxy 'transfer stations', account pools, metadata sanitization and prompt injection to harvest reasoning, coding, agentic and reinforcement-learning capabilities, with techniques mapped to MITRE ATLAS. DeepSeek's R1 and V3 and Alibaba's Qwen families reportedly trained on harvested outputs, and DeepSeek's $5.6 million training-cost claim is disputed as excluding distilled data value. Agencies urge anomaly monitoring, output alteration for suspected extractors, and intelligence sharing across vendors, clouds and aggregators.

  • Advisory AA26-251A issued jointly by NSA, CISA and FBI.
  • Six Chinese firms named; activity operated since at least late 2024.
  • Techniques mapped to MITRE ATLAS: API access, prompt injection, jailbreaks, output collection.
  • DeepSeek's $5.6M training-cost figure disputed as excluding distilled data.
  • Providers told to watch max-quota new accounts and 24/7 repetitive queries.
Full article769 words · extracted from gbhackers.com · click to collapse

U.S. intelligence and cybersecurity agencies have accused six China-based AI companies including DeepSeek, Alibaba, Moonshot AI, MiniMax, StepFun and Z.AI of extracting billions of tokens from leading American AI systems through industrial-scale knowledge-distillation campaigns.

A joint Cybersecurity Advisory, AA26-251A, issued by the National Security Agency, Cybersecurity and Infrastructure Security Agency and FBI, said the campaigns have operated since at least late 2024 and likely occurred with Chinese government awareness.

The agencies characterize the activity as a coordinated effort to turn frontier-model responses into synthetic training data, reducing the cost and time needed to develop competitive domestic models.

Knowledge distillation is a legitimate machine-learning practice in which a smaller “student” model learns from the outputs of a more capable “teacher” model.

The U.S. advisory, however, draws a sharp distinction between authorized research and alleged high-volume, evasive querying intended to reconstruct restricted proprietary functionality.

According to the agencies, the named firms did not merely benchmark U.S. models or use them for ordinary product development.

Instead, they allegedly ran millions of coordinated exchanges to collect outputs across targeted domains, including chain-of-thought-style reasoning, reinforcement-learning optimization, software engineering, agentic workflows, legal tasks, question-answering and creative writing.

The goal was not to copy model weights directly, but to teach domestic models how frontier systems reason, respond, use tools and evaluate quality.

The warning frames that capability extraction as both an intellectual-property and national-security issue.

Large-scale harvesting of model outputs can allow competitors to capture years of expensive research, compute investment, evaluation work and safety engineering without independently bearing the full development cost.

DeepSeek allegedly conducted organized distillation operations against U.S. frontier models beginning in late 2024, focusing on data used to train its R1 and V3 model families.

The advisory claims the company targeted reasoning capabilities, domain-specific functions and specialized optimizations from multiple Claude, GPT, Gemini and Grok variants access.

It further argues that DeepSeek’s widely cited $5.6 million training-cost figure does not account for the value of data acquired through alleged distillation activity.

CISA Researchers said that, the activity allegedly targeted proprietary reasoning, coding, agentic, and specialized capabilities from Claude, GPT, Gemini and Grok model variants.

Alibaba was also accused of using industrial-scale distillation to advance its Qwen model family.

U.S. AI Models Targeted

The advisory says Alibaba allegedly drew from Claude and GPT systems to improve software-engineering functions, customer-service dialogue, virtual-character generation, and training processes involving supervised fine-tuning and reinforcement learning.

Moonshot AI, developer of the Kimi models, allegedly extracted data for coding, mathematics, reinforcement learning and agentic capabilities.

MiniMax reportedly sought chain-of-thought reasoning and code-development functions, while StepFun and Z.AI were accused of targeting advanced coding and reasoning capabilities from newer U.S. model generations.

The agencies said the operations used a gray market of API proxy services, referred to as “transfer stations,” to evade geographical availability controls, obscure customer identity and bypass provider safeguards.

Requests were reportedly routed through native APIs, cloud services, third-party aggregators, account pools and relay infrastructure.

Other alleged tactics include fraudulent or obscured accounts, bulk purchases of premium subscriptions shared among developers, automatic switching to alternative access paths when blocked, metadata sanitization and centralized systems for distributing requests across providers.

Such infrastructure makes detection difficult because no individual account, IP address or platform necessarily reflects the entire campaign.

The advisory maps several techniques to MITRE ATLAS, including AI inference API access, prompt injection, jailbreak activity, collection of model outputs and extraction through inference interfaces.

Of particular concern is attempts to force models to disclose hidden chain-of-thought-style reasoning, which could reveal methods for solving complex coding, logic and agentic tasks.

NSA, CISA and the FBI urged AI providers to monitor anomalous prompts, account behavior, network activity and enterprise-scale throughput patterns.

Indicators include newly created accounts immediately consuming maximum quotas, multiple users sharing subscriptions across diverse IP addresses, uninterrupted 24/7 activity, highly repetitive queries and unusual subscription-to-usage ratios.

The agencies also recommend carefully altering outputs for suspected extraction activity to reduce the value of harvested data, while avoiding disruptions to legitimate users.

Most importantly, they call for intelligence sharing among model vendors, cloud providers, API aggregators and allied governments to expose campaigns distributed across multiple access channels.

The advisory signals that model-output theft is becoming a distinct security discipline: protecting weights is no longer enough.

Frontier AI developers must now defend their inference layer, behavioral data and proprietary capabilities from systematic extraction at internet scale.

Learn 7 Metric-Gated AI SOC Deployment Phases – Download Free AI SOC Deployment Playbook 2026.

Mayura Kathirhttps://gbhackers.com/

Mayura Kathir is a cybersecurity reporter at GBHackers News, covering daily incidents including data breaches, malware attacks, cybercrime, vulnerabilities, zero-day exploits, and more.

Text extracted automatically; images, tables and formatting may be missing. Original: https://gbhackers.com/u-s-ai-models-targeted/