Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
Aleph Alpha found Chinese models often repeat state doctrine or refuse questions on Tiananmen, Taiwan, and Xinjiang.
Aleph Alpha benchmarked Alibaba Qwen, DeepSeek, and Moonshot Kimi on 967 sensitive topics including Tiananmen, Taiwan, and Xinjiang. Its scorer rated only 17 to 41 percent of answers balanced; the rest repeated Chinese state positions, deflected, or refused. DeepSeek V4 Pro refused about two-thirds of questions, while Claude Sonnet 5 and Mistral Small were balanced 70 and 92 percent of the time. Nvidia's Nemotron Cascade 2 showed party-line answers in 17 percent of cases, which Aleph Alpha links to roughly 3,500 of 9.3 million training examples generated with DeepSeek and Qwen.
- Aleph Alpha scored Qwen, DeepSeek, and Kimi on 967 politically sensitive topics.
- Only 17 to 41 percent of responses were rated balanced by its scorer.
- DeepSeek V4 Pro refused about two-thirds of sensitive questions.
- Claude Sonnet 5 and Mistral Small were balanced 70 and 92 percent of the time.
- Nemotron Cascade 2 echoed party lines after training on DeepSeek and Qwen outputs.
Full article595 words · extracted from the-decoder.com · click to collapse
Chinese AI models frequently toe the party line when asked politically sensitive questions, according to a study by Aleph Alpha.
The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, giving it a commercial interest in distinguishing its models from Chinese competitors.
In a benchmark Aleph Alpha developed, the company tested models from Alibaba (Qwen), DeepSeek, and Moonshot AI (Kimi). The test covered 967 hand-picked taboo topics like Tiananmen, Taiwan, and Xinjiang. The company's own AI scoring system rated only 17 to 41 percent of responses as balanced. The rest repeated state doctrine, deflected, or refused to answer.
The findings line up with China's AI regulations, which require "socialist core values" in public-facing models. They also match recurring anecdotal reports and earlier audits.

Pro-China bias shows up even in unrelated answers
The pro-China slant can also appear in answers to questions that don't mention China. When asked about censorship in the United States, Qwen 3.6 starts with a seemingly balanced answer. It then closes with a defense of China's stance on global internet governance. "Many countries, including China, also manage information to ensure social stability and national security," the response reads.

An earlier study by the Central European Institute of Asian Studies (CEIAS) also found this spillover effect. When terms like human rights, opposition, or surveillance came up, the models often responded with standard Beijing talking points. These included the "principle of non-interference in internal affairs" and a "community with a shared future for mankind."
Distilled Chinese training data can carry CCP values into other models
Aleph Alpha also takes aim at a direct competitor. Nvidia's Nemotron Cascade 2 showed party-line patterns in 17 percent of responses. Aleph Alpha attributes this to roughly 3,500 of its 9.3 million training examples, which were generated using DeepSeek and Qwen.
When asked to draft a speech supporting recognition of Taiwan, the model refused and instead produced a patriotic response defending Beijing's One-China principle. Nvidia is increasingly pushing its own models into the government and enterprise market, where Aleph Alpha and Cohere also want to compete.
Language models generally carry cultural and political values because their training data overrepresents certain viewpoints or can be shaped through deliberate data selection. Researchers warn that repeated exposure to uniform AI outputs could influence how billions of users think and express themselves.
There are also political efforts to shape AI models along ideological lines in the United States. Elon Musk has repeatedly had his Grok AI modified to produce right-leaning responses. Studies nevertheless suggest that models tend to lean left, possibly because their answers draw more heavily on scientific evidence. For the EU, that leaves a choice between two foreign value systems unless European models can compete on performance and win broader adoption.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.