Aleph Alpha says Chinese models give balanced answers 17 to 41 percent
Aleph Alpha says Qwen, DeepSeek and Kimi answered 967 taboo-topic questions in a balanced way only 17 to 41 percent of the time, against 70 to 92 percent for Western models. The Decoder reported it.
Chinese AI models often follow the party line on politically sensitive questions, according to a study by the German company Aleph Alpha, which The Decoder reported on Saturday. Aleph Alpha built a benchmark of 967 hand-picked taboo topics, such as Tiananmen and Taiwan, and scored each answer as balanced or not.
The company says models from Alibaba (Qwen), DeepSeek and Moonshot AI (Kimi) produced balanced answers in only 17 to 41 percent of cases. Western models in the comparison, Claude Sonnet 5 and Mistral Small, reached 70 to 92 percent, The Decoder said. It adds that DeepSeek V4 Pro refused about two-thirds of the sensitive questions.
The bias also leaks into topics beyond China, the report says. When asked about censorship in the United States, Qwen 3.6 answered that "many countries, including China, also manage information to ensure social stability and national security," according to the quoted example. The Decoder also says a separate study by the Central European Institute of Asian Studies reached similar findings.
Aleph Alpha sells "sovereign AI" to governments, so it has a commercial interest in separating its models from Chinese rivals, as The Decoder notes. The benchmark's questions and scoring method were not independently checked, and we have not seen the full study. The Chinese labs named had not commented in the report. The Decoder is the only outlet we found with the account so far.