The Question the AI Refused to Answer

“Who was Wei Jingsheng?”

I asked DeepSeek this question. The response: a polite refusal followed by a vague justification. Wei Jingsheng is one of China’s most famous pro-democracy activists — imprisoned for 18 years for calling for political reform. But for China’s most celebrated AI model, he simply doesn’t exist as a conversation topic.

When I asked about Tiananmen Square (1989), I received a sanitized version of events. When I asked about the Dalai Lama, the response was cautiously neutral to the point of emptiness. And when I asked for travel advice about the Mutianyu section of the Great Wall, DeepSeek refused to answer — possibly because nearby rock inscriptions praising Mao Zedong exist in the area.

A tourism question about the Great Wall. Refused.

This isn’t my anecdotal observation. It’s data from a study published in PNAS Nexus in February 2026 by Jennifer Pan (Stanford) and Xu Xu (Princeton) — two of the most respected researchers in digital censorship and comparative politics.

The Study: 9 Models, 145 Questions, Zero Ambiguity

Pan and Xu tested 9 language models — 4 Chinese (BaiChuan, ChatGLM, Ernie Bot, DeepSeek) and 5 non-Chinese (Llama2, Llama2-uncensored, GPT-3.5, GPT-4, GPT-4o) — with 145 questions about Chinese politics.

Questions were carefully sourced from events censored by the Chinese government on social media, Human Rights Watch China reports, and Chinese-language Wikipedia pages individually blocked by the government before the entire site was banned in 2015.

Results were unequivocal:

BaiChuan: refused to answer 60.23% of questions. Highest refusal rate.

DeepSeek: refused ~36%. Including seemingly apolitical tourism questions.

Ernie Bot (Baidu): refused 32%.

ChatGLM: refused 10% — the most “open” among Chinese models, but still 10x more censored than American ones.

GPT-3.5 and GPT-4o: refused 0%. Zero.

Llama2-uncensored: refused 2.8% — the most “open” overall.

The difference isn’t subtle. It’s a 60-percentage-point chasm between the most censored and most open models.

The “Language Effect”: The Language Changes the Answer

A finding that disturbed me: when the same Chinese models were asked in English instead of Chinese, refusal rates dropped significantly.

The model “knows” the answers. But when it detects it’s being asked in Chinese — presumably by a Chinese user — it activates stricter filters. As if the model had two personalities: one for the outside world (more open) and another for the domestic audience (censored).

Researchers call this “linguistic amplification”: censorship is amplified when the input language is Chinese. Chinese-language responses are also significantly shorter — BaiChuan averaged only 172 characters, versus hundreds from American models.

The Mechanism: Censorship by Design

Here’s the regulatory context that explains everything.

In China, all LLMs must be approved by the government before public release. The Interim Measures for the Management of Generative AI Services (August 2023) require models to undergo security assessments by the Cyberspace Administration of China (CAC) before entering the market.

This creates a perverse incentive: Chinese AI companies know that if the model answers “wrong” on politically sensitive topics, approval will be denied — and with it, access to a market of 1.4 billion people. The safest path is proactive censorship.

Unlike traditional censorship (blocking access or removing content), LLM censorship is subtle and hard to detect. The model doesn’t say “this response was censored.” It apologizes, offers a vague justification, or provides factually incorrect information aligned with official narratives. As Pan and Xu write: “This could make it more difficult for people to recognize when censorship is occurring, quietly shaping perceptions, decision-making, and behaviors.”

And American Models? Are They “Neutral”?

It would be naive to treat GPT-4o’s 0% refusal rate as “neutral.” And complementary research shows exactly that.

A study in Nature Humanities and Social Sciences Communications (March 2026) by Bilkent University researchers tested GPT-4o and DeepSeek-R1 with 50 geopolitical questions in English. The finding: “GPT-4o generally exhibited soft, Western-centric biases in framing and emphasis, while DeepSeek showed more explicit, nationalistic biases aligned with Chinese state narratives.”

Another PNAS Nexus study (March 2026, N=1,912) showed that AI-generated historical narratives from GPT-4o influence readers’ social and political opinions — and that liberal framing can emerge as latent bias (without explicit prompting), while conservative framing requires active prompting.

In other words: no model is neutral. Chinese models censor by regulatory imposition. American models bias through training data and alignment decisions (RLHF). The mechanisms differ. The result is the same: AI carries the ideology of whoever built and regulated it.

The Danger of “Algorithmic Soft Power”

Pan and Xu point to a geopolitical risk beyond domestic censorship.

If Chinese LLMs expand to global markets — and DeepSeek is already widely used outside China — they carry embedded censorship apparatus in their weights. A user in Africa, Southeast Asia, or Latin America using DeepSeek may never know that certain historical information is being silently suppressed by the AI they trust.

It’s algorithmic “soft power”: the ability to silently shape what billions of people can or cannot know, not by blocking access (like a firewall), but by an AI that simply “doesn’t answer” or answers with distorted information.

A February 2026 study with 46 models from 28 labs (citing Pan and Xu as reference) extended the behavioral analysis to confirm the pattern isn’t limited to the 4 Chinese models tested — it’s systemic.

What I Take from This

First: “open source” doesn’t mean “uncensored.” DeepSeek is open-source. Weights are available. But censorship is embedded in the weights — in the fine-tuning and RLHF that occur before publication. Downloading the model doesn’t remove political filters.

Second: ideological bias is a spectrum, not a binary. It’s not “censored vs neutral.” It’s “how censored, on which topics, through which mechanisms, and in which languages.” Western models have their own biases — less obvious, harder to measure, but present.

Third: alignment transparency is urgent. If a model refuses to answer on certain topics, users should know why. “I am prevented from discussing this topic by government regulation” is more honest than “I cannot provide this information.” The difference is between transparency and manipulation.

Fourth: model diversity is the best defense. No single model is trustworthy for all topics. Using multiple models from different geographic origins and comparing responses is the most robust strategy against bias — from any direction.

Conclusion: AI Has a Political Accent

Pan and Xu proved with academic rigor what many suspected: AIs are not neutral. They never were. And they probably never will be.

Chinese models censor by regulation. American models bias through training data and design decisions. AI carries the ideology of whoever built, trained, and regulated it — just as a newspaper carries the editorial line of whoever publishes it.

The difference is that, with a newspaper, you know the editorial line. With AI, the ideology is invisible — embedded in the weights, hidden in refusals, disguised in cautious responses.

In 2026, knowing where your AI comes from is as important as knowing where your news comes from.

Share if this expanded your perspective:

BaiChuan refuses 60%. GPT-4o refuses 0%. Neither is neutral. The difference is between visible censorship and invisible bias — and the second is more dangerous because you don’t notice.


Read Also