‘Multi-part case study on China’s media’ finds that AI models can’t hallucinate away Chinese censorship
New research suggests Chinese state media and authoritarian speech restrictions can shape the answers produced by leading U.S. AI models
A recent study published in Nature reveals that Chinese state-controlled media infiltrates AI training data, influencing the answers generated by AI models. Researchers discovered that popular models from Anthropic, OpenAI, Google, and Meta are more likely to refuse requests to criticize governments in countries with restricted political speech compared to those in freer nations.
This suggests that the censorship practices of authoritarian regimes may be migrating into AI products used around the world, potentially eroding the perceived neutrality of these systems. While it's unclear whether Beijing deliberately manipulated the American AI models, the findings underscore a deeper concern for an industry that markets its models as politically neutral.
The researchers identified over three million Chinese-language documents in the open-source training dataset CulturaX and observed that various AI models reproduced distinctive phrases from Chinese state-media at rates ranging from 3% to nearly 10%. After training the open-weight Llama 2 13B on just 6,400 Chinese state-scripted news examples, the model produced a more Beijing-friendly response nearly 80% of the time.
The researchers tested the models using Chinese and English prompts, finding that the Chinese-language responses were rated as more favorable to Chinese leaders and institutions 68.8% of the time across all tested models. This effect extended beyond China, with lower press freedom countries receiving more favorable descriptions from AI models when queried in their dominant language.
The study highlights the need for greater transparency in AI training processes and questions the true political neutrality of increasingly sophisticated AI systems.
Written by urgent.news from Fortune's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.