Urgent.News

What's breaking now, across thousands of outlets.

AI

China faces new AI bottleneck as it runs out of Chinese-language training data

China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data. While the US chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to…

China faces new AI bottleneck as it runs out of Chinese-language training data

China's pursuit of advanced AI models has hit a new obstacle: a shortage of high-quality training data. While the US has struggled with limited access to cutting-edge chips, Chinese AI experts stress that running out of sufficient data could become the next significant hurdle, one that hardware solutions cannot easily overcome. This challenge is faced by AI companies on both sides of the Pacific, with some US firms resorting to aggressive measures to stay ahead.

According to US-based research institute Epoch AI, the global supply of high-quality, publicly available human-generated text may be depleted within six years. OpenAI co-founder Andrej Karpathy has warned of a looming "data wall" by the end of this decade, after which model performance would plateau unless fresh information is provided.

Major US labs are investing heavily in acquiring offline knowledge, sparking ethical debates. Amazon-backed Anthropic, for example, spent millions buying and destroying rare books to create digital versions, causing controversy among authors, archivists, and preservationists.

China faces a unique dilemma due to its limited Chinese-language data, which makes up only 1.3% of the global web's content, compared to languages like Spanish (6%), German (5.9%), and Japanese (5%). In response, Beijing has implemented a nationwide plan to boost the supply, circulation, and commercialization of high-quality AI training data by 2028. The goal is to establish an extensive ecosystem of validated datasets across various sectors, including manufacturing, energy, healthcare, finance, and agriculture.

Experts and scholars are advocating for a coordinated national effort to increase Chinese-language corpora – structured text collections used to train and evaluate AI models. Computer scientist Sun Maosong from Tsinghua University emphasizes the importance of digitizing vast, untapped offline assets, including historical archives, gazetteers, manuscripts, scientific literature, and regional dialects. He also suggests broadening Chinese-language corpora to include dictionaries, audiovisual content, and other resources.

However, efforts to digitize China's offline heritage are met with resistance from publishers and copyright holders, who seek to protect their intellectual property. For instance, a new translation of ancient Daoist texts published by Huaxia Publishing House includes a clause prohibiting its use for AI training, citing legal responsibility for violators. While the publishing industry aims to safeguard their rights, China's government continues to prioritize data as a strategic resource for advancing AI technology.

Written by urgent.news from South China Morning Post's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at scmp.com →

More in AI

More from Saturday 8 August →