Urgent.News

the world's headlines, one feed

AI

China faces new AI bottleneck as it runs out of Chinese-language training data

China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data. While the US chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to…

China faces new AI bottleneck as it runs out of Chinese-language training data

China's pursuit of advanced AI models has reached a critical juncture, with a burgeoning concern over a potential data shortage. While the US holds sway over cutting-edge computing chips, Chinese AI experts warn that the lack of high-quality training data could become the next significant hurdle to the country's technological aspirations. This challenge confronts AI behemoths on both sides of the Pacific, with some US firms employing aggressive tactics to maintain their edge.

A US-based research institute, Epoch AI, predicts that the supply of high-quality, publicly available human-generated text may dwindle to nil within the next six years. Similarly, OpenAI co-founder Andrej Karpathy has cautioned of an impending "data wall" by the end of this decade, beyond which AI model performance could stagnate unless fresh, reliable data is supplied.

In response, major US labs are investing heavily in acquiring offline human knowledge, sparking an ethical debate. Amazon-backed Anthropic, for instance, spent millions acquiring millions of physical books, scanning and digitizing every page before discarding the originals, as reported by The Washington Post in January. The move drew criticism from authors, archivists, and preservationists who accused tech companies of treating human heritage as expendable raw material.

For China, the looming data wall presents a unique threat. While English dominates nearly half of the global web, Chinese accounts for only 1.3 percent, trailing behind languages like Spanish, German, and Japanese. To address this, Beijing is implementing a comprehensive nationwide plan to boost the supply, circulation, and commercialization of high-quality AI training data by 2028.

The plan aims to create an expansive ecosystem of validated datasets across key sectors such as manufacturing, energy, healthcare, finance, and agriculture, as well as frontier areas like embodied AI, autonomous driving, and low-altitude aviation.

Yu Xiaohui, president of the China Academy of Information and Communications Technology, emphasized that the competition in the AI era extends beyond models and computing power, emphasizing the importance of high-quality data supply systems. To bridge the gap, scholars and industry experts advocate for a coordinated national effort to enhance Chinese-language corpora – structured collections of text used for training and evaluating AI models.

Rather than relying solely on web scraping, Chinese computer scientists argue that the country must systematically digitize vast, untapped offline assets. They propose scanning historical archives, local gazetteers, ancient manuscripts, and scientific literature, as well as expanding corpora to include dictionaries, audiovisual content, and regional dialects.

Furthermore, Beijing is encouraging the tech industry to explore simulation and synthetic data generation as additional sources of training data, alongside traditional methods like magazines and programming code.

However, the drive to digitize China's offline heritage faces resistance from publishers and copyright holders, who are drawing clear boundaries around their works. Recent warnings printed on the copyright page of newly translated ancient Daoist texts serve as a stark reminder of the challenges ahead. Huaxia Publishing House, which added a clause prohibiting AI training using its content, asserted that the restriction was imposed at the request of licensing partners for imported titles. Detecting and enforcing AI infringement rights remains a formidable challenge.

Written by urgent.news from SCMP Tech's reporting — not their text. Machine-written; read the original for the full account.

This story

This is one outlet's version. Read the fullest account.

Read the original at scmp.com →

More in AI

AI Won't Kill Your AI Won't Kill Your Motivation. But Mediocrity Might.

My last post did something I didn't expect. I wrote about AI killing my motivation, thinking it might get a few quiet nods from people feeling the same way.

  • AI automates routine tasks, raising developer motivation concerns
  • Skilled developers focus on complex problems, not routine tasks
  • AI should assist decision-making, not replace developer responsibilities

2.Self-Hosted AI: n8n + Ollama, local AI workflows on your Mac

If you want AI agents running on your own machine, with your own models, and no data leaving your computer, this is the article :). This is part three of the series.

  • Install Docker on Mac for local AI setup
  • Set up n8n with Self-hosted AI Starter Kit
  • Connect Ollama to n8n using credentials