{
  "id": 289865,
  "title": "China faces new AI bottleneck as it runs out of Chinese-language training data",
  "url": "https://urgent.news/2026/08/08/china-faces-new-ai-bottleneck-as-it-runs-out-of-chinese-language",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-08T02:00:23.000Z",
  "source": {
    "name": "South China Morning Post",
    "slug": "south-china-morning-post",
    "url": "https://www.scmp.com/tech/tech-trends/article/3363318/china-faces-new-ai-bottleneck-it-runs-out-chinese-language-training-data"
  },
  "original_language": "en",
  "account": "China's pursuit of advanced AI models has hit a new obstacle: a shortage of high-quality training data. While the US has struggled with limited access to cutting-edge chips, Chinese AI experts stress that running out of sufficient data could become the next significant hurdle, one that hardware solutions cannot easily overcome. This challenge is faced by AI companies on both sides of the Pacific, with some US firms resorting to aggressive measures to stay ahead.\n\nAccording to US-based research institute Epoch AI, the global supply of high-quality, publicly available human-generated text may be depleted within six years. OpenAI co-founder Andrej Karpathy has warned of a looming \"data wall\" by the end of this decade, after which model performance would plateau unless fresh information is provided. Major US labs are investing heavily in acquiring offline knowledge, sparking ethical debates. Amazon-backed Anthropic, for example, spent millions buying and destroying rare books to create digital versions, causing controversy among authors, archivists, and preservationists.\n\nChina faces a unique dilemma due to its limited Chinese-language data, which makes up only 1.3% of the global web's content, compared to languages like Spanish (6%), German (5.9%), and Japanese (5%). In response, Beijing has implemented a nationwide plan to boost the supply, circulation, and commercialization of high-quality AI training data by 2028. The goal is to establish an extensive ecosystem of validated datasets across various sectors, including manufacturing, energy, healthcare, finance, and agriculture.\n\nExperts and scholars are advocating for a coordinated national effort to increase Chinese-language corpora – structured text collections used to train and evaluate AI models. Computer scientist Sun Maosong from Tsinghua University emphasizes the importance of digitizing vast, untapped offline assets, including historical archives, gazetteers, manuscripts, scientific literature, and regional dialects. He also suggests broadening Chinese-language corpora to include dictionaries, audiovisual content, and other resources.\n\nHowever, efforts to digitize China's offline heritage are met with resistance from publishers and copyright holders, who seek to protect their intellectual property. For instance, a new translation of ancient Daoist texts published by Huaxia Publishing House includes a clause prohibiting its use for AI training, citing legal responsibility for violators. While the publishing industry aims to safeguard their rights, China's government continues to prioritize data as a strategic resource for advancing AI technology.",
  "summary": "China’s AI race is facing a new bottleneck as the nation runs out of high-quality, Chinese-language training data, according to recent reports. While the US hardware shortage has dominated headlines, Chinese AI experts warn that a lack of data could prove to be the next major hurdle for the nation’s tech ambitions. This challenge is not unique to China, as US companies are also grappling with the \"data wall,\" where model capabilities could plateau due to a lack of fresh, reliable information. To tackle this issue, China is taking aggressive measures to treat data as a strategic asset, with the National Data Administration unveiling a nationwide plan to boost the supply, circulation, and commercialization of high-quality AI training datasets by 2028.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "SCMP Tech",
        "title": "China faces new AI bottleneck as it runs out of Chinese-language training data",
        "url": "https://urgent.news/2026/08/08/china-faces-new-ai-bottleneck-as-it-runs-out-of-chinese-language-290690",
        "published": "2026-08-08T02:00:23.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}