China's latest AI challenge isn't related to chips; it's the depletion of Chinese-language training data.
TL;DR: China is experiencing a significant shortage of AI data. Chinese content makes up only 1.3% of web content, while English constitutes 49%. Beijing aims to establish national datasets by 2028, but platforms like WeChat and Douyin do not share their data with external parties. Publishers are imposing bans on the use of their content for AI training.
China's AI development faces a new limitation, not related to chips. The nation is running low on high-quality Chinese-language training data. While US restrictions on advanced semiconductors have dominated discussions about China's AI abilities, experts within China are increasingly cautioning that data scarcity might be an equally significant obstacle, and unlike chips, there is no hardware alternative. According to internet tracker W3Techs, Chinese content represents merely 1.3% of global web material, whereas English accounts for nearly half, with Spanish at 6% and Japanese at 5%.
This issue is global but has a more severe impact on China. Epoch AI projects that the global supply of high-quality, publicly accessible text could be depleted within six years. OpenAI co-founder Andrej Karpathy has highlighted the potential for a “data wall” by the end of the decade. Chinese developers are already spending more per useful token than their Western counterparts since their models must operate with less native-language content. The digital environment in China exacerbates the data shortage: platforms such as WeChat and Douyin do not share information with third-party developers, forcing AI labs to rely on lower-quality resources for training.
In response, Beijing is treating data as crucial infrastructure. In June, the National Data Administration announced a plan to create validated AI training datasets nationwide by 2028, encompassing sectors like manufacturing, energy, healthcare, finance, agriculture, autonomous driving, and embodied AI. Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, stated, “Competition in the AI era involves not just models and computing power, but also high-quality data supply systems.” Tsinghua University computer scientist Sun Maosong has called on authorities to digitize historical records, ancient texts, scientific literature, and regional dialects.
However, not everyone is eager to digitize. Huaxia Publishing House recently included a disclaimer on a new translation: “Using the content of this book for artificial intelligence training is prohibited. Violators will be held legally accountable.” China is increasingly concerned about American AI models it cannot compete with, and the data shortage partly accounts for this disparity. The US has initiatives like Anthropic’s Project Panama, which involved purchasing and destroying millions of physical books for digitization. In contrast, China, with 1.3% of the web, has a publishing sector that is beginning to close off access.
Other articles
China's latest AI challenge isn't related to chips; it's the depletion of Chinese-language training data.
Chinese represents only 1.3% of the total content on the internet. Beijing is now regarding data as critical infrastructure and has a plan to develop national AI datasets by 2028.
