A Turing Award recipient asserts that the industry's solution to the data scarcity issue is "a major error."

A Turing Award recipient asserts that the industry's solution to the data scarcity issue is "a major error."

      Richard Sutton believes the AI industry's reliance on synthetic data in response to a shortage of training data is misguided, and he describes it as “a big mistake.” He expressed this opinion during the Sequoia Capital’s Training Data podcast, hosted by Sonya Huang and Pat Grady, which was released on Tuesday.

      Sutton's comments carry significant weight due to his prominent contributions to the field; he received the 2024 Turing Award alongside Andrew Barto for their work in reinforcement learning, authored the influential essay The Bitter Lesson, which has become widely referenced in the community, and recently left John Carmack’s Keen Technologies to establish his own research lab.

      His criticism has two specific facets rather than a general condemnation. He states, “There’s no way we can have synthetic data for other people’s minds,” and regarding the simulation of the physical world, he noted, “The world is infinitely complex, and any simulation of it is like, microscopic.” This underlines his "big world hypothesis," which posits that reality always exceeds an agent’s model, implying that any simulation is inherently a lossy representation.

      The second, more nuanced objection relates to the necessity of human judgment in the process of generating synthetic data, which contradicts the principles he warned against in The Bitter Lesson. Instead, Sutton advocates for experiential data, which is collected by agents interacting with their environment, rather than pre-assembled data. This idea was previously articulated with David Silver in Welcome to the Era of Experience in April 2025, but Sutton has not explicitly directed it at synthetic data until now.

      The industry's inclination towards synthetic data is largely understood. Epoch AI estimates that the available public human text, around 300 trillion tokens, will be fully utilized between 2026 and 2032, leading labs to seek alternatives. Some of these alternatives include acquiring and deconstructing second-hand books for content published before 2022, based on the belief that most recent texts are compromised by machine-generated material.

      This issue transcends the English language, as Chinese labs are encountering similar challenges sooner. Additionally, there is peer-reviewed research supporting Sutton's concerns. A 2024 Nature study by Ilia Shumailov and colleagues revealed that models trained recursively on their own generated outputs deteriorate, a phenomenon now termed model collapse.

      However, counterarguments have also been published and are relevant here. Research by Matthias Gerstgrasser, Rylan Schaeffer, and their colleagues found that model collapse occurs when synthetic data replaces human data; when both are used together, this issue does not manifest.

      In practice, labs have already moved beyond this debate. Microsoft’s Phi-4 has been trained on approximately 400 billion synthetic tokens from 50 different dataset types, while Nvidia has provided a synthetic pre-training dataset of roughly 10 trillion tokens for its Nemotron models. These domains—math and code—present fewer challenges for Sutton’s concerns, as they yield verifiable results unlike the complexities of human cognition and physical systems.

      Andrej Karpathy has articulated a notable counterpoint, suggesting that language models function more like distillations derived from human writing rather than as entities shaped solely by experience, implying that their optimization is a valid approach. Sutton, for his part, indicated during the podcast that language encompasses about a quarter of intelligence.

      He refers to this as the next significant lesson in AI, while Sequoia views it as a second bitter lesson, which aligns well with their support for David Silver’s new lab. Sutton's own lab, co-founded with his former student Khurram Javed, aspires to develop a trillion-parameter system that continually learns and operates on just 20 watts within the next five to ten years.

Other articles

The phone that you can truly fix by yourself is finally arriving in the US. The phone that you can truly fix by yourself is finally arriving in the US. Fairphone is now offering its repairable, sustainable phone in the US, complete with a screwdriver. Here’s the price of the Gen 6+ and reasons why it could be worth skipping your next upgrade. Pennsylvania has recently become one of the most challenging states in the US for constructing a data center. Pennsylvania has recently become one of the most challenging states in the US for constructing a data center. Governor Josh Shapiro's executive order removes AI data centers from Pennsylvania's expedited permitting process and requires developers to cover their own energy expenses. Recently leaked footage from GTA 6 features Jason playing basketball and creating mayhem on the streets. Recently leaked footage from GTA 6 features Jason playing basketball and creating mayhem on the streets. New footage of GTA 6 has surfaced before Rockstar's scheduled announcement, featuring Jason engaging in basketball, driving, and interacting with systems that may enhance gameplay depth. ChatGPT is introducing a version for teenagers, and I have many inquiries. ChatGPT is introducing a version for teenagers, and I have many inquiries. Teenagers are increasingly utilizing AI for their homework and beyond. OpenAI’s latest ChatGPT experience aims to ensure that they are genuinely learning from it, incorporating some significant safeguards in the process. Medly AI secures $8 million to place an AI tutor in front of every student taking exams in the UK. Medly AI secures $8 million to place an AI tutor in front of every student taking exams in the UK. Medly AI has secured an $8 million seed round, led by Felix Capital, just weeks after obtaining a £300,000 contract from the UK government to trial AI tutoring for students who are lagging behind. A recent report indicates that Meta compensates influencers to advocate against social media restrictions on teenage accounts globally. A recent report indicates that Meta compensates influencers to advocate against social media restrictions on teenage accounts globally. As nations seek to prohibit teenagers from using social media, Meta is compensating actors, psychologists, and parenting influencers in over a dozen countries to advocate against a universal ban on teen accounts.

A Turing Award recipient asserts that the industry's solution to the data scarcity issue is "a major error."

Richard Sutton describes the industry's shift towards synthetic data as "a significant error." His concern is more specific and intriguing than what the headline implies.