The AI mathematical advancements that are frequently referenced are counterexamples rather than proofs.
A proof and a counterexample represent different achievements. A proof establishes that something is universally true, while a counterexample demonstrates that a statement claimed to be universally true fails by presenting a single instance where it does not hold. Both provide answers to questions, but they require different efforts from those who discover them: one necessitates a comprehensive argument that encompasses all scenarios, while the other only requires one example and the freedom to keep searching until it is found.
This distinction was the focus of a blog post by Cambridge mathematician Timothy Gowers on 12 August. Gowers, who won the Fields Medal in 1998, has engaged with the literature in a way that most commentators on AI and mathematics have not.
He is not dismissive of the advancements in AI. He describes the achievements as “extraordinarily impressive” and states that models can also prove complex concepts. "LLMs excel not only at finding counterexamples but can also discover proofs for intricate statements," he asserts.
Shared features of significant results
OpenAI recently reported the resolution of ten open problems in mathematics and theoretical computer science. The two main problems highlighted included the construction of a non-sofic group, which Gowers considers “one of the most significant unresolved issues in group theory,” and a result indicating that a multicolour Ramsey number increases superexponentially, which Gowers admits he did not anticipate seeing solved in his lifetime.
However, he notes an interesting point that received little attention: most of the prominent results from LLMs have emerged as counterexamples rather than proofs. He lists the above two results, along with the Jacobian conjecture and the unit distance conjecture. His third observation clarifies that while models can effectively prove universal statements, the most noteworthy results they have demonstrated do not align with the most significant things they have disproven.
Reevaluating two results
A counterexample is characterized by disproving something that many believed to be true. By this definition, Gowers reclassifies two notable results, including one that challenges the validity of its own documentation.
Regarding the non-sofic group, he points out that multiple construction methods had been documented previously, and he doubts that many experts firmly believed that all groups were sofic. Therefore, it is more accurately viewed as the first example of a non-sofic group rather than as a counterexample. He directly addresses the tension stemming from OpenAI labeling that section of their paper, “A counterexample to the soficity conjecture.”
The Ramsey result receives similar scrutiny, where Gowers notes that while many anticipated an exponential bound, he himself had worked on an equivalent formulation years earlier, which aligned with the eventual discovery. Thus, for him, it confirmed a mild expectation rather than refuted a previously held belief.
Why machines excel at examples
Gowers outlines eight strategies mathematicians use to search for examples: employing standard examples, constructing one from known components, leaving aspects undefined to adjust as needed, attempting to prove the opposite and interpreting what fails, making guesses and revisions, building the object incrementally, selecting one at random, and choosing a generic one.
Four of these strategies are advantageous for machines because they leverage broad knowledge and the ability to conduct numerous attempts. The methods of checking existing examples, building stepwise, using probabilistic reasoning, and selecting general examples all benefit from large volumes of trials. The other three methods require more complex judgment calls about the viability of the current approach.
This is where Gowers identifies a gap, which he calls a "nose," referring to the intuitive sense of when one is making progress and when to cease pursuing a particular line of inquiry. This intuition allows a human to effectively eliminate unproductive paths that a computer could not exhaust.
Evidence of the gap
His assertions about the gap are somewhat anecdotal, as he admits, based on his experiences with the GPT-5.6 Pro on open problems, which often result in strategies that seem promising yet ultimately fail scrutiny. He comments on a recognizable pattern in responses from the model—it often claims to have narrowed down to a more precise question “which sounds very promising until it has happened five times without any apparent progress.”
Similarly, he notes that experts respond to genuine successes in a predictable manner: initially expressing amazement, followed by a closer examination that reveals an approach uninspired and accessible to a suitably knowledgeable human with minimal hints.
He posits that the lack of an emergent "nose" is primarily due to published mathematics obscuring the search process. Models encounter “tidied up proofs that conceal their discoverers’ thought processes,” leaving little record in training data of rejected paths or reasons behind them. Additionally, a system fast enough to attempt everything lacks motivation to refine its pruning techniques.
The challenge he proposes
Gowers suggests a measurable standard, which is more specific than most commentary. He would consider a model’s achievement significant when it yields a proof as unexpected as the 2016 cap-set solution, which surpassed prior bounds with a methodology
Other articles
The AI mathematical advancements that are frequently referenced are counterexamples rather than proofs.
Fields medalist Timothy Gowers asserts that the renowned AI mathematics findings surfaced as counterexamples rather than proofs. In fact, one of them is not even that.
