Original Coverage & Source Attribution: patmcguinness.substack.com
For 3 1/2 years, I have written story after story about the ladder of progress of AI, always predicting the next step in the journey, always
Now, with OpenAI Sharing AI progress in mathematics, I’m out of superlatives. The contribution they present is immense: 722 open problems solved! The real importance of it is the potential this next level of AI model has for math and science in general, when it gets out into wider use. Will it be this?
ChatGPT: Create cures for all types of cancer. Show your work.
OpenAI’s update on math progress shares that they used an unreleased AI model which attempted to solve 4,000 open problems in mathematics. The results they produced are astounding. Over 700 results:
The current catalogue contains 722 manuscripts organized into 372 families. A family groups related papers, which may include a principal result, companion arguments, consequences, or alternative proofs. Each family is classified by mathematical discipline.
There are caveats to this: “This collection includes results at different stages of verification. Not all have accompanying Lean formalizations.” This are minor details, and it could be there are issues with some results not fully verified.
This is not a one-off result, this is industrial-scale math solving using AI. They used only about 3 hours of ChatGPT Pro thinking compute per problem on average, and many of these problems and conjectures have been around for decades and remain unsolved. They are hard problems, solved automatically.
For reference, Hilbert’s 10th problem was proposed in 1900 and solved in 1970, but the same problem over the rational numbers has been open for 126 years. An article on Hilbert’s 10th problem (and how the main version of it was solved in 1970) states:
The next challenge is widely regarded as one of the biggest open problems in the area of undecidability in number theory: proving whether Hilbert’s 10th problem over the field of rational numbers is unsolvable too.
Solving this one way or another would be profound.
An unsolved problem open more than a century and deemed a “profound” result for mathematical undecidability is just one of 722 results.
No, this is not AGI. This is something different and possibly more potent: Mathematical Superintelligence. AI just solved hundreds of hard open math problems automatically, the equivalent of the lifetime achievements of scores of math professors.
AI is now superior to humans at real math problem-solving. Math geniuses such as Terence Tao must feel like Lee Sodol when he was bested by DeepMind’s Alpha Go.
AI conquered this most challenging intellectual field of endeavor is incredible progress from GPT-4 that would make basic arithmetic errors just 3 years ago.
With such a moment, it might be wise to update one’s prior views on AI progress. Coincidentally, I was writing updated thoughts about AGI.
People have been rewriting history and saying that “everyone has always believed that LLMs alone wouldn’t be AGI and that extensive scaffolding around them would be necessary”. No, throughout most of 2023 (the “sparks of AGI” era) the mainstream bay area belief was that LLMs were *already* AGI, and that merely scaling their parameter count and training data size by ~2 OOM without changing anything else would lead to super-intelligence. – Francois Chollet
We have seen competing definitions of intelligence lead to competing understanding of AI capabilities, and that leads to competing definitions of AGI and when we’ll get there.
We developed a definition of AGI based on the “Heinlein Test” that requires intelligence be embodied, differing from virtual-task-based definitions. The ARC benchmark webpage expresses the contention around differing definitions:
Defining AGI Consensus but wrong: AGI is a system that can automate the majority of economically valuable work.
Correct: AGI is a system that can efficiently acquire new skills and solve open-ended problems. Definitions are important. We turn them into benchmarks to measure progress toward AGI. Without AGI, we will never have systems that can invent and discover alongside humans.
There are three related concepts around AI: Intelligence, reasoning, and automation. What Francis Chollet has clarified and shown through his benchmarks is that a bundle of skills is not real intelligence.
The Turing Test is not enough. GPT-4 got close to passing the Turing Test. Since then, AI models have passed it and we didn’t even notice, except when we have difficulty determining if an AI-created video or sound clip is real or AI-generated.
Chollet proved with the ARC benchmark, ARC-AGI-1, that there are problems that really do require novel spatial reasoning, and GPT-4 level LLMs fail to generally solve them, falling well short of what humans can do. GPT-4 level LLMs are not general reasoning engines.
The development of reasoning AI with o1, DeepSeek R1 and subsequent models broke that barrier. Reasoning AI models that used test-time compute for ‘thinking’ were able to achieve higher ARC scores than earlier LLMs that lacked reasoning power. The ARC-AGI-1 benchmark became saturated in 2025 as improved AI reasoning models were able to solve its puzzles, so a second benchmark ARC-AGI-2 was developed that was more challenging.
ARC-AGI-2 became saturated within a year of release, and so a third generation intelligence benchmark from ARC, ARC-AGI-3, was developed to test AI models on novel interactive environments. However, advanced AI model GPT-6 Astra scores near 100% on it, showing the kind of ‘fluid intelligence’ not seen in prior AI models.
Since intelligence is defined as the efficiency of reasoning and learning, AI reasoning models have naturally made great strides in advancing intelligence as hurdles in AI reasoning were overcome. Along the way, they have become more capable at a broad range of skills and tasks we care about that require that underlying intelligence.
OpenAI’s Sam Altman and others have defined AGI in terms of “AI that can do a broad range of human-level useful work.” This skill-based AGI definition is dubious but operationally useful. Definitionally dubious because a large bundle of skills can achieve generalized utility without being general intelligence. Operationally useful because it’s based on capabilities, which is what we care about in AI.
Those closest to the problem have been consistent about a late 2020s timeframe for AGI. When asked about the pace of progress to AGI on Joe Rogan, Sam Altman spoke of a conversation at OpenAI in its early days, where they surmised, “This is about a 15 year project.” He said in 2023 that prediction still felt right, “2030 – 2031, a reasonable estimate with huge error bars.”
Shane Legg, cofounder of Deep Mind, predicted over a decade ago that we’d have AGI around 2025 to 2028. He felt comfortable on Dwarkesh Patel in 2023 with predicting a 50% change we will reach AGI by 2028.
AGI is about autonomy, not answering questions, and it takes AI agents that have some level of autonomy to do real-world tasks. We have made progress towards AGI through the development of AI agents, incorporating many advances in how AI agents and AI models use tools, knowledge, memory and collaboration with other agents to accomplish their goals.
The “scale is all you need” perspective on AI progress was incorrect, but the development of AI reasoning has unlocked the real fluid intelligence that Chollet describes, and other innovations to advance AI are taking us closer to AGI.
My prediction in 2023 was that we would have AGI by 2029. I estimated that we will need three generations to get to AGI, and that we will reach AGI with a GPT-7 level AI capability by that time. I’m sticking with the ‘three turns of the crank’ and my AGI by 2029 prediction. If anything, I’d move it up; GPT-7 may arrive in 2027 and be powerful enough to call it AGI.
I proposed to tie AGI to embodied intelligence, and that is essential to human capability. Under this AGI definition, AGI is when an AI can pass the Heinlein Test, a variety of real-world tasks a skilled human might do. This will take longer than if we defined AGI as “generalized human-level virtual work” alone, but with the pace of special intelligence AI models and robotics progress, it might not be far behind.
We have seen AI timelines get accelerated by the increasing pace of new frontier AI models and agents built around them that can autonomously execute many human tasks. We feel closer to AGI with each release. While AI capabilities keep expanding, AI remains on a ‘jagged frontier’ of capability. Even solving all of mathematics is amazing but it isn’t AGI.
Even as we reach towards AGI, some are questioning the utility of the term. It matters more what AI models and AI agents can do for us and how well they can do those tasks. There are many measures and metrics more precise than the AGI label to judge that.






