
AI is making impressive progress in mathematics
Varacolaci/Alamy
OpenAI has dropped 722 mathematical papers proving and disproving a range of problems – a scale that academics say is as impressive as it is baffling.
For months, the ability of AI models to tackle mathematical problems has been increasing. As recently as 2019, AI models struggled to achieve a passing grade on GCSE mathematics papers. Last month, however, OpenAI solved a problem related to the Navier-Stokes equations of fluid dynamics, which was one of the toughest and most enduring puzzles in mathematics.
Having achieved impressive depth, OpenAI has seemingly now opted for breadth, releasing hundreds of mathematical discoveries in one fell swoop. The company didn’t name the AI made the discoveries, saying only that it was an “internal frontier model”.
Advertisement
Francis Johnson at University College London worked for many years on Wall’s D(2) problem – one of the hundreds of puzzles solved in OpenAI’s tranche of papers.
“I worked on this problem for 25 years. I produced two books on it. I, personally, gave up,” says Johnson. “I’m surprised that [AI] has done it quite so quickly, but I’m not surprised that it’s done it.”
Johnson recounts working with AI models in recent months and being “astonished with the sophistication and clarity of the analysis that they gave”.
“It puts us all in a very strange position,” says Johnson. “Let’s face it, the genie is out of the bottle now. We’re going to have to live with it. Human beings are supposed to be adaptable, so we’re going to have to adapt. I think for the moment, we just stand back and be astonished.”
But Kevin Buzzard at Imperial College London says we should exercise caution. He says the release included 30 papers that were relevant to his field of number theory, but only seven of those seemed impressive and only one was formally verified in Lean – a type of computer analysis that can prove mathematical results beyond reasonable doubt.
“Unfortunately, acceptance of these results by the community will take time, and journalists are going to have to wait while the mathematicians do their job,” says Buzzard. “The six unformalised results will have to wait until either an expert is motivated to read and check the text, or a Lean formalisation is produced.”
But while it is important not to get ahead of ourselves on the validity of the results, Buzzard says the general upwards trend in AI mathematics is astonishing. He says that, assuming the new results turn out to be mostly true, then we will get some kind of idea as to what the new normal is. “It has been a long time since there were humans who were experts in all of mathematics, but now we seem to have machines with this property,” he says.
Ben Allanach at the University of Cambridge says that AI companies have been hiring mathematicians and investing time and resources into research, and that impressive results have been forthcoming. “It’s [mathematical discoveries] seen as an intellectual trophy, an achievable intellectual trophy, by the companies,” he says.
It’s “confusing and disrupting”, says Allanach. “I’m certainly glad I’m not a pure mathematician because you’d be saying, ‘Well, what’s the point’.”
But Allanach is also critical of the scale of the released results and the lack of effort made to explain them. “Is it coming into the cadre of human knowledge?” he asks. “Typically, mathematicians like what they do and they like the process of solving puzzles, and just checking or understanding machines’ working is probably not as attractive.”
Unusually for mathematical research, the papers were published on GitHub, an online database more commonly used to share computer code. Earlier this week, the scientific pre-print server arXiv announced it was introducing upload limits to stem the tide of “low-value” submissions.
OpenAI has faced criticism in recent months over the way it has disclosed mathematical research. Terence Tao at the University of California, Los Angeles, said last month that the way in which discoveries are being made and “dumped” on the mathematical community may harm the field. In response, an independent Advisory Group on Mathematics and Artificial Intelligence was set up that would advise AI companies on best practice.
OpenAI has previously said that it would take this advice on board, but its most recent release seems to go against many of the group’s suggestions. For instance, the group recommends that the model’s name, the prompts used and the estimated cost of computation be released. It also suggests that, as far as possible, a proof released by an AI lab should be formalised. The group also suggests that companies disclose how many other problems of comparable difficulty their models tried and failed to solve, as well as how the problems were chosen.
Lindsay McCallum Rémy at OpenAI told New Scientist in an email announcing the results that “we’re continuing to explore other community-hosted alternatives for this release which meet the committee’s guidelines. For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding. We want to give the mathematical community time and space to assess this work.”