OpenAI’s mathematical breakthrough has begun to raise questions. Scientists compared the published proof of the Navier-Stokes equations with its algorithmic Lean version and found discrepancies: in several key places, the machine version, verified on a computer, contains weaker statements. This is not a refutation yet, but the ground for doubt is already quite fertile

Image source: Thomas T /unsplash.com
The discrepancy doesn’t mean the evidence is wrong or that OpenAI got the problem wrong, mathematicians say, but it does call into question whether mathematical results generated by artificial intelligence models can always be relied upon.
“What you have to do with all these big proofs generated from a language model is that they have to be read by people, and that puts a huge extra burden on mathematicians,” said team leader Anders Hansen.
The process of finding discrepancies took the team about two weeks. In this case, mathematicians used ChatGPT hints. Let us remind you that it took OpenAI 88 hours to solve the problem.
Mathematicians pointed out a discrepancy in a part of the proof called Lemma 8.6. In a natural language proof, the equation in this part requires that a certain value be less than m + 4, where m is an integer, but in the Lean version it is less than m + 5, which is mathematically weaker because it allows for more possible solutions and is not equivalent to the first one.
Imagine that you are asked to solve the equation x + 3 = 6, the answer to which is x = 3. You can write a proof that x must be less than 4, and also that x must be less than 5. Both of these mathematical statements are absolutely true, but the latter allows for more possible answers for x, which makes it mathematically weaker.
According to Hansen, the mistranslation could occur because the AI strives to ensure that the computer code is completely self-consistent and does not produce errors. If, during the automatic formalization process, the AI encounters a part of the proof that does not compile, it will try to find a workaround, even if this means deviating from the proof written in natural language.
Kevin Buzzard of Imperial College London reported that it is possible to reliably formalize the theorem itself in Lean and then check whether the proof compiles. But this does not guarantee that the text proof in the PDF correctly conveys the logic. That is, the Lean code can be internally consistent, but not correspond to what is written in the article. “I am confident that the Navier-Stokes problem was solved correctly,” says Buzzard. — I am much less confident that the evidence described in the text is correct.”
OpenAI told the resource New Scientistthat she is aware of the inconsistency between the natural language and code proofs, and that this does not mean that neither is valid.
If you notice an error, select it with the mouse and press CTRL+ENTER.





