Urgent.News

What's breaking now, across thousands of outlets.

AI

Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be…

Autoformalisation, a process used to verify mathematical texts including those generated by AI, faces significant challenges that undermine confidence in the original natural language (NL) arguments. The AI system translates the NL text into a formal language like Lean, which allows for mechanical verification of the argument. However, the accuracy of this process is compromised due to the difficulties in translating mathematical NL text semantically faithfully.

The problem of resolving ambiguities in mathematical NL text is arbitrarily high in the Solvability Complexity Index (SCI) hierarchy or arithmetical hierarchy, with SCI equal to infinity. This means that providing semantically faithful AI autoformalisation is more complex than any computational problem, including the Halting problem, which also has SCI equal to 1.

To illustrate the impact of this issue, several examples of AI mistranslations of NL statements and proofs into Lean were provided, leading to mismatches between NL proofs and their Lean verification. One such example is OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. The formalised Lean proof does not correspond to the original NL proof, demonstrating that the process of AI autoformalisation does not guarantee the correctness of natural language proofs.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at arxiv.org →

More in AI

More from Thursday 8 October →