Skip to content
Breaking:

OpenAI’s Math Proofs Draw Scrutiny Over Transparency and Translation Gaps

Mathematicians highlight missing reasoning chains, unformalized proofs, and code discrepancies in the lab’s latest batch of solutions.

By The Company Wire3 min read
Share
OpenAI — OpenAI’s Math Proofs Draw Scrutiny Over Transparency and Translation Gaps
OpenAI — OpenAI’s Math Proofs Draw Scrutiny Over Transparency and Translation Gaps. Photo: TechCrunch AI.

OpenAI’s release of hundreds of claimed solutions to complex mathematical problems has drawn scrutiny from the research community for falling short of guidelines established by leading mathematicians, according to a report from TechCrunch AI (https://techcrunch.com/2026/10/08/openais-math-solutions-arent-meeting-the-fields-standards-yet/).

The frontier lab sought input from an advisory panel to prevent controversy following previous automated proof announcements. However, several core recommendations outlined in late September by the Advisory Group on Mathematics and Artificial Intelligence (AGMAI)—a nine-member panel hosted by Princeton University’s Institute for Advanced Studies—were not met in OpenAI's latest release.

A central recommendation from AGMAI urged AI labs to stop testing advanced mathematical challenges on closed, proprietary models. OpenAI's release explicitly noted that it evaluated proprietary models against open research problems. In addition, while the lab released its results promptly, it included chain-of-thought reasoning for only 10 of the 719 published manuscripts.

Academic guidelines also advised that proofs should be formalized into code when human understanding is lacking. Approximately 42% of the proofs published by OpenAI had not undergone formalization. Furthermore, OpenAI did not include machine-readable metadata linking natural language proofs with formal code artifacts, another measure requested by AGMAI.

The friction between natural language and automated verification was underscored in a study published by mathematicians at the University of Cambridge and King’s College London. AI models typically generate a natural language explanation before attempting to translate it into Lean, a programming language designed to verify mathematical proofs. The study identified at least two discrepancies between OpenAI’s natural language proof and its Lean code for a problem derived from the Navier-Stokes equations governing fluid behavior.

The study's authors warned that because of translation discrepancies, OpenAI's natural language and autoformalized Lean proofs should not be accepted without standard peer review. Prominent mathematician Terence Tao echoed concerns about automated solver workflows, writing on social media that AI prompters often lack the domain interest or understanding required to discuss results or interact with the field. Harvard University mathematics professor Melanie Wood told TechCrunch AI that when models output solutions without human understanding at release, "now the work begins."

Sources

  1. TechCrunch AI

Company: OpenAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.