OpenAI keeps testing proprietary models on math despite its advisers’ request

- OpenAI posted 719 math manuscripts from an unreleased internal model on GitHub on October 6.
- Its advisory group had asked labs on September 29 to stop testing advanced math problems on proprietary models.
- A paper found OpenAI’s Lean code for its Navier-Stokes proof proves weaker statements than the written version.
On October 6, OpenAI pushed 719 math papers from an unreleased internal model to GitHub.
One week earlier, its own advisory group asked labs to stop testing advanced math problems on proprietary models.
Over 600 survey responses shaped the IAS panel’s September 29 guidance
The Advisory Group on Mathematics and Artificial Intelligence is based at the Institute for Advanced Study in Princeton, New Jersey. On September 29, it published recommendations grounded in more than 600 survey responses.
“We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models,” the group wrote.
The same document warns that private internal models risk creating a two-tier system where labs outrun the rest of the field.
As part of model development, OpenAI says in its repository, the company tests its models on open research problems. It extended those tests after its existing math benchmarks saturated.
OpenAI said it conferred with the group and drew on its recommendations for the release. In a statement released October 6, the group said its advisory role is not “an endorsement of the process by which OpenAI obtained them.”
The group entrusted the verdict on compliance to the wider mathematical community.
The Navier-Stokes Lean code proves weaker claims than the written paper
OpenAI grouped the 719 papers into 372 math families, each built around a main result and its related proofs. It gave the proprietary model about 4,000 problems, and each result took about three hours of ChatGPT Pro thinking compute on average.
The group asked AI labs to publish the model name, prompts, a summarized chain of thought, time taken, and compute cost for each result.
OpenAI published 10 reasoning summaries, and about 42% of top-line results carry Lean formalizations.
Lean is a programming language that can be used to check a proof by a computer. If the code compiles, it doesn’t mean that the code and written argument are the same. It only means that the formal statement is true.
Alexander Bastounis, Fabian Circelli, and Anders C. Hansen examined that gap in a paper posted on October 6. They find that the written Navier-Stokes paper from OpenAI makes stronger claims than what is proven by its Lean code, and one estimate changes the order of the input derivatives.
The authors detected similar mistranslations in OpenAI’s Euler proof. OpenAI’s written proof and other autoformalized Lean proofs “should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to,” they wrote.
The Navier-Stokes result was announced by OpenAI on September 8. According to Cryptopolitan, NYU mathematician Tristan Buckmaster then disputed its description of the minimal human input involved in the work.
The company pledged cash for workshops, conferences, and special programs on major AI results.
On October 6, Terence Tao wrote that AI prompters often solve problems and then lose interest in the field. Many cannot answer questions or give talks on the result, he wrote. His post did not name OpenAI.
The company has said that its advisers have no say over the pace of its research, Cryptopolitan reported in September.
If you're reading this, you’re already ahead. Stay there with our newsletter.
FAQs
What did OpenAI release on October 6, 2026?
719 math manuscripts from an unreleased internal model, with Lean formalizations for about 42% of top-line results.
What did the advisory group ask OpenAI to do?
Stop testing advanced math problems on proprietary models, and publish each result's prompts and summarized chain of thought.
Why are mathematicians worried about the Navier-Stokes proof?
A paper found OpenAI's Lean code proves weaker statements than its written proof, and changes a derivative order in one estimate.
Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses
Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.
















