OpenAI’s math solutions aren’t meeting the field’s standards yet
OpenAI published hundreds of claimed solutions to difficult math problems this week, a release the company said followed advice from an advisory group of mathematicians. Researchers and the advisory group say OpenAI fell short of key recommendations — notably around ensuring human understanding, providing formal verification, and avoiding testing proprietary models on open mathematical problems.

Why It Matters
The episode highlights a broader tension between rapid AI-driven result generation and the traditional mathematical norms of verification, transparency and community engagement. Gaps between natural-language explanations and machine-checked formalizations raise concerns about relying on models to autonomously produce dependable mathematical knowledge.
Key Facts
- Number of manuscripts released: 719
- Advisory group: Advisory Group on Mathematics and Artificial Intelligence (AGMAI), nine researchers, hosted by Princeton IAS
- AGMAI guideline release: End of September (guidelines published)
- AGMAI first request: Stop testing advanced mathematical problems on proprietary models
- Chain-of-thought disclosures: 10 of 719 manuscripts included the model's chain of thought (explanatory traces)
OpenAI this week published hundreds of solution manuscripts for challenging mathematical problems, saying it had consulted an advisory group of leading mathematicians. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton’s Institute for Advanced Study and composed of nine researchers, issued a set of recommendations in late September intended to guide frontier labs working on mathematical problems.
AGMAI’s guidance emphasized that newly claimed results should be accompanied by human-understandable explanations and, where necessary, formal verification. The group also urged labs to stop testing hard mathematical problems on proprietary models. OpenAI’s release followed some of these practices, such as rapid disclosure and providing information about model reasoning, but it did not meet others: only 10 of the 719 manuscripts included the models’ chain-of-thought traces, and AGMAI’s requested machine-readable metadata linking natural-language and formal artifacts was not provided.
Formal verification was another focal point. AGMAI recommended formalizing proofs that are not immediately comprehensible to humans; according to TechCrunch’s reporting, 42% of OpenAI’s released proofs had not undergone that formalization process. Independent researchers from the University of Cambridge and King’s College London documented at least two discrepancies between the natural-language explanations and the Lean formal code for a problem derived from the Navier–Stokes equations, showing that automatic translation from prose to formal proof can produce mismatches.
Mathematicians quoted in coverage stressed that when humans publish new results they also take responsibility for explaining, defending, and integrating them into the broader literature through talks and peer review. Critics including Terence Tao have warned that AI-driven workflows can produce solutions without corresponding human understanding or community engagement — leaving significant work for mathematicians to validate, interpret, and disseminate any claimed advances.
Keep Reading
Asos confirms breach of customer data after hackers send rogue app notification

New York alleges TikTok gave teens, children a placebo safety feature instead of a real one

Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

Anthropic changes usage policy to ban model abuse and election interference
Original source: TechCrunch