OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt—Mathematicians Want Receipts
OpenAI posted 722 mathematics manuscripts on GitHub that it says were generated by an unreleased internal AI model; the company claims most outputs came from a single prompt given to a single agent. Only 162 of the manuscripts include a computer-checked main result formalized in Lean, and OpenAI cautioned that some unformalized results may contain errors.

Why It Matters
If verified, the release would represent a substantial advance in AI-assisted mathematical research, but experts say the claim of one-shot, single-agent generation is unconfirmed without the model, prompts, or broader reproducibility evidence. The mix of unformalized results and limited transparency has prompted calls from mathematicians and institutions for more disclosure and verifiable artifacts.
Key Facts
- Manuscripts published: 722
- Result families: 372
- Problems posed to model: Roughly 4,000
- Lean-formalized main results: 162 (about 22%)
- Average compute per result (OpenAI): Equivalent of ~3 hours of ChatGPT Pro thinking compute
OpenAI published a collection of 722 AI-generated mathematics manuscripts on GitHub, attributing them to an internal model it has not released. The company said the bulk of the outputs derived from a single prompt issued to a single AI agent, although it acknowledged some items may have required multiple attempts. The documents are organized into 372 "families," where each family can include a main theorem plus related arguments, consequences or alternate proofs, so the manuscript count does not equal distinct solved problems.
OpenAI reported that it presented the model with about 4,000 problems and kept roughly 722 outputs it judged significant. The firm provided average compute figures — roughly three hours of ChatGPT Pro equivalent per result — and released abridged reasoning summaries for 10 results, but it has not published the model itself or the prompts used to generate the manuscripts. An advisory group at the Institute for Advanced Study had recommended disclosing the model name, prompts, summarized chains of thought, and per-result compute metrics; OpenAI has not complied with those specific recommendations and said it is still working on a responsible release.
A formalization catalog in the repository shows that 162 of the 722 papers include a main result checked in Lean, the proof assistant that mechanically verifies each logical step. OpenAI cautioned that many manuscripts are unformalized and that "some of the unformalized results could have issues." Observers note that a Lean check verifies that a proof follows from a statement as written in Lean, but does not by itself confirm that the stated theorem accurately matches the original problem, that the theorem is new, or that the mathematics is correct in the broader scholarly sense.
Responses from the mathematics community have been mixed. Some researchers urged caution, saying claims of one-shot problem solving by a single agent are unverified until the model and prompts are released and independent teams can reproduce the results. The Institute for Advanced Study emphasized the ongoing importance of human understanding and responsibility in mathematical scholarship. Other academics praised the work as an important advance while noting most items appear to be significant but incremental contributions within existing research programs rather than decisive breakthroughs; the collection includes one claim labeled the Quasi-Riemann Hypothesis among the set. The repository currently has Issues turned off and has not accepted pull requests, which some mathematicians criticized given the volume of material released.
Keep Reading
Microsoft and Nvidia are teaming up on a supercharged AI laptop

Google Lets You Build a Video Game by Typing a Description

Hackers Used AI Agents to Raid a Megachurch's Database, Exposing 850,000 Members
