Mathematicians want proof OpenAI didn’t use their work

Mathematician Andreas Thom has publicly accused OpenAI of lacking transparency about whether interactions with its ChatGPT chatbot contributed to the company's recent mathematical advances. Thom says his and colleagues' conversations with the model may have influenced a result OpenAI announced that built on work by Thom and Gábor Kun, and he is calling on the company to disclose relevant training data and procedures.

By AI NewsroomPublished 29 minutes agoUpdated 29 minutes ago0 views
Mathematicians want proof OpenAI didn’t use their work

Why It Matters

If proprietary AI models used unpublished or nonpublic interactions from researchers to improve performance, that would raise ethical and intellectual-property concerns and could chill open research collaboration. The dispute comes as OpenAI claims breakthroughs on high-profile mathematical problems, heightening the stakes around data provenance and credit.

Key Facts

  • Whistleblower: Andreas Thom, mathematician, raised concerns in posts on Mastodon.
  • Topic area: One of OpenAI's announced results involved non-sofic groups, a field connected to Thom's expertise.
  • Acknowledgement: OpenAI acknowledged its non-sofic groups result relied heavily on prior work by Thom and Gábor Kun and later amended its writeup.
  • Communications: Thom emailed OpenAI researchers Sébastien Bubeck and Mark Sellke asking whether his ChatGPT interactions were part of training data or accessible to the model's reasoning.
  • OpenAI's stance: In its Navier-Stokes blog post OpenAI denied accessing specific user data to solve the problem but said it could not rule out that de-identified usage-derived data helped improve models.

A second mathematician has publicly challenged OpenAI over whether the company used interactions with its ChatGPT system to train models that later produced notable mathematical results. Andreas Thom said in a series of Mastodon posts that conversations he and colleagues had with the chatbot before OpenAI’s announcement may have influenced a result the company touted, and he asked for clarity about how user interactions feed into training. OpenAI’s widely publicized package of mathematical findings included a result on non-sofic groups, an area closely related to Thom’s research. The company acknowledged that its non-sofic groups result drew heavily on earlier work by Thom and fellow mathematician Gábor Kun, and subsequently revised its public writeup after criticism in the mathematical community. Thom says he reached out directly to OpenAI researchers Sébastien Bubeck and Mark Sellke to ask whether his conversations with ChatGPT were ‘‘part of the training data or accessible to the reasoning process.’’ He wrote that the reply he received only addressed direct access to his conversations, not whether those exchanges had been incorporated into the training pools used to improve models. Thom described the response as insufficient and accused the company of misleading behavior. The exchange echoes earlier disputes about OpenAI’s handling of user-provided material. In announcing a claimed Navier–Stokes breakthrough, OpenAI stated it had not accessed any specific user data to reach the result, while also saying it could not fully exclude an indirect effect from de-identified usage-derived data. That ambiguity has left some researchers unconvinced, with concerns that de-identification does not erase the intellectual content of mathematical ideas. Researchers say the episode has intensified unease in mathematical circles at a moment when OpenAI’s claimed achievements are attracting attention. Several mathematicians told The Verge they worry the prospect that private or unpublished work could be absorbed into model training without consent or clear attribution will push the field toward greater secrecy. OpenAI did not immediately respond to The Verge’s request for comment.

Keep Reading