"Plainly dishonest": mathematician says OpenAI lied to him about training on his private chats

Andreas Thom, who's a group theorist at TU Dresden, has publicly questioned whether his private ChatGPT sessions fed into the OpenAI result that produced the first known non-sofic group — one of ten problems OpenAI said Astra settled on August 1.
In a lengthy three-part post on Mathstodon, Thom wrote that he emailed OpenAI researchers Mark Sellke and Sebastien Bubeck shortly after that announcement. He and a Dresden colleague had spent months working through the expander matching problem and extensions of his work with Gabor Kun inside ChatGPT, and asked two things: whether those exchanges entered training data, and whether the system could reach them while working on the proof. Sellke's reply, as Thom quotes it, was a single line: that did not happen. Thom argues the answer addressed only the 'direct access' part. He now describes it as unjustifiably broad and, in hindsight, dishonest.
The exchange resurfaced this week alongside OpenAI's Navier-Stokes announcement. There, OpenAI said neither its researchers nor its agents saw the pair's work before publication and that no specific user data was accessed. It also added that it could not rule out that de-identified data derived from their use of its products had helped improve its models.
Two other points support his argument as well. He switched off model training on 29 June. Model training is a setting that users can't audit — it says nothing about earlier chats. Most researchers favoured quantum-games methods over the approach he took with Kun, so he was surprised Astra chose the same route as him.
As of writing, Thom has stopped short of claiming proof. Instead, he has asked that OpenAI discloses the basis for its denial.
OpenAI has not publicly addressed his posts.
Source(s)
@andreasthom on mathstodon via @ValerioCapraro on X
Featured image by Levart_Photographer on Unsplash







