A second mathematician accuses OpenAI of dishonesty over its training data

A second mathematician has publicly accused OpenAI of dishonesty over the origins of the training data behind its recent string of mathematical breakthroughs — an escalating credibility dispute that arrives just weeks after the company's celebrated announcement of a Millennium Prize-level result.
In a series of posts on Mastodon, mathematician Andreas Thom said he began questioning his own interactions with ChatGPT after Tristan Buckmaster, a mathematics professor at New York University, publicly raised similar concerns about OpenAI's Codex. One of the ten results OpenAI announced last month fell squarely in Thom's specialty — so-called non-sofic groups, infinite mathematical structures that cannot be approximated by finite ones — and built heavily on prior work by Thom and fellow mathematician Gábor Kun.
A quietly amended writeup and an unsatisfying answer
OpenAI's original announcement of the non-sofic groups result was criticized in mathematical circles for failing to credit Thom and Kun's recent contributions; the company quietly revised its writeup afterward. Thom said he was also struck by what he called OpenAI's “detailed command” of techniques that were neither the most obvious nor most promising approaches available at the time — a detail he found suspicious enough to email OpenAI researchers Sébastien Bubeck and Mark Sellke directly, asking whether his ChatGPT conversations had entered the company's training data or reasoning pipeline.
The response he received, he said, addressed only whether his conversations could be directly accessed — not whether they had been folded into the broader training corpus used to improve OpenAI's models. “No such qualification, explanation, or evidence was given,” Thom wrote. “I take this as dishonesty to say the least.”
The same obfuscation pattern, twice
Thom drew a direct parallel to how OpenAI handled Buckmaster's concerns over the Navier-Stokes result, part of the company's Millennium Prize announcement. In that case, OpenAI's blog post stated flatly that “no specific user data was accessed” to solve the problem, while conceding it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Thom called this the identical obfuscatory move the company made in its private correspondence with him: “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.”
Thom argues the burden of proof sits with OpenAI, not with the researchers trying to reverse-engineer an opaque training pipeline. “Only OpenAI has the relevant data for that,” he wrote, calling on the company to disclose the datasets and settings necessary to conclusively rule out indirect use of unpublished mathematical work. He characterized the underlying scenario — nonpublic research improving models that then compete with the same researchers for publication credit — as “ethically indefensible” if confirmed. OpenAI did not immediately respond to a request for comment from The Verge.
A chilling effect on open collaboration
The dispute compounds pressure on OpenAI at what should be an unambiguous high point: its claimed solution to a Navier-Stokes Millennium Prize problem, if independently verified, would be a landmark achievement. But researchers told The Verge the recurring credit and provenance disputes risk pushing the mathematics field toward greater secrecy, since even informal rumors of progress on a hard problem could now trigger a well-resourced AI lab racing to solve it first using indirectly absorbed insights. As first reported by The Verge.
Originally reported by The Verge. Read the original article for additional details.
View original source