Bankless · Friday, September 11, 2026
Mathematicians Tristan Buckmaster and Levent Alapataji allege that OpenAI used their input data from using OpenAI's Codex models to solve a million-dollar Millennium Prize math problem. While OpenAI denies directly using their chat logs, they admit de-identified data may have helped improve their models. This situation highlights concerns about data sovereignty and the 'AI arms race' where companies prioritize data acquisition for model training.
“The allegation is that OpenAI front-ran these mathematicians, took their proprietary information, they thought this was a private chat, and then actually used it in order to get the prize, capture the notoriety, solve the problem, and front-run their user base.”
“OpenAI came out and said, we didn't actually use any of these chat logs. They said, we cannot rule out that de-identified data derived from their usage of our product helped improve our models.”
“The claim is that OpenAI took the inputs that Tristan and his co-mathematician Levant were putting into OpenAI codecs, and they were using that as a jumping off point And then OpenAI used a ton of compute to basically finish the job and truly solve the problem.”