When Mathematicians Hand Unpublished Proofs to AI, Is the Platform Still Just a Tool?
An unresolved dispute in mathematics has brought a harder question into the open: what happens to the data boundary when genuinely valuable work moves onto an AI platform?

I came across a remarkable story today.
Mathematicians Tristan Buckmaster and Levent Alpöge used models including Claude and Codex while working on three results involving fluid equations. The one drawing the most attention concerns finite-time blowup for the three-dimensional incompressible Euler equations with a smooth forcing term. They also formalized and verified the proof in Lean.
The boundary matters here: this is an Euler result with smooth forcing. It is not a result for the unforced Euler equations, and it does not solve the full Navier-Stokes Millennium Prize problem.
Nor did the two authors, or a language model, invent this line of attack from scratch. In his original statement, Buckmaster credits the foundational idea to Diego Córdoba and Luis Martínez-Zoroa. Building on earlier work, Buckmaster and Alpöge extended the forcing-based construction to additional equations. Terence Tao also noted that these results remain some distance from Navier-Stokes.
That alone would already be important.
But a more complicated story emerged outside the mathematics itself. It is not mainly about whether a model can do mathematics. It is about who can see a research process once it enters an AI platform, and who might benefit from what that process reveals.

As of September 8, 2026, based on public results, Buckmaster’s account, and Bubeck’s public denial. The colors distinguish different kinds of evidence; they do not assign a verdict.
An unpublished hundred-page proof
For now, this part of the story rests largely on Buckmaster’s account.
According to him, word had begun circulating that “Anthropic had solved a major mathematical problem.” Alpöge was also warned that news of their progress had reached OpenAI.
On September 3, Buckmaster emailed a mathematician at OpenAI. He confirmed that he and Alpöge had a result coming out and explained that it was a private collaboration, not work on behalf of either employer. He did not disclose the proof. The researcher replied the same day and asked for enough information to avoid having the two sides compete on the same problem.
Two calls followed on September 6. Buckmaster says he was told that an internal OpenAI model had produced a proof of roughly one hundred pages for finite-time blowup in Navier-Stokes with smooth forcing. Buckmaster has not seen that proof, and OpenAI has not made it public.
He then pressed for an answer about when the work had begun. Buckmaster says it was eventually acknowledged on the calls that the first prompts were sent only after news of his and Alpöge’s progress had reached OpenAI. He also says the effort involved a team, multiple attempts, and substantial compute.
Buckmaster further claims that OpenAI proposed two publication arrangements. Under one of them, he would be the sole named author presenting OpenAI’s result, while Alpöge, who works at Anthropic, would be excluded. Buckmaster declined.
Even if the story ended there, it would raise questions about research priority, credit, and commercial competition. But the question Buckmaster asked next is the one I find more important.
Even without exchanging proof details, the knowledge that someone is close to a result, or that a particular direction deserves concentrated resources, can itself be valuable. That does not prove OpenAI used data it should not have used. It does show why verbal assurances alone are unlikely to settle doubts when a platform and its users are conducting research in the same field.
They put the entire research process into Codex
Buckmaster and Alpöge used models including Claude and Codex throughout the project. In his statement, Buckmaster says every draft was placed into Codex conversations.
So he asked OpenAI directly: had its internal model been trained on, or otherwise accessed, those conversations?
According to Buckmaster, he was told that models do not proactively search user data. When he followed up by asking whether the data had been used for training, he received no answer.
Those are different questions. Retrieval at inference time, use in training or evaluation, and access by internal staff are three separate permissions. Answering one does not answer the other two. At the same time, the absence of an answer is not evidence that OpenAI used the data.

Inference-time retrieval, training and evaluation, and staff access are separate permission gates. Retention, deletion, and audit logs run across the entire data lifecycle.
Buckmaster is careful about the limits of what he knows. He has not seen OpenAI’s proof. He does not know what the internal model actually did, and he does not know whether the researchers’ Codex data was used. He therefore says he has no evidence on which to accuse anyone.
As of publication on September 8, OpenAI’s Sébastien Bubeck had issued a brief denial, calling the allegations false and inflammatory and saying that a fuller response would follow. That fuller response had not yet appeared.

Bubeck’s initial response on X. The image preserves the account information and first sentence of the denial; the full post is linked above.
For now, then, we can confirm only that Buckmaster published his account and Bubeck denied it. Whether OpenAI’s proof is valid, how the internal work actually began, what was said on the calls, and whether the Codex conversations were used all remain unknown.
An AI conversation, or a laboratory notebook?
The question left behind by this dispute is not limited to mathematics.
If a researcher gives an AI every unpublished conjecture, draft, failed path, and derivation, is that material ordinary product data, or is it a laboratory notebook?
Those categories should receive very different levels of protection.
People may accept that ordinary product data is used to improve a service. A laboratory notebook, however, can contain unpublished findings, evidence of research priority, and the direction of several years of future work. Even if a final proof is produced independently, simply knowing which path deserves concentrated compute may have real value.
OpenAI’s current public policy is more specific than the answer Buckmaster says he received on the call. Content from individual ChatGPT and Codex accounts may be used for training, and users can opt out through data controls; the full Codex environment also has a separate setting. By contrast, inputs and outputs from ChatGPT Business, Enterprise, and the API are not used for training by default unless an organization explicitly opts into data sharing.
But those public rules still do not answer the question in this particular case. We do not know which product Buckmaster used, what type of account he had, what settings were active, or how those conversations were actually handled.
That is why an enterprise buying AI cannot stop at the phrase “not used for training by default.”
When the tool provider does the same kind of work
What companies put into large models is no longer limited to chat messages.
It includes source code, product plans, tender proposals, customer records, and even next year’s strategic bets.
Companies used to worry mainly about whether that material might leak to outsiders. Now they may need to ask another layer of questions. Who inside the platform can access the data? Can it enter training, evaluation, or human review? Are different models, teams, and environments genuinely isolated? How long does deletion take? If a dispute arises, can the provider produce a complete access and processing log?
Those questions matter even when a platform only provides a general-purpose tool. They become more sensitive when the same platform also writes software, builds products, conducts research, and enters specialized industries of its own.
The tool provider and the user may be collaborators and potential competitors at the same time.
Enterprise buyers will not compare large-model vendors on capability and price alone. Training boundaries, internal access, data isolation, retention and deletion, and auditability will become as important as accuracy.
It is not realistic for every company to stop using cloud-hosted models. What needs to change is how data is classified and how these tools are purchased.
Authorized public information that contains nothing sensitive can go into general-purpose tools. Ordinary internal material needs explicit data controls. Core source code, unpublished research, customer secrets, and strategic plans belong in environments backed by contractual commitments, access controls, retention policies, and audit capability. At higher levels of risk, organizations should consider controlled deployments, isolated environments, or simply withholding the complete material from the model.
The more useful AI becomes, the less ambiguous the boundary can be
Determining who was right in this dispute will require a fuller response and more evidence.
The validity of the mathematical result and whether the alleged conduct occurred are also separate questions. Even if OpenAI’s proof turns out to be correct, that would not answer how the data was obtained or how authorship was handled. Conversely, a dispute over communication cannot show that the mathematical proof is wrong.
But the episode is already a warning. Once AI enters genuinely valuable work, the data boundary can no longer be left to a user agreement that almost nobody reads.
Researchers need to know whether they are handing over a single question or an entire unpublished line of inquiry.
Companies need to know whether they are buying a tool or giving their most valuable work process to a platform that may one day compete with them.
The boundary must be clear before the work begins, and it must be verifiable.
Reference
Discussion