OpenAI says an artificial intelligence system has solved one of mathematics' most famous unsolved problems, but a dispute involving two mathematicians who used the company's Codex tool is raising a separate question: what happens when researchers share unpublished work with an AI platform that is also developing competing research capabilities?
The company announced September 8 that an internal AI system had produced a proposed solution to the Navier-Stokes existence and smoothness problem, one of seven Millennium Prize Problems established by the Clay Mathematics Institute.
OpenAI says its system produced both a mathematical proof and a formalized version in the Lean proof language showing that smooth fluid motion can develop a singularity in finite time. The effort involved roughly 10,000 AI agents working concurrently.
The result has not been formally recognized by the Clay Mathematics Institute. Under its rules, a proposed solution must be published through a qualifying outlet, remain public for at least two years and gain general acceptance in the mathematics community before it can be considered for the $1 million prize. OpenAI says it does not intend to claim the award.
The legal issue emerged because New York University mathematician Tristan Buckmaster and Levent Alpöge, a mathematician who works for Anthropic, had been pursuing closely related research while using AI tools, including OpenAI's Codex.
Buckmaster said the pair entered drafts from their work into Codex while developing results involving fluid-dynamics equations.
After learning that OpenAI was devoting major computing resources to the Navier-Stokes problem, Buckmaster said he asked company researchers whether its internal model had been trained on or otherwise had access to their Codex sessions.
He stopped short of accusing OpenAI of taking the work.
"I do not know whether our data was used," Buckmaster wrote in a public statement.
OpenAI says neither its researchers nor the AI agents working on the problem accessed the mathematicians' specific prompts, drafts or other user data. The company says its researchers did not see Buckmaster and Alpöge's work until it was publicly released.
OpenAI did acknowledge one possibility.
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models," the company said.
That distinction raises a broader legal question as researchers increasingly use generative AI before their work becomes public.
Under U.S. copyright law, an unpublished mathematical paper or written proof can receive protection once it is fixed in a sufficiently permanent form. Publication is not required.
Copyright, however, protects the author's expression, not the underlying mathematical idea, discovery, process or principle.
That means copying protected portions of another person's written proof could potentially raise copyright concerns. Independently arriving at the same mathematical result generally would not.
Scientific priority also does not automatically create an intellectual property right. Academia may place enormous importance on who discovered a result first, but copyright law does not give the first researcher exclusive ownership of the underlying concept.
Trade-secret law could potentially protect some unpublished research, though the requirements are different. Federal law generally requires the information to have economic value from remaining secret and requires its owner to take reasonable steps to preserve that secrecy.
The terms governing the AI platform may therefore become just as important as traditional intellectual property law.
OpenAI's consumer Terms of Use state that users retain ownership rights in material they submit. The terms also permit OpenAI to use content to operate, maintain, develop and improve its services, while allowing users to opt out of having their content used for model training.
Codex practices vary by account type. OpenAI says Business, Enterprise and Edu customer content is not used to improve its models by default. Plus and Pro users can disable model training through ChatGPT's data controls.
There is no public information establishing which terms and data settings applied to every Codex interaction involving Buckmaster and Alpöge's work.
Ownership also does not necessarily mean a platform has no permission to use submitted material. A person can retain rights in content while granting a company contractual permission to process or use it for certain purposes.
That makes disputes involving AI-assisted research more complicated than a conventional copying case, particularly when the original material does not appear verbatim in a later AI-generated result.
Buckmaster has not filed a lawsuit accusing OpenAI of copyright infringement, trade-secret misappropriation or another legal violation. OpenAI maintains that its proposed proof differs significantly from the mathematicians' work and that its researchers and agents did not access their specific material.
The dispute may never reach a courtroom, but the underlying issue is likely to become more common.
Scientists, engineers, programmers and other researchers increasingly enter unpublished ideas, preliminary results and source code into AI systems operated by companies developing their own increasingly capable research tools.
The Navier-Stokes controversy shows why ownership may be only part of the question. As AI becomes embedded in the research process, the more difficult issue may be what rights a platform receives to use what researchers share with it, and how anyone could prove that unpublished work later influenced an AI discovery.