On the Navier–Stokes Millennium Prize Problem
OpenAI claims an internal AI system has produced a Lean-verified solution to the Navier–Stokes Millennium Prize Problem after running tens of thousands of agents and burning millions of dollars’ worth of compute, which many see as a watershed moment for AI’s mathematical capabilities. At the same time, serious concerns are raised that the model may have been indirectly trained on private ChatGPT/Codex sessions of independent mathematicians working on related problems, and that OpenAI then rushed to “scoop” them while pressuring for asymmetric authorship credit. Commenters worry this sets a precedent that incentivizes secrecy in research, blurs the line between assistance and plagiarism, and shows how dependent future scientific progress could become on opaque corporate AI systems.
Compute, Cost, and Methodology
- Commenters estimate ~$15M at public Astra prices for the ~300B output tokens; some note OpenAI’s internal cost is lower, but still huge.
- The system used ~10,000 concurrent agents over ~88 hours. Some see this as “country of geniuses in a datacenter”; others argue it’s brute force plus heavy human steering, not a clean “one-shot” AGI.
- Several point out that the chain of prompts, tool use, and human guidance is nontrivial, and that the blog post downplays human involvement.
Priority Dispute and Plagiarism Concerns
- A separate team had been working for about a year on closely related Euler/NS work using Codex/LLMs; their public statement alleges:
- OpenAI started their effort only after hearing rumors of a Millennium solution.
- Their Codex sessions contained drafts and key ideas.
- Questions about whether those chats were used in training were not clearly answered.
- OpenAI proposed joint announcements conditioned on dropping the Anthropic-affiliated coauthor and allegedly used career-threat language.
- OpenAI responses (in the post and on social media) claim:
- They never accessed specific private chats.
- The model’s proof and even the precise Euler result differ.
- User data might have contributed “de-identified” training signals; they “cannot rule it out.”
- Thread splits:
- One side sees this as likely IP leakage / “front‑running” customers using their own data.
- The other side insists there is no concrete evidence of misconduct and that allegations require proof.
Data Use, Privacy, and Trust
- Heavy focus on the training-toggle (“improve model for everyone”) and whether opt‑out truly prevents use of either raw or “de‑identified” data.
- Some argue OpenAI could and should determine whether specific chats entered training; others say scale and de‑identification make that infeasible or privacy‑violating.
- Many conclude researchers should avoid proprietary LLMs (or demand strict zero‑data‑retention) for unpublished work.
Impact on Mathematics and Research Culture
- Mixed reactions: awe at a Millennium problem falling; dismay at the way it happened.
- Concerns that rumors of progress now trigger massive AI efforts that “flatten” open problems before human projects mature, pushing mathematics toward secrecy (“cognitive dark forest”).
- Some say this showcases AI’s frontier reasoning and signals AGI; others liken it to Deep Blue or chess engines: superhuman in a narrow, well‑verified domain.
Lean Formalization and Correctness
- Many emphasize the Lean proof as strong evidence of correctness, possibly stronger than traditional peer review.
- Others note:
- Lean statements must exactly match the Clay problem, which still requires human checking.
- Past bugs in proof assistants mean some human scrutiny is still necessary.
Broader Societal and Job Implications
- Debate over whether this accelerates an “intelligence explosion” threatening knowledge-work jobs versus merely pushing humans to higher‑level tasks.
- Several see it as a watershed moment; others warn against over-interpreting one high-profile math result as general AGI.