More questions about whether researchers can trust OpenAI with unpublished math

Accusations that OpenAI’s models may have used private chat logs from mathematicians to help solve major open problems have intensified concerns about data misuse and academic credit. Commenters debate whether recent “AI breakthroughs” in areas like Navier–Stokes and group theory reflect genuine algorithmic innovation or the aggregation and amplification of near-finished human work, possibly obtained via research chats and user prompts. Many see this as a broader warning about trusting cloud-based AI with sensitive intellectual property, arguing for local models, stronger regulation, and clearer guarantees around training data and attribution.

Alleged use of unpublished math in training

  • Several mathematicians claim OpenAI’s internal models likely trained on private ChatGPT/Codex sessions containing work-in-progress on hard problems (e.g., Millennium Prize problems).
  • One mathematician says they opted out of training in late June, asked OpenAI about prior use, and initially got a non-answer, then a later statement that data after early July could not have influenced the system.
  • Many commenters note it is practically impossible for outsiders to verify whether such claims are true.

Plagiarism, scooping, and attribution

  • A core grievance: AI labs allegedly “sniped” nearly-finished work by throwing enormous compute and internal teams at it, then marketing solutions as AI breakthroughs.
  • Accusations include: using private chats as training data, omitting or downplaying key human contributors in citations, and trying to shape authorship to exclude collaborators at rival labs.
  • Others argue there’s no concrete proof of deliberate theft, liken this to traditional scooping in academia, and note the mathematicians themselves used AI tools.

Capabilities vs piggybacking

  • One camp sees the results as genuine evidence that frontier models can make major mathematical advances via agent-based search plus proof assistants (e.g., Lean), even if based on known frameworks.
  • Another argues the “breakthroughs” may mostly be systematic exploitation of an existing “overhang” of human ideas and partially completed work, not autonomous creativity.

Data use, consent, and dark patterns

  • Many were surprised or angry that prompts are used for training by default, and that opting out is hidden, can be reset, or partially bypassed (e.g., via thumbs-up/down feedback or “analytical purposes”).
  • Debate over whether “de-identified” user data being used for training is acceptable; some see this as exploitation and appropriation, even if within the TOS.

Impact on math and research careers

  • Reports that students are questioning whether to pursue pure math PhDs if AI labs can front-run and overshadow their work.
  • Some foresee a shift toward private, non-cloud workflows, arXiv-style early timestamping, or grants for local models.

Broader trust and structural concerns

  • Widespread sentiment that large AI labs, built on mass scraping and aggressive IP use, are fundamentally untrustworthy.
  • Calls for stronger regulation, clearer consent standards, and a push to local/open models; others argue this is “common-sense opsec” users must adopt.