How we measured AI writing across arXiv, and where the measurement breaks
An analysis of 12,750 arXiv papers from 2021–2026 claims that up to 39% of recent submissions, and 65% in computer science, show statistical signals of AI-generated writing, compared with 0.4% before ChatGPT’s release. Commenters question the reliability and transparency of AI detectors, share examples of false positives, and note how evolving writing styles and widespread use of AI “polishing” blur the line between human and machine text. The thread wrestles with whether AI-assisted papers are a problem at all, raising concerns about trust, fraud, reviewer overload, and the erosion of writing as a quality signal in science.
Detector methodology & reliability
- Many commenters question the detector’s rigor, interpretability, and reproducibility, especially the way multiple sub-scores are combined.
- Users report high “machine” scores on clearly human, pre-LLM work (theses, older papers, Stack Overflow answers), suggesting nontrivial false positives.
- Others see the detector performing well on their own texts and view a calibrated pre-ChatGPT false-positive rate (~0.4%) as evidence it “works on average.”
- Concerns include:
- Training data leakage and domain drift (new jargon, shifting styles).
- Sensitivity to LaTeX/formatting.
- Theoretical limits of text-only detection and convergence of human and AI styles.
- Commercial tools like Pangram and GPTZero are cited as both “quite accurate” and “trivially fooled,” underscoring disagreement on what “accuracy” means in practice.
Prevalence estimates & where they might fail
- The reported rise to ~39% AI-flagged papers overall and ~65% in CS is seen by some as striking, and roughly aligned with independent checks on subsets of arXiv.
- Others argue the baseline is shaky: language and fashions change, and detectors may be picking up “modern style” or topic drift (e.g., LLM jargon) rather than genuine AI authorship.
- Suggestions: include older preprints (pre-2020), break down by field, and apply detectors to presumably higher-curation venues (e.g., top journals) as a control.
“So what?” Impact on science and trust
- One camp sees little inherent problem: if the science is sound and papers become clearer, AI assistance is acceptable; the real issue is underlying research quality, not who wrote the sentences.
- The opposing view stresses:
- LLMs’ tendency to hallucinate facts, citations, and even entire related-work narratives.
- Erosion of trust and quality signals (polished prose no longer implies effort or rigor).
- Increased reviewer burden and rising signal-to-noise, with analogies to spam and “enshittification.”
- Examples are given of papers with entirely hallucinated citations and of conference reviewing already strained and low-effort.
Use of LLMs for polishing and non-native authors
- Non-native speakers describe using LLMs to turn “basic” English into conventional scientific prose; many see this as necessary and fair, provided scientific content is checked.
- Critics argue that:
- “Polishing” often introduces new claims or subtly changes meaning.
- Chasing a generic “scientific style” is itself harmful and unnecessary.
- There is disagreement over whether this is a minor tool use or a substantive co-authorship that undermines personal accountability.
Incentives, spam, and detection arms race
- Commenters link rising AI usage to structural incentives: publish-or-perish, hatred of writing, and cheap text generation.
- Some expect more “spam papers” and advocate treating AI-style text as a coarse filter, even if it occasionally rejects good work.
- Others highlight the danger of misusing detectors (especially in education) and note that AI-generated and AI-evading text can both be produced easily, making any stable equilibrium in detection unlikely.