The AI bullshit singularity

Fears are growing that large language models will trigger an “AI bullshit singularity,” where the web becomes saturated with low-quality, AI-generated text and images that future models are then trained on, degrading their capabilities over time. Commenters debate whether synthetic data and self-training inherently cause this downward spiral or whether careful curation, human feedback, alternative data sources (like code, math, and video), and reputation systems can keep models improving. The exchange highlights broader concerns about trust, gatekeeping, and how to preserve high-quality human creativity and information in an environment increasingly flooded by machine output.

Self-Training, Synthetic Data, and “Incest” Analogies

  • Many compare repeated training on AI-generated data to inbreeding: errors compound, diversity shrinks, and models may converge on a “bullshit” local maximum.
  • Others argue synthetic data can be as good or better if paired with strong external signals (e.g., game outcomes, human votes, task success).
  • Some note that models are already trained on other models’ outputs, and claim targeted synthetic data is preferable to random internet text.

Internet Quality and “Bullshit Singularity”

  • Several say the internet was mostly low-quality or SEO spam long before LLMs; AI just scales it.
  • Concern: AI will massively increase hay (junk) around the needle (good content), making quality harder to find.
  • Counterpoint: individuals can still “pull” directly from trusted sources; system-wide spam doesn’t stop that.

Curation, Reputation, and Gatekeepers

  • Expectation that curation, identity, and reputation will become more central as quality filters.
  • Some foresee cryptographic signatures and institutional certification of “human-made” content; others doubt crypto’s real-world robustness.
  • Worry that stronger curation leads to new gatekeepers and brand/reputation capture; optimism that honest curators will still exist and be valuable.

Art, Images, and Creative Work

  • Fear that image models will reinforce clichés as new scrapes are polluted with AI art; early models trained on “untainted” data may have a moat.
  • Others insist pure synthetic hill-climbing plus human fitness signals (what gets shared/kept) can still yield progress.
  • Some think abundance of generic AI output will ultimately increase the value and distinctiveness of genuinely original human work.

Singularity, Scaling, and Limits

  • Debate over whether self-improvement inevitably slows (diminishing returns, fixed compute) or can keep delivering >50% gains across many dimensions.
  • Some think humans are nowhere near the physical limits; others stress eventual ceilings and carrying capacities.

Hallucinations, Verification, and Intelligence

  • Skepticism that hallucinations can ever be fully removed from probabilistic text generators.
  • Suggestions: external tools (web search, code execution, robotics, math, simulations) and “verifier AIs” could constrain errors.
  • Ongoing argument over whether LLMs “understand” or merely do sophisticated pattern-matching; comparisons with humans, parrots, and animal cognition surface repeatedly.

Data, Human Feedback, and Economics

  • Concern that high-quality RLHF/SFT data at expert level is expensive, and current annotation markets race to the bottom on quality.
  • Speculation that AI companies may eventually subsidize or even pay royalties to writers/artists to obtain premium training data.

Overall Mood

  • Thread mixes strong anxiety (polluted knowledge ecosystem, loss of creativity, “life thieves”) with optimism (better tools, new forms of curation, continued advances via synthetic data and new methods).
  • Many note that predictions are highly uncertain; even six-month forecasts feel unreliable.