Word2Vec received 'strong reject' four times at ICLR2013

Word2Vec, a now‑foundational technique for learning word embeddings in NLP, was repeatedly rejected by an academic conference in 2013, prompting debate over how well peer review handles novel or impactful ideas. Commenters argue over whether the reviews actually improved the paper’s rigor or reflected systemic biases toward incremental, easily benchmarked work, touching on issues like reviewer quality, incentives, and the reproducibility crisis. Several propose alternative models for scientific communication — from open platforms like arXiv and OpenReview to more transparent, GitHub‑like workflows — while others defend peer review as a necessary but flawed quality filter.

Role and Limits of Peer Review

  • Many see peer review as ill-suited to recognizing novel ideas; it rewards “peer think,” incremental work, and good formatting over innovation.
  • Others argue the process often improves clarity and rigor, even for ultimately influential papers, and that almost any manuscript can be strengthened via review.
  • Several note that reviewers are not supposed to validate correctness, only assess whether work is clear and reproducible; replication and follow-up are meant to test truth.

Word2Vec Reviews and Title Nuances

  • Commenters who read the ICLR reviews find them generally reasonable: they criticize minimal model description, unclear comparisons, and missing explanations.
  • The thread notes that four “strong reject” entries appear to be duplicates from a single reviewer, making the headline potentially misleading.
  • Some argue the reviews pushed the authors to clarify and extend the work; others say important but rough ideas get blocked because the bar for rigor is too high relative to novelty.

Innovation vs Rigor, Quality vs Impact

  • Debate over whether “quality” (clarity, rigor) should be judged independently from potential impact; some think this separation is a core flaw.
  • Others emphasize that future influence is extremely hard to predict, so process should focus on clear exposition and solid methodology.
  • Several claim conferences overemphasize novelty, new architectures, and SOTA metrics, undervaluing simplification, negative results, and replications.

Systemic Problems in ML Conferences

  • Complaints include too many submissions, overreliance on inexperienced reviewers, generic or erroneous reviews, and near-zero-sum accept/reject decisions.
  • There is concern about low accountability for reviewers and lack of real feedback mechanisms on the review system itself.

Credit, Memory, and Trust

  • The discussion covers disputes over who originated certain neural translation ideas and frustration over missing acknowledgments, seen as symptomatic of money- and prestige-driven behavior.
  • Some recount unreproducible results in related work and unresponsiveness from coauthors, feeding distrust.

Reform Proposals and Alternatives

  • Suggestions include GitHub-like systems with code/data verification, OpenReview-style open commenting, and internet forums as de facto post-publication review.
  • Others warn that platforms with voting (e.g., Reddit) exhibit severe groupthink and popularity bias, so they’re no panacea.