Now is the time to give LLMs access to the ACM digital library
Opening up the ACM Digital Library to large language models is prompting clashes over who should control and benefit from decades of scholarly work. Many argue the library should be fully open to humans and machines alike, especially given public funding of research and the likelihood that AI companies have already scraped much of the content, while others focus on copyright, author consent, and fears that paywalled licensing to big AI firms will deepen inequalities and further erode an already fragile peer‑review system. Underneath is a broader worry that LLMs could entrench new gatekeepers over scientific knowledge rather than expanding genuine access.
Open access and ACM’s policies
- Several argue ACM’s digital library should have been fully open to humans long before discussing LLM access.
- Others note ACM has already committed to making all publications open access from Jan 2026, though some find current “open access” confusing, paywalled, or tied to author fees.
- Some see ACM’s move as a “marketing” definition of open access, pointing to Sci‑Hub/LibGen as de facto open access.
LLMs, access, and research usefulness
- Some researchers report that LLMs dramatically increased their use of primary literature by helping trace claims back to original papers.
- Others think LLMs have almost certainly already ingested most ACM content via arXiv, Sci‑Hub, and other PDF sources; a minority disputes how complete that is.
- A recurring view: blocking LLMs only hurts rule-followers; bad actors will scrape anyway.
Licensing, ownership, and compensation
- One camp sees ACM’s licensing plans as hypocritical: authors did the work, publishers capture the LLM revenue, and authors face growing burdens (student AI use, broken peer review).
- Others counter that scholarly authors historically did not expect royalties; by publishing, they chose prestige and knowledge dissemination, not payment.
- Some suggest: free access for open‑weight/non‑profit models, paid licenses for closed, commercial models.
Copyright, fair use, and legality
- Debate over whether training on ACM papers is copyright infringement or fair use:
- One side stresses that copyright covers expression, not ideas; reading and using ideas (human or machine) has always been allowed.
- The other highlights verbatim or near‑verbatim regurgitation, derivative works, and recent lawsuits; argues that LLM outputs can substitute for the original and thus infringe.
- There is no consensus; posters note courts are still clarifying boundaries and often disagree on what current rulings actually imply.
Impact on peer review and academia
- Some think LLMs will “wreck” peer review (except for hard‑experimental work) and that peer review is already deeply flawed or largely symbolic.
- Others say it was broken long before AI; preprint culture (arXiv) is cited as evidence.
- A few advocate burning down predatory publishing structures if they become AI gatekeepers instead of enabling open scholarship.
Equity, power, and social impact
- Concerns that selling access to big AI labs will deepen concentration of power and exclude smaller/open players.
- Skepticism that AI-driven “abundance” will reach everyone; many foresee more inequality, potential authoritarian control, or social unrest rather than broad prosperity.