Better Call GPT: Comparing large language models against lawyers [pdf]

Large language models are increasingly being tested against human lawyers for tasks like contract review, with some studies and startups claiming comparable accuracy at a fraction of the cost. Commenters see major potential for cheaper access to legal information and routine document analysis, but highlight serious concerns around hallucinations, lack of liability, regulatory barriers, and the inability of current models to handle high‑stakes, fact‑intensive work such as criminal defense or complex family law. Many expect LLMs to become powerful tools that augment, rather than replace, lawyers—especially for standardized documents and preliminary research—while warning that guild-like professional structures and legal restrictions may slow deeper disruption.

Promise of AI for Affordable Legal Help

  • Many see LLMs as a near-future “lawyer for the masses”: reviewing personal documents (tax, medical, insurance, marriage), explaining options, and giving low‑stakes guidance without $500/hr billing.
  • People already use LLMs for early research on licenses, contracts, and software legal issues to learn terminology and frame questions before hiring a professional.
  • Similar excitement appears for medical and financial advice, and for consumer contract summaries (leases, ToS, insurance).

Reliability, Liability, and Regulation

  • Core concern: LLMs can be confidently wrong, with no malpractice liability or bar oversight.
  • Lawyers have licensing, ethics rules, malpractice insurance, and reputational risk; LLMs and vendors currently do not.
  • Ideas raised: forcing “legal GPTs” or their operators to carry malpractice‑like insurance; new “AI advice” insurance; but enforcement and collectability are unclear.
  • Unlicensed practice of law rules and lawyer‑dominated legislatures are seen as likely barriers to fully automated legal services.

Capabilities: Where LLMs Help vs. Fail

  • Strongest current use: contract/document review, clause extraction, summarization, and “is this standard?” checks, especially for high‑volume, low‑risk agreements.
  • Multiple practitioners report LLMs are better at analyzing than generating contracts; standardized templates plus light AI editing are preferred over fully AI‑drafted documents.
  • Weaknesses: complex, fact‑specific work (criminal defense, family law, bespoke deals), subtle contingencies (e.g., trust beneficiaries), and negotiation strategy.
  • LLMs lack intentionality and can’t provide nuanced risk assessments like “it’s defensible but gray,” which is central to real practice.

Impact on Lawyers, Access, and Power

  • Some expect major efficiency gains and fewer junior roles; others think guild‑like protections will preserve high‑end work and pricing.
  • There is tension between empathy for displaced professionals and the huge unmet need for affordable counsel.
  • Worry: governments might use LLMs to replace human defense counsel for the poor, weakening Sixth Amendment protections.

Technical & Product Considerations

  • RAG, anonymization, narrow scopes, and human‑in‑the‑loop validation are seen as pragmatic patterns.
  • Users observe safety “nerfing” and context limits; others note newer models benchmark better but don’t solve truthfulness.
  • Several startups are already selling AI‑assisted review and Q&A, but monetization and trust remain challenging.

Critiques of the Paper

  • The study reportedly used only 10 procurement contracts and a single review playbook; commenters find this too narrow for “industry‑level” claims.
  • Some view it as more marketing than rigorous benchmarking, especially given commercial ties.