| Dimension | Score |
|---|---|
| Output reliability | 2/5 |
| Workflow fit | 3/5 |
| Context handling | 4/5 |
| Learning curve vs. payoff | 3/5 |
| Failure transparency | 2/5 |
| Overall | 2.8/5 |
What AI does well
First-pass contract review. Spellbook inside Word and Claude with a document dump both flag missing clauses, unusual terms, and deviations from a playbook. Observed gain: 2-4 hours a week on routine review. Caveat: the tools are better at flagging what is present and odd than at noticing what is absent entirely.
Summarizing long records. Deposition transcripts, medical records, and discovery productions are exactly the long-document work Claude handles best: a 400-page record becomes a dated chronology in minutes. Caveat: the summary occasionally drops the one entry that matters, so it is a reading guide, not a substitute for the read.
Drafting routine documents. Engagement letters, discovery request shells, deadline reminder letters, and privilege log descriptions draft well in Harvey or ChatGPT from a matter summary. Caveat: the boilerplate arrives in the wrong jurisdiction’s flavor unless you say so explicitly.
Intake and file organization. Classifying, renaming, and indexing an unstructured production, and building a first-pass privilege log, is tedious, low-risk work the tools do reliably.
Where it fails
Legal research and citations. The unsolved failure of the category. Asked for supporting authority, models still produce case names with real-sounding parties, a real reporter, and a volume number that does not exist. Courts have sanctioned filings with AI-invented citations every year since 2023, and 2026 has not broken the streak. A concrete example: a cite-check request returned five authorities; three were real, one was a real case that said the opposite, and one did not exist at all. The output looked identical for all five. This is why output reliability and failure transparency both score 2.
Privilege and confidentiality judgment. Whether a document is privileged depends on facts the model cannot know: who was in the room, why the email was sent, what the engagement covers. A misclassified privileged document in a production is not an efficiency problem; it is a waiver problem.
Jurisdiction-specific procedure. Local rules, filing formats, and deadlines. A hallucinated deadline is a malpractice event, and models state wrong deadlines with the same fluency as right ones.
My take
My take (August 2026, Bernat Sampera)
The SERP for this query is vendor pitches and listicles, and paralegals deserve a straighter answer than either, because no profession on this site has a wider gap between the real gains and the real liability. The gains are concentrated where the work is reading: review, summarization, organization. The liability is concentrated where the work is citing: research, filings, deadlines. The tools do not tell you which mode you are in; the same chat box does both. The verdict is falsifiable: log one month of AI-suggested authorities and record the verification failure rate. If it is zero, the citation problem is solved and my 2.8 is too harsh. Published sanctions cases say it is not zero. The working rule: treat AI as a very fast reader with no license, let it read everything and cite nothing, and keep the Westlaw check for every authority exactly where it has always been.
What I would build
Every failure above is a context failure. The model does not know this matter’s verified authorities, the judge’s local rules, or what the firm has already filed, so it improvises. What I would build is a persistent matter file for AI-assisted legal work: the verified citation list, the procedural calendar with its sources, the privilege decisions already made, stored where every AI session reads them first and where anything the model cites gets checked against the verified list automatically. Legal tech is closer to this than most industries (Harvey and the research platforms hold pieces of it), but the chat tools paralegals actually use day to day still start every session knowing nothing. Until that changes, the workaround is a per-matter context document, maintained by hand, pasted into every session.
Verdict (August 2026, Bernat Sampera): AI does first-pass document review and summarization well enough to save a paralegal several hours a week, and remains unfit for anything citable: fabricated authorities are still routine in 2026, so every case, statute, and clause reference must be verified in Westlaw or Lexis before it leaves your desk. Overall: 2.8/5.