Cases
Blog / AI and research 10 min read

Can ChatGPT Do Legal Research? What It Gets Right and Where It Fails

Last updated July 2026 · Cases

Try
ranked by relevance

How courts have ruled

Sample results, illustrative only. Informational research, not legal advice. Verify every citation.

ChatGPT can help you think about a legal problem, but it cannot be trusted to find case law. A general-purpose chatbot generates text that looks like a citation rather than retrieving one from a database, which is why it invents cases that do not exist and why hundreds of American lawyers have been caught filing them. Used as a brainstorming partner it is genuinely useful. Used as a research tool without verification it is a malpractice risk.

That is the whole answer. The rest of this piece explains why the failure happens, what the actual numbers look like as of mid-2026, what ChatGPT is legitimately good at in a legal workflow, and what to use instead when you need a real case with a real citation.

Why does ChatGPT make up cases?

Because inventing a plausible case is what a language model does when it does not know. A large language model predicts likely next words from patterns in its training data. It has seen tens of thousands of real citations, so it has learned the shape of one: an italicized party name, a volume number, a reporter abbreviation, a page, a court, a year. When you ask for authority supporting a proposition and no memorized case fits, the model does not stop and say it has nothing. It produces something with the right shape. Thompson v. Reliance Insurance Co., 412 F.3d 887 (9th Cir. 2005) looks exactly like a real case. It reads correctly. It is formatted correctly. It just does not exist.

This is a structural property of the technology, not a bug that a bigger model quietly fixes. Newer models hallucinate less often, which is arguably worse for a careless user, because a tool that is wrong one time in twenty trains you to stop checking. The confidence never changes. A fabricated citation is delivered in exactly the same tone as a real one.

How many lawyers have been sanctioned for using ChatGPT?

Enough that it is now a documented genre of judicial opinion. The best-known tracker of AI-fabricated citations in court filings, the AI Hallucination Cases database maintained by researcher Damien Charlotin, catalogued roughly 1,500 decisions worldwide by mid-2026, with more than a thousand of them in the United States. Several hundred involved licensed attorneys rather than self-represented litigants.

The landmark is Mata v. Avianca, the 2023 case in which a New York lawyer filed a brief citing six nonexistent decisions that ChatGPT had produced, then doubled down when challenged by asking ChatGPT whether the cases were real (it said yes). Judge Castel imposed a 5,000 dollar sanction and required letters to every judge falsely named as the author of a fabricated opinion. Sanctions since have escalated. In 2026 a federal court in Oregon fined two lawyers a reported 110,000 dollars combined over 23 fabricated citations, the largest such penalty reported in the United States to date.

The pattern in these opinions is consistent, and it is worth internalizing: the sanction is rarely for using AI. It is for not checking. Courts have said plainly that the tool is not the violation. Signing a filing under Rule 11 without verifying what is in it is the violation. We go deeper on this in AI hallucination in legal cases.

What is ChatGPT actually good at in legal work?

Quite a lot, as long as nothing it produces goes into a filing unverified. The reliable uses share one trait: you supply the facts and the law, and the model works with what you gave it rather than reaching into memory for authority.

  • Explaining a doctrine you half-remember. Ask it what the elements of promissory estoppel are and it will give you a serviceable answer. Confirm against a real source, but as a starting orientation it is fine.
  • Summarizing a case you paste in. Give it the opinion text and ask for the holding. It is working from your document, not its memory, so the hallucination risk drops sharply.
  • Drafting and rewriting. Tightening a paragraph, restructuring an argument, generating a first pass at a client letter, turning a rambling fact pattern into a clean chronology.
  • Brainstorming issues and counterarguments. Ask what an opponent would say to your theory. It is a decent devil's advocate.
  • Translating jargon. Turning a dense clause into plain English a client can follow.

The line is clean. Reasoning about material you provide: usually fine, still check it. Retrieving authority from its own memory: not safe, ever. The same boundary applies when you ask a chatbot to review a contract clause by clause, where purpose-built tools that read your actual document beat a model reciting from memory.

Does ChatGPT with web search fix the problem?

It helps, and it does not solve it. When a model is connected to live search, it can retrieve real pages and cite them, which removes the worst failure mode: the citation that points at nothing. That is a genuine improvement, and it is why browsing-enabled assistants are more useful for legal questions than the offline versions were.

Three problems survive. First, retrieval quality. Web search finds pages that rank, not the case that controls your jurisdiction, and a lot of what ranks for legal questions is secondary commentary rather than the opinion itself. Second, currency. A real case that has been overruled or superseded is still a real case, and a web-connected model will happily cite it because nothing in its pipeline runs a citator. Third, jurisdiction. The model has no reliable sense of whether the authority it found binds your court or is a persuasive decision from three states away, which is the distinction that decides whether the case is worth anything to you. Our guide to binding versus persuasive precedent covers why that distinction matters more than most non-lawyers assume.

So a browsing chatbot moves you from "this case may not exist" to "this case exists, but I have no idea if it is good law or if it binds you." That is progress. It is not research.

How reliable are purpose-built AI legal research tools?

Better, and not perfect, and anyone who tells you otherwise is selling something. A widely cited 2024 Stanford study tested the AI research products from the major legal publishers and found that while they hallucinated substantially less than general chatbots, they still produced incorrect or unsupported statements at a meaningful rate. The honest summary is that a purpose-built tool grounded in a real case law database is a large improvement over a general chatbot and still requires you to read the case.

The architectural difference is what matters. A general chatbot generates a citation from memory. A grounded research tool retrieves opinions from an actual database and shows you which ones it used, so every result has a citation that points at a real document you can open. That does not make the summary right. It makes the summary checkable, which is the property that keeps you out of the sanctions database.

This is the design principle behind verifiable citations in Cases. You ask your question in plain English, you get on-point precedents back as headnote cards with a plain-English summary and the holding, and every card carries a real citation you confirm in the official reporter. The tool gets you to the case. You remain the one who decides it says what you think it says.

How to use AI for legal research without getting sanctioned

A short discipline covers nearly all of it.

Never cite anything you have not opened. Not the summary, not the quote, not the pin cite. Open the opinion. If you cannot find the opinion, the case does not go in the brief. This one rule would have prevented essentially every sanction in the database.

Verify the citation at the source. Confirm the case in the official reporter, and confirm the quote actually appears on the page you cite. Fabricated quotes from real cases are now a common variant, and they are harder to catch than fabricated cases. A case citation checker handles the first half of that check, telling you whether the citation points at an opinion that exists at all.

Run a citator. A real case can be overruled, superseded, or quietly gutted. No general chatbot checks this. See how to check whether a case is still good law.

Confirm the jurisdiction binds you. A perfect case from the wrong state is a persuasive footnote at best. Start from the courts that actually control your matter, which is what our case law search by state guides are organized around.

Never delegate the signature. Rule 11 attaches to the human who signs. Courts have been unambiguous that "the AI produced it" is not a defense, and several sanctioned lawyers found that out expensively.

So, can ChatGPT do legal research?

It can do the thinking parts. It cannot do the finding part, because it does not retrieve, it predicts. Treat it as a very fast, very well-read associate with a serious and undisclosed habit of making things up when it does not know, and you will use it about right: for structure, for explanation, for drafting, for arguing with. When you need a case that exists, in a jurisdiction that binds you, that is still good law, use a tool that retrieves from a real database and hands you a citation to check.

Cases is informational research, not legal advice. Every result comes with a real citation precisely so you can confirm it yourself, because in the end the only verification that counts is the one done by the lawyer who signs the brief.

Search case law in plain English

Ask a legal question the way you would say it out loud and get on-point precedents with plain-English summaries, holdings, and citations you can check. Informational research, not legal advice.