A citation invented by an AI model looks exactly like a real one — same format, same confident tone. Courts have sanctioned lawyers over $145,000 in the first quarter of 2026 alone for filing them. Here is what actually happened, and where cross-checking a citation against multiple AI models catches the fabrication before it reaches a filing — and where it does not.
We are not claiming 6 AI models means zero fabricated citations. If every model shares the same gap in its training data, they can all produce a similarly plausible fake, and agreement between them proves nothing. What multi-model comparison reliably catches is the far more common case: one model invents a citation the others do not recognize. See the FAQ below for the honest limit of this method.
New York, July 2026
Attorney Michael Sanders sanctioned $10,500 after filing a brief with AI-generated citations to cases that do not exist.
Sixth Circuit Court of Appeals
Two attorneys sanctioned — opposing counsel's fees plus $15,000 in punitive damages each — over more than 24 fabricated citations in an appellate brief.
Federal immigration case (habeas corpus)
Citations fabricated by an AI model were filed in a federal habeas petition; the attorney was sanctioned once the cases were shown not to exist.
Périgueux, France — December 18, 2025
First documented French case of a lawyer's filing containing fictitious references generated by AI, identified by the tribunal judiciaire.
Q1 2026, nationwide (US)
Legal research trackers (LexisNexis, Bloomberg Law) recorded roughly $145,000 in cumulative sanctions and over 700 court filings affected by AI hallucinations in a single quarter.
Sanction amounts and case counts are as reported by legal-industry trackers (LexisNexis, Bloomberg Law) and court filings as of Q1–Q2 2026; figures are cited for reference and should be verified against the primary court record before republication.
A single model has no incentive to flag the citations it invented. The fabricated case name, docket number, and holding are generated with the same fluency as a real one — there is no stylistic tell. A lawyer under deadline pressure, reading a citation that reads exactly like every other citation in the brief, has no local signal to know which one is fake.
Cross-checking the same citation across independent models changes what is visible before the citation reaches a filing: if five models cannot corroborate a case that one model just cited with confidence, that divergence is the flag a single-model workflow never produces — even though it still does not replace checking the primary source.
Multi-model comparison cannot catch a fabrication that every model makes the same way — a shared gap in training data produces shared confidence, not disagreement. It also does not replace a primary-source check: a citation that all six models "agree" exists still needs to be pulled from Westlaw, Lexis, or the court's own docket before it goes into a filing. Treat cross-model agreement as a reason to move faster to that verification step, not a reason to skip it.
Why do AI models invent legal citations?
Large language models generate text that matches the statistical shape of a real citation — case name, reporter, year, docket format — without verifying the case exists. The output reads exactly like a real citation, which is what makes it dangerous in a filing: nothing about the format signals that it was fabricated.
Does asking a more advanced model fix this?
No single model is immune. Fabricated citations have been documented across GPT-4, Claude, and other frontier models. A stronger model reduces the rate of fabrication; it does not eliminate it, and a single model has no built-in way to flag its own fabrication as uncertain.
Does using 6 AI models guarantee no fabricated citation will ever get through?
No, and we do not claim that. If a citation error is a shared blind spot — none of the models has seen the real case and all generate a similarly plausible fake — cross-checking will not catch it by itself. What cross-checking reliably catches is the much more common failure mode: one model invents a citation the others do not recognize or cannot corroborate. That divergence is the signal a single-model workflow never produces.
What should a lawyer actually do before filing?
Cross-check every AI-suggested citation against a primary legal database (Westlaw, Lexis, Légifrance, a court's own docket system) before it goes into a filing. Multi-model comparison is a fast way to flag which citations are suspect and deserve that primary-source check first — it is a triage step, not a replacement for verification.
Is this only a US problem?
No. The Périgueux tribunal case (France, December 2025) shows the same failure mode outside common-law jurisdictions — a lawyer filed AI-generated fictitious references in French civil proceedings.
Satcove — A product by Abyssal Group