Will the machine invent a clause?
That is the objection standing in front of every contract someone wants a machine to review, and it is a fair one. A hallucination is output that is linguistically flawless and factually invented — a case number that does not exist, a section that reads differently, a clause that is not in the contract. It does not stand out, because it sounds exactly like the real thing.
The cause is structural. A language model computes probabilities over sequences of tokens. It does not query a database and has no notion of whether something exists. Where information is missing, something that looks like the information appears in its place — and the more specific the question, the more plausible the invention. Legal questions are maximally specific.
How often — and why one number tells you nothing
Frequency is the second most asked question about hallucinations, and the honest answer is uncomfortable: published measurements span two orders of magnitude, because they measure different tasks.
From which follows the only durable statement about hallucination rates: a rate quoted without its task is not information. For your own application the number that counts is the one measured on your own documents — German contracts and manuals behave differently from the English benchmark corpora the leaderboards are built on.
What the courts have done about it
This stopped being hypothetical some time ago. The database maintained by legal researcher Damien Charlotin recorded roughly 1,490 court decisions worldwide as of May 2026 — over a thousand of them in the United States — in which a party relied on fabricated material and a court responded. Secondary counts disagree with each other, so the order of magnitude is the honest figure here rather than a number that pretends to precision.
The direction is not ambiguous. Sanctions ran to a few thousand dollars in 2023 and reach six figures now: the largest known US penalty is $110,204.38 in Couvrette v. Wisnovsky, where the briefs contained fifteen non-existent cases and eight fabricated quotations. Beyond money, 2026 has added suspensions, revoked admissions, referrals to regulators and personal liability for supervising partners.
Europe has its own record. In Germany the Kammergericht Berlin admonished a lawyer whose brief in a family-law matter cited case law that does not exist — decision of 20 November 2025, Az. 17 WF 144/25, headed "Erfundene Rechtsprechungszitate in anwaltlichem Schriftsatz". The German page sets outwhat that means under § 43a BRAO. The common finding across jurisdictions is the same: the duty to check sits with the professional, not with the tool.
What actually reduces it
It cannot be switched off; it follows from how the models are built. It can be pushed down to a level that carries document work. What does the pushing is not the choice of model but the architecture around it.
The model answers from passages retrieved out of your documents, not from memory, and every statement carries its citation: document, page, paragraph. What cannot be cited is not asserted.
The first pass finds the passages; the second checks the answer against those same passages. Most invented clauses die in the second pass, because they are simply not in the text.
A system obliged to answer will invent. One allowed to say "this is not in the documents" invents far less often. That is a design decision, not a model choice.
The gain is not replacing the reading. It is finding the eight pages out of two hundred that have to be read — with the citation attached, so checking takes seconds.
What helps less than commonly claimed: a better prompt. "Do not invent anything" lowers the rate measurably and does not remove it, because the model does not know when it is inventing. An instruction in the prompt is a request; what you need is a check.
Our measurement
We did not assert this, we measured it — on three public datasets that test exactly the hard case: questions whose answer is spread across several documents, which is where simple retrieval returns part of the answer and the rest gets invented.
F1 score, higher is better. The full breakdown is on theARGUS page. For your own material the same caveat applies as to every leaderboard on this page: the number that counts is measured on your documents, not on a benchmark corpus.
Common questions
What is an AI hallucination?+
Why do AI models hallucinate?+
What is the hallucination rate of AI models?+
Can AI hallucinations be fixed?+
How do you detect a hallucination?+
Does a better prompt help?+
Which AI hallucinates least?+
Settle it on your own documents
The durable answer to the hallucination question is a measurement on your material, not a ranking. In Discovery we take one real matter from your office and show what can be answered with a citation and what cannot.