Clef · Try the live demo →
Hallucinations

Will the machine invent a clause?

That is the objection standing in front of every contract someone wants a machine to review, and it is a fair one. A hallucination is output that is linguistically flawless and factually invented — a case number that does not exist, a section that reads differently, a clause that is not in the contract. It does not stand out, because it sounds exactly like the real thing.

The cause is structural. A language model computes probabilities over sequences of tokens. It does not query a database and has no notion of whether something exists. Where information is missing, something that looks like the information appears in its place — and the more specific the question, the more plausible the invention. Legal questions are maximally specific.

How often — and why one number tells you nothing

Frequency is the second most asked question about hallucinations, and the honest answer is uncomfortable: published measurements span two orders of magnitude, because they measure different tasks.

< 1.5 %
Summarising a document you supply. Top models.
4.6–6.1 %
The 2026 frontier cohort across a mixed corpus.
> 33 %
Tasks that need several reasoning steps.
58–88 %
Specific legal queries put to general-purpose models — measured by Stanford RegLab and HAI.

From which follows the only durable statement about hallucination rates: a rate quoted without its task is not information. For your own application the number that counts is the one measured on your own documents — German contracts and manuals behave differently from the English benchmark corpora the leaderboards are built on.

What the courts have done about it

This stopped being hypothetical some time ago. The database maintained by legal researcher Damien Charlotin recorded roughly 1,490 court decisions worldwide as of May 2026 — over a thousand of them in the United States — in which a party relied on fabricated material and a court responded. Secondary counts disagree with each other, so the order of magnitude is the honest figure here rather than a number that pretends to precision.

The direction is not ambiguous. Sanctions ran to a few thousand dollars in 2023 and reach six figures now: the largest known US penalty is $110,204.38 in Couvrette v. Wisnovsky, where the briefs contained fifteen non-existent cases and eight fabricated quotations. Beyond money, 2026 has added suspensions, revoked admissions, referrals to regulators and personal liability for supervising partners.

Europe has its own record. In Germany the Kammergericht Berlin admonished a lawyer whose brief in a family-law matter cited case law that does not exist — decision of 20 November 2025, Az. 17 WF 144/25, headed "Erfundene Rechtsprechungszitate in anwaltlichem Schriftsatz". The German page sets outwhat that means under § 43a BRAO. The common finding across jurisdictions is the same: the duty to check sits with the professional, not with the tool.

What actually reduces it

It cannot be switched off; it follows from how the models are built. It can be pushed down to a level that carries document work. What does the pushing is not the choice of model but the architecture around it.

Bind every answer to a source

The model answers from passages retrieved out of your documents, not from memory, and every statement carries its citation: document, page, paragraph. What cannot be cited is not asserted.

Two passes, not one

The first pass finds the passages; the second checks the answer against those same passages. Most invented clauses die in the second pass, because they are simply not in the text.

Let the answer be missing

A system obliged to answer will invent. One allowed to say "this is not in the documents" invents far less often. That is a design decision, not a model choice.

Leave the judgement with the reader

The gain is not replacing the reading. It is finding the eight pages out of two hundred that have to be read — with the citation attached, so checking takes seconds.

What helps less than commonly claimed: a better prompt. "Do not invent anything" lowers the rate measurably and does not remove it, because the model does not know when it is inventing. An instruction in the prompt is a request; what you need is a check.

Our measurement

We did not assert this, we measured it — on three public datasets that test exactly the hard case: questions whose answer is spread across several documents, which is where simple retrieval returns part of the answer and the rest gets invented.

HotpotQA
multi-step retrieval
naive retrieval 38
ARGUS 83.1
MuSiQue
chains of three to four steps
naive retrieval 20
ARGUS 68.4
2WikiMultiHopQA
compound questions
naive retrieval 42
ARGUS 89.2

F1 score, higher is better. The full breakdown is on theARGUS page. For your own material the same caveat applies as to every leaderboard on this page: the number that counts is measured on your documents, not on a benchmark corpus.

Common questions

What is an AI hallucination?+
Output that is linguistically flawless and factually invented: a case number that does not exist, a section that reads differently, a clause that is not in the contract. It does not stand out, because it sounds exactly like the real thing.
Why do AI models hallucinate?+
A language model computes probabilities over sequences of tokens; it does not query a database and has no notion of whether something exists. Where information is missing, something that looks like the information appears in its place — and the more specific the question, the more plausible the invention.
What is the hallucination rate of AI models?+
It depends entirely on the task, which is why a single number misleads. Top models stay under 1.5 % on simple summarisation. Multi-step reasoning exceeds 33 %. On specific legal queries, Stanford RegLab and HAI measured 58 to 88 % for general-purpose models. Anyone quoting you 3 % is quoting a summarisation benchmark.
Can AI hallucinations be fixed?+
Not switched off — they follow from how the models are built. Pushed down to a level that carries document work, yes: by binding answers to cited passages, by verifying in a second pass against those same passages, and by letting the system return nothing when the documents say nothing.
How do you detect a hallucination?+
By making it checkable. If every statement carries a citation, verification is a lookup rather than a judgement call. Without citations you are left comparing the answer against the whole document, which is the work you were trying to avoid.
Does a better prompt help?+
Somewhat, not enough. "Do not invent anything" lowers the rate measurably and does not remove it, because the model does not know when it is inventing. An instruction in the prompt is a request, not a check.
Which AI hallucinates least?+
The wrong question for document work. Leaderboards rank models on one task, usually summarisation of English text, and your documents are neither. What decides the outcome is the architecture around the model — retrieval, citation, verification — and a pilot on your own material settles it faster than any ranking.

Settle it on your own documents

The durable answer to the hallucination question is a measurement on your material, not a ranking. In Discovery we take one real matter from your office and show what can be answered with a citation and what cannot.