Hallucination risk in professional documents, and how firms control it
A language model does not look things up. It generates text that is statistically plausible given what it has seen, which means it can produce a confident, correctly formatted, entirely invented citation. The trade calls this hallucination, and it has already embarrassed professional firms in public. In a widely reported 2023 case in the US courts, lawyers were sanctioned after filing a brief containing case citations that a chatbot had fabricated; the cases simply did not exist. Similar incidents have surfaced since, in several jurisdictions, and the pattern is always the same: model output pasted into a formal document as if it were a source.
The lesson is not "do not use AI in documents". Firms are using these tools daily for bids, reports and correspondence, and the productivity is real. The lesson is that model output is a draft, never a source, and that firms which use AI safely have specific controls between the draft and the signature.
This post sets out those controls. None of them is exotic. Most of them are the same review discipline a good firm already applies to a junior's first draft.
Why do models invent things at all?
Because generating plausible text is the whole mechanism. A model that has read thousands of legal briefs knows exactly what a citation looks like: party names, a year, a reference in the right format. When asked for one it does not have, it can produce one anyway, because producing the right-shaped text is what it does. The same applies to clause numbers, British Standards references, product certifications and financial figures.
This is an active research area, not a solved one. Model providers publish regularly on making outputs more honest and grounded (Anthropic's announcements cover their work in this area), and current models fabricate less than 2023-era ones. Less is not never. The control environment, not the model version, is what makes the risk manageable.
Which failure modes should you actually plan for?
The useful move is to stop thinking about "hallucination" as one risk and break it into the specific ways it reaches your documents. Each one has a matching control.
| Failure mode | What it looks like | The control |
|---|---|---|
| Fabricated citation | A case, standard or regulation that does not exist, correctly formatted | Every citation checked against the source before the document leaves |
| Wrong reference | A real clause or standard number attached to the wrong content | Retrieval: the model works from your verified documents, not memory |
| Confident wrong summary | A summary of a contract or report that states something the document does not say | Reviewer reads the summary against the source, not on its own |
| Invented specifics | Figures, dates, product names or capabilities filled in to complete a sentence | Numbers and names only enter via checked source material, never generated |
| Stale knowledge | Guidance or regulation that has since changed, stated as current | Date-sensitive claims verified against the current published version |
Two things about that table. First, every control is a human or a process, not a setting. Second, the left column is exactly the list of things a diligent reviewer checks in a human-written document too. AI did not create the need for review; it created a fluent drafter whose errors are harder to spot because they are never clumsy.
What does retrieval change?
It changes what the model is allowed to answer from.
A model answering from its training data is recalling an approximation. A model doing retrieval is handed your actual documents (the contract, the spec, your method statements, your previous bids) and instructed to answer from them, quoting where it found things. Fabrication drops sharply because the model is completing a reading task rather than a memory task, and crucially, the output can carry references you can click and check.
This is one of the strongest arguments for a private AI setup over a public chatbot: a system built on your own verified document library, with sources shown, rather than a general tool answering from the open internet's average opinion. It also changes which model you need, because a mid-tier model reading the right document beats a frontier model guessing.
Where does the human sign-off sit?
At the same gate it always sat: nothing leaves the firm without a named person having checked it.
The practical version for AI-assisted documents is a short pre-issue checklist. Every citation opened and confirmed real. Every number traced to a source document. Every summary spot-checked against what it summarises. The person signing is accountable for the content exactly as if they had written it, because professionally, they did. This matters most in the documents with consequences: bids, contractual correspondence, anything with a clause number in it. The approach in AI bid writing without the robot voice is built on the same principle: the model drafts, the human owns.
It is worth writing this into your AI usage rules explicitly, alongside your data handling rules (the ICO's guidance for organisations is the reference point for the data side). One line does most of the work: "model output is a draft, and drafts get checked".
Is the risk a reason to wait?
No. It is a reason to build the checking in from day one.
The firms that were embarrassed did not fail because they used AI. They failed because they skipped the review step they would never have skipped for a trainee's work. Treat the model as a fast, tireless, occasionally overconfident junior: brilliant on a first draft, never the final word on a fact. AI Metric builds document systems around exactly these gates, retrieval from verified sources and sign-off included, but the discipline costs nothing and starts today. The signature on the document is yours. Keep the checking that goes with it.