Why AI Chatbots Make Things Up, and Say It With a Straight Face
A chatbot can invent a fact, a quote or an entire court case and present it with total confidence. Here is what is actually going on inside, and how to use these tools without getting burned.
In this story 4 sections
In 2023, a federal judge in New York was handed a legal brief that cited court decisions, complete with names, quotes and docket numbers. Several of them did not exist. A lawyer had asked ChatGPT to help with research, the chatbot had produced realistic-looking cases, and nobody had checked them. In June that year, Judge P. Kevin Castel fined two lawyers and their firm $5,000 in the case, Mata v. Avianca.
It is the most famous example of what researchers call an AI hallucination: a confident, fluent answer that is simply wrong. It is also one of the most misunderstood. The chatbot was not lying, and it was not broken in the usual sense. It was doing exactly what it was built to do.
What a chatbot is actually doing
The systems behind tools like ChatGPT are large language models. During training they read an enormous amount of text and learn one core skill: given the words so far, predict what comes next. Do that one piece at a time, very well, and you get paragraphs that read as if a knowledgeable person wrote them.
What a language model does not have, by default, is a lookup table of verified facts. It has patterns. Many of those patterns line up with reality, which is why chatbots are so often right. But when the model reaches a gap in what it learned, it does not stop. It keeps producing the most plausible-sounding continuation.
A language model is built to sound right. Being right is a separate job.
Why fake answers look so real
Court citations are a perfect trap. They follow a rigid, recognisable format: party names, a volume number, a reporter abbreviation, a page, a year. A model that has seen thousands of real citations can produce a new one that has every feature of the real thing except existence. The same goes for academic references, product specifications and quotes attributed to real people.
The confidence comes from the same place. Most of the text a model learns from is written in a confident register, so that is the register it produces. The tone of the answer tells you very little about whether the content underneath is solid.
Known vs. still being worked out
Researchers broadly agree that hallucinations come from how these models generate text, rather than from a single bug that can be patched. How best to measure them, and how far they can be reduced, is still an active area of research.
What makes it better, and what does not
| Situation | Risk of made-up answers | What helps |
|---|---|---|
| Well-known, widely written-about facts | Lower | Still worth a quick check for anything important |
| Specific citations, quotes, statistics | High | Ask for sources, then open every one |
| Very recent events | High | Use a tool that searches the web and shows its sources |
| Niche or local information | High | Treat the answer as a lead, not a finding |
Many newer chatbots can search the web or a set of documents before answering and then point to what they used. That grounding helps a lot, because the model is summarising text it has just retrieved rather than relying on memory. It does not make mistakes impossible: a model can still misread a source or attach the wrong one to a claim. The fix is the same as it was for the lawyers in New York. Open the source and read it.
How to use chatbots without getting burned
Use them for what they are good at: drafting, summarising text you supply, explaining concepts and suggesting where to look. Be most careful exactly where they sound most authoritative, with names, numbers, dates and references. If a chatbot gives you a source, the source is the answer, not the chatbot.
In this story 4 sections
Keep Reading
All storiesThe Odd List · Fridays