Real Cases, Real Numbers

AI Misinformation: The $110,000 Mistake and the Security Risk Everyone Got Wrong

In May 2026, a federal judge in Oregon issued the largest AI hallucination penalty in American legal history — $110,000, for 23 fabricated citations and eight invented quotations. It's the newest entry in a pattern nearly 900 cases deep. And OWASP's own security practitioners ranked this exact risk near the bottom of their list — until the real evidence proved them wrong.

$110K
Largest AI hallucination fine in US legal history, May 2026
~900
Documented AI hallucinations in US court filings since 2023
#1
Widest gap between practitioner vote and real incident data, any category

The Fine That Broke the Record

In May 2026, a federal judge in Oregon fined two lawyers a combined $110,000 for submitting 23 fabricated citations and eight invented quotations in court filings — the largest AI hallucination penalty in American legal history to date. Not one bad citation slipping through. Twenty-three fabricated cases and eight invented quotes, treated by the court as attorneys standing by fictional legal authority as if it were real.

This wasn't an isolated event. It's the latest, largest data point in a pattern that's been accelerating since 2023 — approximately 900 AI hallucinations have now been documented in US court filings alone.

Where It Started: The Case Every Lawyer Now References

The origin story is Mata v. Avianca. In 2023, New York attorney Steven Schwartz used ChatGPT to help research a personal injury claim against a Colombian airline. The tool generated several supporting case citations — Martinez v. Delta Air Lines, Zicherman v. Korean Air Lines, and others — that looked completely legitimate. Schwartz asked the chatbot directly whether the cases were real. It confirmed they were.

None of them existed. Some misidentified judges; some involved airlines that had never been party to any such case. Judge Kevin Castel didn't mince words in his ruling: the attorneys had "abandoned their responsibilities when they submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT, then continued to stand by the fake opinions after judicial orders called their existence into question." Schwartz and his colleague were fined $5,000.

At the time, it read as an outlier — a cautionary tale about one lawyer's carelessness. Three years and roughly 900 documented cases later, it reads as the first data point in a trend.

The Ruling That Changed Who's Responsible

The more important legal precedent isn't about lawyers using AI carelessly — it's about who's on the hook when a company's own AI misleads a customer directly. In Moffatt v. Air Canada, decided by a British Columbia tribunal in early 2024, a customer relied on the airline's own website chatbot for information about its bereavement fare policy. The chatbot gave him false information. When Air Canada was sued, the airline's defense was genuinely remarkable: it argued the chatbot was "a separate legal entity that is responsible for its own actions."

The tribunal's actual ruling

The tribunal rejected that defense outright. Air Canada was held liable for what its own chatbot told a customer — full stop. The presiding judge's own words in the ruling are worth reading directly: "As this case has unfortunately made clear, generative AI is still no substitute for the professional expertise that the justice system requires... Competence in the selection and use of any technology tools, including those powered by AI, is critical." A company doesn't get to disclaim its own AI's mistakes as someone — or something — else's problem.

The Pattern Hasn't Slowed Down

Every subsequent case follows the same basic shape, and the consequences have only gotten more serious. A British Columbia lawyer was formally reprimanded in February 2024 for the identical mistake — inserting fake, ChatGPT-invented cases into court filings, with the presiding judge noting he didn't believe there was intent to deceive, but was "troubled all the same." In February 2026, in US v. Heppner, a federal judge went further than any prior ruling — finding that documents generated using consumer AI tools aren't protected by attorney-client privilege or work-product doctrine at all, adding a whole new category of exposure on top of the sanctions risk already established. Then, three months later, the Oregon case broke every record that came before it.

The trend line is not subtle: more cases, larger fines, broader legal consequences, and it's still accelerating.

Why This Is a Security Risk, Not Just an Accuracy Problem

OWASP's Top 10 for LLM Applications 2026 formally classifies misinformation as LLM07 — one of the ten most critical risks in LLM-powered systems. That's a deliberate framing choice worth understanding: an AI producing confidently wrong output stops being an accuracy issue and becomes a security failure the moment that output drives a real decision, a tool call, or an automated workflow. Steven Schwartz's fake citations weren't a security breach in the traditional sense — nothing was hacked. But the outcome was functionally identical to one: false information, generated with total confidence, was trusted and acted on, with real financial and professional consequences.

The real story behind this risk's ranking — and why it matters

OWASP's 2026 edition was the first to test practitioner opinion against real incident data — 7,714 real cases, 6,639 detailed enough to classify. Practitioners voted misinformation near the bottom of the list. The actual incident record placed it near the top — the widest gap, in the direction that made the risk look more serious than experts assumed, of any category on the list. OWASP's own explanation is straightforward: when a model's fluent, confident-sounding output drives a decision or a tool call, a wrong answer becomes a wrong action — and the real-world record shows that happening more often than the security community had been assuming.

Why This Gets More Dangerous as AI Gets More Autonomous

Every case above involves a human still reading the AI's output before acting on it — a lawyer filing a brief, a customer reading a chatbot's answer. That's actually the less dangerous version of this risk. The more serious version is what happens when an autonomous AI agent generates a wrong answer and then acts on it directly, with no human review step in between at all — approving a transaction based on a hallucinated policy detail, or calling the wrong tool because it confidently misremembered what a previous step actually returned. A misinformation problem in a chatbot produces an embarrassing conversation. The identical failure inside an autonomous agent can produce a real, unauthorized business action.

How to Actually Defend Against This

Frequently Asked Questions

What is the largest AI hallucination fine in US legal history?
In May 2026, a federal judge in Oregon fined two lawyers a combined $110,000 for submitting 23 fabricated citations and eight invented quotations generated by AI — the largest AI hallucination penalty in American legal history to date, and part of a documented pattern of roughly 900 AI hallucinations found in US court filings since 2023.
Can a company be held legally liable for its own AI chatbot's false statements?
Yes — this was directly tested and settled in Moffatt v. Air Canada (2024). Air Canada's chatbot gave a customer false information about its bereavement fare policy. The airline argued the chatbot was a separate legal entity responsible for its own words. A tribunal rejected that defense outright and held Air Canada liable, establishing that a company is responsible for what its own AI tells customers.
Why does OWASP treat misinformation as a security risk, not just an accuracy problem?
Because the moment a model's confidently wrong output drives a real decision, a tool call, or an automated workflow, an accuracy failure becomes a security failure — the model didn't need to be hacked or manipulated, it just needed to be trusted. OWASP's real incident data pushed this risk up sharply in its 2026 ranking specifically because that pattern shows up more often in production systems than practitioners had assumed.
Did OWASP practitioners underestimate this risk?
Yes, measurably. In OWASP's 2026 Top 10 for LLM Applications, practitioners voted misinformation near the bottom of the list. But when OWASP checked that vote against a real corpus of 7,714 incidents, misinformation showed the widest gap between what was predicted and what the evidence actually showed — in the direction that made the risk look more serious, not less.
Is this risk unique to lawyers and legal filings?
No — legal filings simply produce the clearest, most publicly documented paper trail, since fabricated citations are easy to verify and courts formally sanction the failures. The identical underlying pattern — confident, plausible-sounding, false AI output getting trusted and acted on — shows up anywhere an AI's claims aren't independently checked before being relied on, including medical, financial, and autonomous-agent contexts where the consequences are often just as serious but far less visible from the outside.

Related Reading