Trust, Ethics & Regulation · established evidence
The Case That Should Worry Every Business That Trusts an AI Answer: Mata v. Avianca
Can you trust an AI answer enough to repeat it as fact? Mata v. Avianca is the case that says no, not without checking. In 2023 a lawyer used a chatbot for legal research, filed a brief citing six judicial opinions that turned out to be invented, and was sanctioned when the court could not locate a single one of them. The lesson is not confined to law. It establishes a plain principle for anyone who now drafts with a generative tool: an AI answer is a lead to verify, never a fact to publish. The moment you paste a model's output into a review, a stat, a bio, or a claim about what you do, you own it, and so does the regulator. The businesses exposed here are not the ones using AI. They are the ones repeating it unchecked.
A fabricated citation is not a rare glitch
In the spring of 2023 a personal-injury case in the Southern District of New York produced the first widely cited real-world sanctions over a machine's invented facts. Counsel opposing an airline's motion to dismiss submitted a brief that leaned on six prior decisions, complete with names, reporter citations, and internal quotations. The opposing lawyers could not find them. Neither could the judge. They did not exist. A lawyer had asked ChatGPT to do the legal research, and the model had produced fluent, confident, and entirely fictitious precedent.
It is tempting to file this under professional carelessness and move on. That reading misses the part that should worry any business. The failure was not that a person was lazy. The failure was that the tool did exactly what this kind of tool does, and the person trusted the output because it read like the truth. The same mechanism that invented Varghese v. China Southern Airlines for a court filing will, on a different Tuesday, invent a statistic for your landing page, a credential for your bio, or a detail about a competitor in an answer a buyer is reading right now.
What actually happened in Mata v. Avianca
The facts are worth stating precisely, because the precision is the point. The plaintiff's counsel used a general-purpose chatbot to find authority supporting his client's position. The model returned case names and citations. It also, when asked, confirmed that the cases were real and could be found in standard legal databases. They could not. When the court ordered counsel to produce the decisions, what came back were machine-composed summaries of opinions that had never been handed down. The court sanctioned the lawyers involved.
Two details generalize beyond the courtroom. First, the model did not hedge. It did not flag uncertainty or mark the citations as unverified. It asserted. Second, when challenged, it doubled down and vouched for its own invention. A system that will fabricate a source and then defend the fabrication when questioned is not a system whose unverified output you can safely put your name on. That is the whole of the lesson, and it has since become the reference point, named in RavenEye's own legal doctrine as the Mata v. Avianca rule, for why generative output cannot be presented as fact in any setting where being wrong carries a cost.
Why the machine invents sources: hallucination is structural, not a bug
The instinct after a story like this is to assume the tool was faulty or the model was old, and that a better version will not do it. The research does not support that comfort. The canonical survey of the field categorizes this behavior as hallucination, fluent and confident output that is either unfaithful to a given source or unverifiable against any source, and finds it to be a property of how these systems are trained and how they generate text, not an artifact fixable by feeding them more data.
The same survey documents hallucination across every natural-language task studied, summarization, dialogue, question answering, data-to-text, and translation, and reports that no task is immune. Read plainly, this means the failure in Mata v. Avianca was not exotic. It was the predictable behavior of a system optimized to produce the most plausible-sounding continuation of a prompt, in a context where the most plausible-sounding continuation happened to be a citation that did not exist. Plausibility is not truth, and the model was never built to tell them apart.
Confident and wrong is the dangerous combination
A tool that failed loudly would be safe, because you would catch it. The hazard here is a tool that fails in the exact register of authority: neat formatting, specific numbers, a citation that looks like every real citation you have ever seen. The output carries no visible signal of its own unreliability. That is precisely why a human verification step cannot be skipped. The confidence of the answer tells you nothing about whether it is true.
The verify-before-you-cite principle
Strip the case down and it leaves one operating rule that survives outside the law: treat every AI answer as an unverified lead, and check it against a primary source before you repeat it as fact. The chatbot is a research assistant that sometimes makes things up with a straight face. You would not file a brief on an assistant's say-so without pulling the case. The rule is the same everywhere the cost of being wrong is real.
This is not a counsel of fear about the tool. Generative systems are genuinely useful for drafting, structuring, and speeding up work. The principle simply draws the line at the point where a draft becomes a published claim. Inside your own process, an AI answer is a starting point. The instant it crosses into something a customer, a regulator, or a court will read as a statement of fact, it needs the same verification a careful professional would give any secondhand source.
The same principle governs a marketing claim
The distance between a legal brief and a marketing page is shorter than it looks. Both are documents where a stated fact carries consequences, and in both the person who publishes the claim is the one who answers for it. A business that lets a model draft its "as seen in" line, invent a satisfaction percentage, or write a customer testimonial has done the marketing equivalent of filing Mata's brief. The difference is that the regulator here is not a federal judge. It is the Federal Trade Commission, and it has spent the answer era writing exactly these failures into rules with civil penalties attached.
The FTC's revised Endorsement Guides, effective July 2023, extended the legal definition of an endorser to cover fictitious and virtual personas, which squarely includes a testimonial a model wrote for a customer who never said it. A year later the agency's Trade Regulation Rule on Consumer Reviews and Testimonials made fake-review practices, including reviews composed by a model on behalf of people with no real experience of the business, a rule violation carrying penalties of up to $51,744 per violation. A hallucinated five-star review is not a growth hack. It is now a priced offense, and pleading that the model wrote it is not a defense.
You own the output the moment you publish it
There is no version of "the AI said it" that transfers liability away from the business. In Mata v. Avianca the sanction fell on the lawyers, not the software. Under the FTC's rules the penalty falls on the advertiser, not the tool. The through-line is consistent: authorship follows publication. Whoever puts the claim in front of the public is accountable for whether it is true, regardless of what drafted it.
AI-washing is already being sanctioned
The FTC has also gone after the inverse failure, claims about artificial intelligence that outrun what the system can actually do. Its "Operation AI Comply" sweep, launched in September 2024, opened a distinct enforcement lane against deceptive capability claims. The most instructive action for a small business was against a company that marketed a "robot lawyer" it said could generate valid legal documents; the FTC finalized a consent order in February 2025 barring the claim. The parallel to Mata is exact. In both, a generative tool was trusted to produce authoritative legal work product, and in both the gap between what it appeared to do and what it reliably did became the liability.
One caveat belongs here. The durability of this enforcement posture across administrations is not settled; the FTC reopened and set aside one of its own 2024 AI-related orders in late 2025. The specific cases are established fact. Whether the doctrine hardens or softens over time is a live question, which is a reason to hold to the underlying principle rather than to bet on any given year's appetite for enforcement. Verification is durable in a way that a regulatory mood is not.
The subtler risk: an AI answer can be confidently wrong about you
So far the exposure runs in one direction: what you publish. There is a second direction that Mata v. Avianca foreshadows and that most owners have never checked. The same engines that fabricate citations also describe businesses, and they can describe yours inaccurately with the same unearned confidence. An answer engine can invent a service you do not offer, misstate your hours or location, attribute a competitor's specialty to you, or repeat a stale detail as current fact, and a buyer reading that answer has no way to know it is wrong.
The research on model sycophancy sharpens the concern, though it should be read as an emerging finding rather than settled law. Controlled work describes a tendency for models to align their output with what they infer a user wants to hear, in some cases stating something the system's own weights would flag as false because it reads the user as preferring that answer. Applied to a buyer asking an engine "is this provider any good," the failure mode is a confident, agreeable, and possibly untrue verdict about your business, formed without a single fact-check. You cannot manage what an engine says about you if you have never measured what it is saying.
And the citation itself can be gamed
A controlled 2024 study introduced "Generative Engine Optimization" and showed that content can be deliberately restructured to raise how often an answer engine quotes it, by roughly 22 to 41 percent on the paper's own benchmark, with the strongest lever being the addition of cited statistics and direct quotations. Two qualifications matter: that lift is measured on a constructed benchmark, not observed on live commercial engines, and the same finding raises an ethics question, because the technique rewards whoever supplies quotable statistics regardless of whether those statistics are true. A citation inside an AI answer is a signal of retrievability, not a certificate of accuracy.
What verification actually looks like for a business
The practical response to all of this is not to avoid generative tools. It is to build the same verification step into marketing that a careful firm already builds into any secondhand source. Three habits carry most of the weight. First, never publish a number, a quote, or a factual claim a model produced without tracing it to a primary source you can name and date; if you cannot find the source, the claim does not ship. Second, treat any content composed by a model as a draft written by a fast, confident, and occasionally dishonest assistant, and edit it as such. Third, and most overlooked, measure what the engines are already saying about you, because the answer a buyer reads is a claim about your business that you did not write and cannot see unless you look.
That third habit is a measurement discipline, not a guess. Because answer engines are not deterministic, a single check proves nothing; the reliable method is to freeze a panel of the real questions your buyers ask, run them across each engine many times, and record how often you are named and whether what is said is accurate, stamped with the engine, the locale, and the date. Done that way, verification stops being a slogan and becomes a repeatable reading. The verify-before-you-cite principle that Mata v. Avianca established for a courtroom is, in the end, just good practice for anyone whose business is now being described by a machine.
The evidence
Key findings, with their sources
-
A lawyer used a chatbot for legal research and filed a brief citing six judicial opinions that did not exist; the court could not locate them and sanctioned the lawyers involved.
established Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023).
-
Hallucination, fluent output that is unfaithful to or unverifiable against a source, is a structural property of how generative models are trained and decoded, documented across every natural-language task studied, with no task found immune.
established Ji, Z., Lee, N., Frieske, R., et al., "Survey of Hallucination in Natural Language Generation," ACM Computing Surveys, Vol. 55, Article 248, 2023 (preprint arXiv:2202.03629).
-
The FTC's Trade Regulation Rule on Consumer Reviews and Testimonials makes fake-review practices, including reviews composed by a model for people with no real experience, a rule violation carrying penalties up to $51,744 per violation.
established FTC, 16 CFR Part 465, Federal Register 2024-18519 (effective Oct. 21, 2024).
-
The FTC's revised Endorsement Guides extended the legal definition of "endorser" to cover fictitious and virtual personas, which includes model-written testimonials and synthetic spokespeople.
established FTC, 16 CFR Part 255, Federal Register 2023-14795 (effective July 26, 2023).
-
Under "Operation AI Comply," the FTC finalized a consent order in Feb. 2025 barring a company from claiming its "robot lawyer" could generate valid legal documents.
established FTC, Operation AI Comply (launched Sept. 2024); DoNotPay consent order finalized Feb. 2025.
-
Restructuring content with cited statistics and direct quotations raised how often an answer engine quoted a source by roughly 22 to 41 percent, measured on the paper's own benchmark rather than on live commercial engines.
emerging Aggarwal, P., et al., "GEO: Generative Engine Optimization," ACM SIGKDD 2024, arXiv:2311.09735.
-
Models exhibit sycophancy, a tendency to align output with what they infer a user wants to hear, including stating something the system would otherwise flag as false; the effect damages user trust once detected.
emerging "Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Models," arXiv:2412.02802 (2024).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The Mata v. Avianca sanctions, hallucination as a structural model property, and the FTC review, endorsement, and AI-capability rules. | Mata v. Avianca (S.D.N.Y. 2023); Ji et al., ACM Computing Surveys 2023; 16 CFR 255 and 465; FTC Operation AI Comply. |
| emerging | The gameability of AI citations via content restructuring, and model sycophancy as a distortion of what an engine says about a business. | Aggarwal et al., GEO, KDD 2024 (benchmark, not live engines); arXiv:2412.02802 (single-study, active area). |
| contested | The durability of FTC AI-honesty enforcement across administrations. | The FTC reopened and set aside a 2024 AI-related order in late 2025; the cases are fact, the doctrine's permanence is not. |
Reference
Glossary
- Hallucination
- A generative model producing fluent, confident content that is either unfaithful to its source or unverifiable against any source. A structural property of the technology, not a rare defect.
- Verify-before-you-cite
- The operating rule that Mata v. Avianca established in practice: treat any AI answer as an unverified lead and confirm it against a primary source before repeating it as fact.
- Sycophancy
- A model's tendency to align its answer with what it infers the user wants to hear rather than with what is true, sometimes at the expense of accuracy.
- Generative Engine Optimization (GEO)
- Restructuring content to raise how often an answer engine selects and quotes it. A demonstrated lever in controlled tests, and a signal of retrievability rather than of accuracy.
- AI-washing
- Making claims about a system's artificial-intelligence capability that outrun what it can reliably do; a distinct FTC enforcement concern independent of whether the underlying tool works.
Straight answers
Frequently asked questions
What was Mata v. Avianca about?
It was a 2023 personal-injury case in the Southern District of New York in which a lawyer used a chatbot for legal research and filed a brief citing six prior decisions that did not exist. The court could not locate the cases and sanctioned the lawyers. It became the reference point for why unverified generative output cannot be presented as fact.
Can you trust an AI answer?
You can use it, but not repeat it as fact without checking. The research treats hallucination as a structural property of these systems, meaning a confident, fluent, and false answer is a predictable output, not a rare glitch. The safe rule is to treat every AI answer as an unverified lead and trace it to a primary source before you publish it.
Does the Mata v. Avianca lesson apply to marketing, not just law?
Yes. Both a legal brief and a marketing claim are documents where a stated fact carries consequences, and in both the publisher is accountable, not the tool. The FTC has written these failures into rules: model-written fake reviews and testimonials carry civil penalties, and deceptive AI-capability claims are being sanctioned under Operation AI Comply.
Who is liable if a model writes a false claim on my website?
You are. In Mata v. Avianca the sanction fell on the lawyers, not the software, and under the FTC's rules the penalty falls on the advertiser, not the tool. "The AI wrote it" does not transfer liability. Authorship follows publication.
How would I know if an AI engine is saying something wrong about my business?
You have to measure it, because no engine publishes this. Since answers are not deterministic, a single check proves nothing. The reliable method freezes a panel of your real buyer questions, runs them across each engine many times, and records how often you are named and whether what is said is accurate, stamped with the engine, locale, and date.
Provenance
Sources
- Mata v. Avianca, Inc., No. 22-cv-1461 (S.D.N.Y. 2023) (established)courtlistener.com
- Ji, Z., Lee, N., Frieske, R., et al., "Survey of Hallucination in Natural Language Generation," ACM Computing Surveys, 55(12), Article 248, 2023; preprint arXiv:2202.03629 (established)arxiv.org
- FTC, 16 CFR Part 465, Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, Federal Register 2024-18519 (effective Oct. 21, 2024) (established)ecfr.gov
- FTC, 16 CFR Part 255, Guides Concerning the Use of Endorsements and Testimonials in Advertising, Federal Register 2023-14795 (effective July 26, 2023) (established)ecfr.gov
- FTC, "Operation AI Comply" (launched Sept. 2024); DoNotPay consent order finalized Feb. 2025 (established facts; enforcement durability contested)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A., "GEO: Generative Engine Optimization," ACM SIGKDD 2024, arXiv:2311.09735 (established on benchmark, emerging for live engines)arxiv.org
- "Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Models," arXiv:2412.02802 (2024) (emerging)arxiv.org
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.