Buyer Behavior Science

The Credence-Good Problem in an AI World

When a buyer can't judge your work until it's too late, AI answers aren't closing that old trust gap, they're moving it, and the habit of checking a second source is fading right when it matters most.

Original research by Chandranshu Kumar, Founder, Raveneye Global. Published 2026-07-28. · 11 min read

Part of Choice Science in the Insights library.

Abstract

Economists have known for fifty years that services like auto repair, medicine, and law are "credence goods," work whose quality a buyer often can't judge even after paying for it. The evidence gathered here shows AI answer engines have stepped directly into that old gap: chatbot use is climbing fast, but trust in what they say still lags trust in the humans they're displacing, the AI answers themselves are not yet reliable in the highest-stakes categories, and the habit of clicking through to a second source, the very behavior that used to let a buyer catch a bad answer, is measurably declining. Nothing in the record says the old asymmetry is closing. It looks like it's relocating.

49% of US adults now use AI chatbots, up from 33% a year earlier; ChatGPT use specifically rose from 18% to 44% Pew Research Center, June 2026
8% vs 15% share of Google result-page visits ending in a click to a traditional web link, with an AI summary present versus without one Pew Research Center, July 2025
68.01% of US Google searches in early 2026 ended with no click to any external site at all SparkToro, 2026
90% vs 76% diagnostic-reasoning accuracy of an LLM used alone, versus physicians using that same LLM as a decision aid (barely above the 74% they scored without it) Goh et al., JAMA Network Open, October 2024
~33% of queries on which Westlaw's AI-Assisted Research tool hallucinated in independent testing, roughly double the rate measured for Lexis+ AI, a competing legal-research tool Stanford RegLab, 2024
How the market is evolving

The shift is real and it is uneven across markets. In the US, chatbot use jumped from 33% of adults to 49% in a single year, with ChatGPT specifically more than doubling from 18% to 44%, according to Pew Research Center's 2026 survey, and a fifth of chatbot users already ask for medical advice or diet and fitness information, both classic credence-good territory. Globally the pattern has a different shape: the Reuters Institute's Digital News Report 2026 found weekly AI chatbot use for news rose from 7% to 10% year on year, with 17% of 18-to-24-year-olds already weekly users, but that growth concentrates in developing markets across Asia and Africa and in countries like South Korea and Greece, where usage roughly doubled, while the US and UK stayed flat and low, around 4%. Underneath both trends sits a quieter and larger shift in how people search at all: SparkToro's 2026 analysis of US Google data found 68.01% of searches now end with no click to any external site, and Pew's own browsing-panel work found that when an AI summary appears on a results page, only 8% of visits end in a click to a traditional link, versus 15% when no summary is present. The surface people used to search across is being replaced, for a large and growing share of queries, by a single synthesized answer.

What it does to buyers

That single-answer habit is changing how buyers discover, trust, and choose, but trust is not keeping pace with use. Only 20% of people trust news delivered through AI chatbots, against 37% overall trust in news, per the Reuters Institute, though regular chatbot users trust it far more (44%) than the general population, and the UK records the lowest AI-chatbot trust of any market measured, just 6%. Pew's health-information research shows the same split inside a single category: among adults who get health information from chatbots, slightly more call it inaccurate (23%) than call it highly accurate (18%), yet frequent users report far higher perceived accuracy (45%) than infrequent users (13%). That gap tracks habit, not verified accuracy, which is the classic shape of an overconfidence loop. It runs alongside a separate and older behavioral quirk: people readily trust algorithmic advice in one-off decisions (what researchers call algorithm appreciation), but professionals who make similar judgments routinely tend to underuse it, and that underuse costs them. In a JAMA randomized trial, physicians given LLM access scored an average of 76% on diagnostic reasoning versus 74% without it, a difference close to noise, while the same LLM used alone, with no physician in the loop, scored roughly 90%. The professionals in the loop did not capture the tool's advantage. For a buyer choosing a doctor, lawyer, or mechanic, the practical result is that trust in the machine is rising fast, use of the machine as an aid by the expert is uneven, and neither one is reliably tracking what actually works.

What it means for the attention terrain

This is exactly where attention and the ROI of being findable are moving, and it is also exactly the gap in the evidence. Nobody has yet built a direct census of how often a buyer consults an AI answer at the specific moment of a credence-good decision, choosing a surgeon, a lawyer, a mechanic, a financial advisor, as opposed to general information seeking. What the record does show is the terrain shifting under that decision: fewer clicks to a second source (Pew, SparkToro), rising habitual trust in the first answer given (Pew, Reuters Institute), and, in the one healthcare field experiment available so far, AI advice measurably shifting who holds authority in the room, physician or patient, before a human even weighs in. For a business that sells something a buyer can't fully judge in advance, that means the old game of being findable across many surfaces and letting the buyer cross-check you is being replaced by a narrower, higher-stakes moment: whatever the engine says when it's asked. Mapping exactly where that moment lives for a given market, and what it currently says, is the open question a Visibility Corpus is built to answer, and it has not yet been fielded for this category.

The data, in one read

Diagnostic Accuracy: Physician vs. LLM
Physicians alone (no AI)
74%
Physicians + LLM as aid
76%
LLM alone (no physician)
90%
establishedIn Goh et al.'s randomized clinical trial, the LLM scored 90% on its own, but physicians using it as a decision aid scored only 76%, barely above the 74% they scored with no AI at all, meaning the humans in the loop never captured the tool's advantage. Source: Goh et al., JAMA Network Open, October 2024.

The trust problem is older than the chatbot

Economists gave this problem a name in 1973, long before anyone imagined asking a machine for a second opinion. Darby and Karni defined a credence good as one where the seller, a doctor, a mechanic, a lawyer, knows more about what you actually need than you do, and where you often can't verify the quality of what you got even after you've paid for it. Dulleck and Kerschbamer's 2006 survey of that literature identified the three forces that keep this kind of market honest when it works at all: the seller's liability if something goes wrong, how verifiable the work actually is after the fact, and how much competition and reputation are on the line. Take any of the three away and the seller has room to exploit the gap. This is not a theoretical risk. A field experiment using real taxi rides in Athens found drivers cheated passengers in exactly the ways the theory predicts, taking longer routes with riders who didn't know the area and manipulating fares with riders who didn't know the tariff system. Closer to home for most buyers, AAA's 2016 national survey found two out of three US drivers didn't trust auto repair shops in general, citing recommendations for unnecessary work (76%) and overcharging (73%) as the top reasons, and roughly 75 million motorists said they'd never found a shop they trusted at all. This is the market AI answer engines are now entering, not a clean slate.

Two out of three American drivers already didn't trust the person diagnosing their car, years before a single AI chatbot entered the picture.

Adding AI to the room doesn't automatically close the gap

There are two well-documented, and slightly contradictory, human responses to algorithmic advice. In one-off decisions like forecasts or estimates, people tend to trust an algorithm's answer more than another person's, what researchers Logg, Minson, and Moore called algorithm appreciation. But professionals who make similar judgments routinely show the opposite pattern: they underuse algorithmic advice relative to laypeople, and that underuse measurably hurts their accuracy, an effect the same research calls an expertise paradox. A 2024 randomized clinical trial in JAMA Network Open put a version of this to the test with real diagnostic reasoning. Physicians given access to an LLM alongside their usual resources scored 76% on average, against 74% for physicians using conventional resources alone, a gap close enough to be noise. Used completely on its own, with no physician involved, the same LLM scored roughly 90%. The tool was meaningfully better than the humans using it as an aid, and the humans did not capture that advantage. For a buyer who assumes their doctor "already uses AI to check," the answer from this trial is that having the tool in the room is not the same as the expert actually deferring to it when it's right.

The AI scored 90 on its own. Add a physician to the loop and the combined score falls to 76, barely above the 74 physicians scored with no AI at all.

Trust is rising faster than anyone has verified the accuracy

Usage and trust are both climbing, but they are not climbing together. Pew's 2026 survey found 49% of US adults now use AI chatbots, up from 33% the year before, and a fifth of users already turn to them for medical advice or diet and fitness information. A companion Pew study of health-information seekers found a near-even split on accuracy, 23% call chatbot health information inaccurate against 18% who call it highly accurate, but frequent users report far higher perceived accuracy (45%) than infrequent users (13%). That is a habit signal, not an accuracy signal: the more someone uses the tool, the more they trust it, independent of whether the tool actually got better. The Reuters Institute's Digital News Report 2026 finds a similar split at the level of trust itself: only 20% of people trust news delivered via AI chatbots against 37% overall trust in news, though trust climbs to 44% among regular chatbot users, and the UK posts the lowest AI-chatbot trust of any market measured, just 6%. Read together, the picture is that trust in AI answers is rising, usage is rising faster in most markets, but nothing here shows accuracy rising alongside either one.

The second opinion is disappearing

The credence-good problem has always had one partial defense: a buyer who gets a second opinion, or clicks through to check a claim, closes some of the information gap on their own. That behavior is what appears to be eroding fastest. In a Pew browsing-panel study of 900 US adults, when an AI summary appeared on a Google results page, only 8% of visits ended in a click to a traditional web link, versus 15% on pages without a summary; only 1% of visits clicked through to one of the summary's own cited sources, and users ended their entire browsing session after 26% of AI-summary pages, compared with 16% of standard pages. SparkToro's 2026 analysis of Similarweb search data puts a number on how far that has spread across search generally, 68.01% of US Google searches in early 2026 ended without any click to an external site, with mobile zero-click behavior near 77%. SparkToro is upfront that its figure isn't strictly comparable to earlier years because the underlying data panel changed, so it sits at an emerging tier rather than an established one, but it points the same direction as Pew's controlled study: whatever a buyer is told first, they are markedly less likely to go verify it against anything else.

When an AI summary appears, roughly ninety-nine visits out of a hundred never open the source it cited.

The answers themselves are not solved yet, especially where the stakes are highest

None of this would matter as much if the answers were reliable, and in the credence-good fields tested so far, they are not. Dahl and colleagues tested general-purpose language models against a large set of verifiable legal questions and documented substantial hallucination rates that varied sharply by model. The researchers found the models frequently showed no awareness of their own error, reinforcing a wrong legal premise rather than flagging uncertainty. The exact per-model figures aren't repeated here because they couldn't be independently reconfirmed against the original study text, though the core finding, that legal answers from general-purpose models hallucinate at meaningful, model-dependent rates, is well established in the literature. A follow-up Stanford RegLab study went further, testing commercial, retrieval-augmented legal research tools marketed specifically as hallucination-resistant. Westlaw's AI-Assisted Research product hallucinated on roughly a third of tested queries, about double the rate measured for Lexis+ AI, a competing product. These are paid, professional-grade tools built for exactly this use case, and the problem is not solved in them either. Law is one of the clearest credence goods there is; a buyer usually cannot judge a legal answer's correctness at all, before or after. That is precisely the domain where an unreliable intermediary does the most damage.

Regulators are writing the rulebook while the market runs ahead of them

Enforcement exists for the clearest cases of overclaiming. The FTC finalized a 5-0 order against DoNotPay, which had marketed itself as "the world's first robot lawyer," after finding the company never tested whether its chatbot's legal document generation and advice actually matched a human lawyer's standard, and never hired or retained attorneys to check its accuracy. The order required $193,000 in monetary relief, notice to subscribers from 2021 through 2023, and bars the company from claiming its service performs like a real lawyer without evidence. But a general accuracy standard for AI answers, one that would apply beyond a single company's specific false claim, is still being drafted. Following a December 2025 executive order, the FTC opened a public comment period, closing at the end of July 2026, on a policy statement addressing how existing unfair-or-deceptive-practices law applies to the accuracy of AI outputs generally. That comment period is a real, ongoing process rather than a finished standard, and it's held here at a lower confidence level pending independent confirmation, but its direction is consistent with the DoNotPay action: the enforcement tools exist, the general standard for AI accuracy claims does not yet.

Who actually holds the authority in the room is already shifting

The clearest early evidence that AI advice changes the relationship itself, not just the information in it, comes from a July 2026 field experiment at a Chinese hospital, currently a preprint that has not completed peer review. Patients given AI chatbot access before their visits received advice that was systematically directional: it routinely cautioned against medications, especially traditional medicine and antibiotics, while issuing clean, confident recommendations for diagnostic testing. That pattern measurably shifted physician behavior, fewer prescriptions and more ordered tests, especially among physicians who were more receptive to patient input, and it reduced patient compliance and satisfaction. This is a single, unreviewed study and should be read as early and directional rather than settled. But it is the first piece of evidence in this record showing that AI advice doesn't just add a new information source to a credence-good decision, it can move who has the final say between buyer and expert before either one has fully weighed in.

Where this leaves the buyer, and the business they're trying to judge

Put together, nothing here shows AI answers closing the old credence-good gap. Trust in AI answers is rising, but it tracks habit more than verified accuracy. The tools themselves are not yet reliable in the highest-stakes categories tested. And the behavior that used to let a buyer catch a bad answer, clicking through to check it against something else, is declining fastest in exactly the moments an AI summary appears. The buyer's question used to be "can I trust this doctor, lawyer, or mechanic." Increasingly it is "can I trust this answer," asked once, with a shrinking chance anyone checks a second source. For a business whose service a buyer cannot fully judge in advance, whether that's a clinic, a law practice, a repair shop, or a financial advisory firm, that single moment, whatever an engine says when asked, is doing more of the work that used to be spread across a whole search session. Nobody has yet mapped, for any given market, exactly where that moment lives or what it currently says. That is the frontier a Visibility Corpus is built to chart.

The evidence, in numbers

Key findings, dated and sourced

  • Credence goods were formally defined in economics as goods or services where an expert seller knows more about the quality or treatment a buyer needs than the buyer does, and where the buyer often cannot verify quality even after consumption. This founding taxonomy underlies the entire credence-good literature.

    established University of Chicago Press / Journal of Law and Economics, Darby, M.R. and Karni, E., 'Free Competition and the Optimal Amount of Fraud', Journal of Law and Economics, 16(1), pp. 67-88 (1973)

  • The canonical survey of credence-goods economics establishes that fraud, overtreatment, and undertreatment in expert markets (doctors, mechanics, computer specialists) is theoretically driven by three levers: seller liability, verifiability of the service performed, and market competition and reputation.

    established American Economic Association, Dulleck, U. and Kerschbamer, R., 'On Doctors, Mechanics, and Computer Specialists: The Economics of Credence Goods', Journal of Economic Literature, 44(1), pp. 5-42 (2006)

  • A field experiment using real taxi rides in Athens found drivers cheat passengers in systematic, information-dependent ways: passengers with weaker route knowledge were taken on longer detours, and asymmetric knowledge of the local tariff system led to manipulated bills.

    established Oxford University Press / Review of Economic Studies, Balafoutas, L., Beck, A., Kerschbamer, R. and Sutter, M., 'What Drives Taxi Drivers? A Field Experiment on Fraud in a Market for Credence Goods', The Review of Economic Studies, 80(3), pp. 876-891 (2013)

  • Two out of three US drivers said they do not trust auto repair shops in general, citing recommendation of unnecessary services (76%) and overcharging (73%) as top reasons; roughly a third of motorists, about 75 million, had never found a repair shop they trusted at all.

    established AAA (American Automobile Association), AAA national survey on auto repair trust (2016-12)

  • In single-encounter decisions like forecasts or estimates, laypeople rely more on algorithmic advice than on advice from another person (algorithm appreciation). This reverses with expertise: professionals who make similar judgments routinely relied on algorithmic advice less than laypeople did, and this underreliance measurably hurt their accuracy.

    established Elsevier / Organizational Behavior and Human Decision Processes, Logg, J.M., Minson, J.A. and Moore, D.A., 'Algorithm Appreciation: People Prefer Algorithmic to Human Judgment', Organizational Behavior and Human Decision Processes, 151, pp. 90-103 (2019)

  • 49% of US adults report using AI chatbots, up from 33% in 2024, and 44% specifically use ChatGPT (up from 18% in 2023). Among chatbot users, 20% report using them for medical advice and 20% for diet or fitness information.

    established Pew Research Center, 'Americans' Views on AI Chatbots, Smart Devices and AI's Impact' (2026-06-17)

  • Among US adults who get health information from AI chatbots, slightly more say the information is not too or not at all accurate (23%) than say it is highly accurate (18%), yet frequent chatbot users report much higher perceived accuracy (45%) than infrequent users (13%).

    established Pew Research Center, 'Health information from social media and AI rated more convenient than accurate' (survey of 5,111 US adults, fielded October 20-26, 2025) (2026-04-07)

  • Global weekly use of AI chatbots for news rose from 7% to 10% year on year, with 17% of 18-24-year-olds already using them weekly for news; growth is concentrated in developing markets across Asia and Africa and in countries like South Korea and Greece (usage roughly doubled), while the US and UK show flat, low adoption around 4%.

    established Reuters Institute for the Study of Journalism, University of Oxford, Digital News Report 2026 (survey of nearly 100,000 people across 48 markets) (2026-06-16)

  • Only 20% of respondents trust news delivered via AI chatbots, versus 37% overall trust in news; trust is markedly higher among regular chatbot users (44%) than the general population, and the UK records the lowest trust in AI-chatbot answers of any market measured (6%).

    established Reuters Institute for the Study of Journalism, University of Oxford, Digital News Report 2026 (2026-06-16)

  • In a browsing-panel study of 900 US adults (March 2025), only 8% of visits to a Google results page with an AI summary present ended in a click to a traditional web link, versus 15% on pages without one; just 1% of AI-summary visits clicked one of the summary's own cited sources, and users ended their browsing session entirely after 26% of AI-summary pages versus 16% of standard pages.

    established Pew Research Center, 'Google users are less likely to click on links when an AI summary appears in the results' (2025-07-22)

  • Analysis of US Google desktop and mobile search data for January-April 2026 found 68.01% of searches ended without any click to an external site, with mobile zero-click behavior around 77% versus roughly 50% on desktop. The authors caution the figure is not strictly comparable to prior years because underlying panel providers changed.

    emerging SparkToro, 'In 2026, Less than One Third of Google Searches Still Send a Click'

  • In a randomized clinical trial, physicians given LLM access alongside conventional resources scored an average of 76% on diagnostic reasoning vignettes versus 74% for physicians using conventional resources alone, while the LLM used alone, without any physician in the loop, scored approximately 90%.

    established American Medical Association / JAMA Network Open, Goh, E., Gallo, R., Hom, J. et al., 'Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial', JAMA Network Open, 7(10), e2440969 (2024-10)

  • Testing general-purpose LLMs against a large set of verifiable legal queries, researchers found substantial hallucination rates that varied sharply by model, and documented that models often showed no self-awareness of the error, reinforcing incorrect legal premises rather than flagging uncertainty. The precise per-model percentages are not restated here pending independent verification against the primary source; the study and its general finding are well corroborated in the literature.

    established Oxford University Press / Journal of Legal Analysis (Stanford RegLab / Stanford HAI research), Dahl, M. et al., 'Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive', Journal of Legal Analysis, 16, pp. 64-93 (2024)

  • A follow-up Stanford RegLab study of commercial, retrieval-augmented legal-AI research tools marketed as hallucination-resistant found they still hallucinate on a meaningful share of queries. Westlaw's AI-Assisted Research product hallucinated on approximately 33% of tested queries, roughly double the rate measured for Lexis+ AI.

    established Stanford RegLab / Stanford Institute for Human-Centered AI, 'Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools' (2024)

  • The FTC finalized (5-0 vote) an order against DoNotPay, which marketed itself as 'the world's first robot lawyer,' after finding the company never tested whether its chatbot's legal document generation and advice matched the standard of a human lawyer and never hired or retained attorneys to check accuracy. The order imposed $193,000 in monetary relief, required notice to subscribers from 2021-2023, and bars DoNotPay from claiming its service performs like a real lawyer absent evidence.

    established Federal Trade Commission, 'FTC Finalizes Order with DoNotPay That Prohibits Deceptive "AI Lawyer" Claims, Imposes Monetary Relief, and Requires Notice to Past Subscribers' (2025-01-16)

  • Following a December 2025 executive order, the FTC opened a public comment period (closing July 31, 2026) on a forthcoming policy statement addressing how the FTC Act's unfair-or-deceptive-practices prohibition applies to the accuracy of AI model outputs. A general federal accuracy standard for AI answers is still being built, not yet in force. This item is held at an emerging tier because it has not yet been independently reconfirmed against a live source.

    emerging Federal Trade Commission, 'FTC Seeks Public Comment on Policy Statement Addressing AI Accuracy' (2026-07)

  • A large-scale field experiment at a Chinese hospital found AI chatbot access before patient visits was directionally biased, routinely cautioning against medications (especially traditional Chinese medicine and antibiotics) while issuing clean recommendations for diagnostic testing. This shifted physician behavior (reduced prescriptions, increased diagnostic testing, especially among physicians receptive to patient input) and reduced patient compliance and satisfaction. As a July 2026 preprint it has not completed peer review.

    emerging arXiv preprint, Chen, Y., Li, H., Meng, L., Qiu, X. and Yang, Q., 'Directional AI Advice: Experimental Evidence from Healthcare', arXiv:2607.08706 (2026-07-10)

Learning outcomes

What this study teaches

  1. The moment a prospective client asks an AI engine whether they can trust you may now be the only checking they do; the old habit of clicking through to a second source is declining fastest exactly when an AI summary is present.
  2. Trust in AI answers is climbing faster than any evidence of their accuracy in high-stakes categories, which means being the source an engine actually draws from now matters more than it did when buyers still cross-checked multiple sites.
  3. Overclaiming what an AI-assisted service can do carries real regulatory risk, not just a reputational one; the FTC has already finalized a six-figure enforcement action against a company that marketed unverified AI legal advice as equivalent to a human lawyer.
  4. Even paid, professional-grade AI research tools still produce meaningfully wrong answers in high-stakes fields like law, so a business built on expert judgment should treat AI-assisted answers as a starting point for a client to verify, not a replacement for your own credibility signals.
  5. Nobody has yet mapped where your specific buyers' attention and trust-checking behavior actually sits today, across search, AI answers, reviews, and forums; that gap is the frontier worth understanding before you invest further in being found there.

Honest limits

What this does not yet settle

  • No independent census yet measures how often AI answer engines are consulted specifically at the moment of a credence-good purchase decision, choosing a surgeon, a lawyer, a mechanic, a financial advisor, as opposed to general information seeking or news use.
  • No large-sample, peer-reviewed study directly compares consumer verification behavior, such as seeking a second opinion or cross-checking a quote, before versus after AI-chatbot adoption within a single credence-good category. The closest evidence is a July 2026 healthcare working paper that is a preprint and has not been peer reviewed.
  • Legal-AI hallucination research is concentrated on case-law and citation retrieval; no equivalent controlled, published hallucination-rate study yet exists for medical-advice-only chatbots, or for auto-repair or home-services diagnoses produced by an AI engine, specifically.
  • Click-through and zero-click statistics for AI summaries come from at least three different methodologies (a Pew browsing panel, a SparkToro/Similarweb clickstream panel, and other vendor studies) that are not reconciled with each other and use different baselines. The direction, sharply fewer clicks when an AI answer is present, is well established; a single authoritative percentage is not.
  • No independent measurement yet exists of where attention and verification activity for a credence-good buyer journey actually sits today across search, AI answers, review platforms, and forums, or how that terrain is shifting. This is the exact primary-research gap a Visibility Corpus mapping is designed to fill, and it has not yet been fielded for this domain.

This is a synthesis of dated, attributed evidence, not a census. The AI-answer layer in particular has no independent, Nielsen-grade measurement yet, so readings of it are directional and named as a frontier, never presented as settled.

Straight answers

Frequently asked questions

Are AI chatbots actually replacing where people check on a doctor, lawyer, or mechanic before they hire one?

Chatbot use is climbing fast: Pew Research Center found 49% of US adults now use AI chatbots, up from 33% a year earlier, and a fifth of users already ask them for medical advice or diet and fitness information. But no one has yet built a direct census of how often a buyer consults an AI answer at the specific moment of a credence-good decision, choosing a surgeon, lawyer, mechanic, or financial advisor, as opposed to general information seeking, so that exact swap is not yet measured.

If AI answers are more accurate than what a human search used to surface, does that close the old trust gap in these industries?

No. The evidence shows trust in AI answers rising faster than any proof of their accuracy. Pew found that among people who get health information from chatbots, slightly more call it inaccurate (23%) than call it highly accurate (18%), even though frequent users report much higher perceived accuracy (45%) than infrequent users (13%), a pattern that tracks habit rather than verified quality. In law, Stanford RegLab found Westlaw's AI-Assisted Research tool hallucinated on roughly a third of tested queries, about double the rate of a competing tool, in paid, professional-grade products built for that exact use case.

Do buyers still cross-check an AI answer against a second source the way they used to click through search results?

That habit is fading right when it matters. Pew's browsing-panel study found that when an AI summary appears on a Google results page, only 8% of visits end in a click to a traditional link versus 15% when no summary is present, and just 1% of visits click through to one of the summary's own cited sources. SparkToro's 2026 analysis found 68.01% of US Google searches ended with no click to any external site at all, though SparkToro flags that figure as emerging tier because its underlying data panel changed year over year.

Does having a professional use AI as a decision aid protect a client from a bad judgment call?

Not automatically. In a JAMA Network Open randomized trial, physicians given LLM access scored 76% on average on diagnostic reasoning versus 74% without it, a difference close to noise, while the same LLM used alone, with no physician in the loop, scored roughly 90%. The physicians in the study did not capture the tool's advantage, so a client assuming their doctor already uses AI to double check isn't guaranteed the benefit that AI alone demonstrated.

Is there legal risk in a business claiming its AI-assisted service performs as well as a human expert?

Yes, and it has already been enforced. The FTC finalized a 5-0 order against DoNotPay, which marketed itself as the world's first robot lawyer, after finding the company never tested whether its chatbot's advice matched a human lawyer's standard and never had attorneys check its accuracy; the order required $193,000 in monetary relief and bars the company from making that claim without evidence. A broader federal standard for AI accuracy claims is still being drafted, the FTC opened a public comment period on the topic that closed at the end of July 2026, so the general rule isn't finalized yet even though the enforcement tools clearly exist.

Provenance

References

  1. Darby, M.R. and Karni, E. (1973). 'Free Competition and the Optimal Amount of Fraud.' Journal of Law and Economics, 16(1), 67-88. https://econ.vt.edu/content/dam/econ_vt_edu/seminars/spring-2020/02-24-20Karni.pdf
  2. Dulleck, U. and Kerschbamer, R. (2006). 'On Doctors, Mechanics, and Computer Specialists: The Economics of Credence Goods.' Journal of Economic Literature, 44(1), 5-42. https://www.aeaweb.org/articles?id=10.1257%2F002205106776162717
  3. Balafoutas, L., Beck, A., Kerschbamer, R. and Sutter, M. (2013). 'What Drives Taxi Drivers? A Field Experiment on Fraud in a Market for Credence Goods.' The Review of Economic Studies, 80(3), 876-891. https://ideas.repec.org/a/oup/restud/v80y2013i3p876-891.html
  4. AAA (American Automobile Association) (2016). National survey on auto repair trust. https://newsroom.aaa.com/2016/12/u-s-drivers-leery-auto-repair-shops/
  5. Logg, J.M., Minson, J.A. and Moore, D.A. (2019). 'Algorithm Appreciation: People Prefer Algorithmic to Human Judgment.' Organizational Behavior and Human Decision Processes, 151, 90-103. https://www.jennlogg.com/uploads/2/8/9/2/2892148/algorithm_appreciation__logg_minson_moore_2019_.pdf
  6. Pew Research Center (2026-06-17). 'Americans' Views on AI Chatbots, Smart Devices and AI's Impact.' https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact/
  7. Pew Research Center (2026-04-07). 'Health information from social media and AI rated more convenient than accurate.' https://www.pewresearch.org/science/2026/04/07/users-of-social-media-and-ai-chatbots-for-health-information-are-more-likely-to-say-they-are-convenient-than-accurate/
  8. Reuters Institute for the Study of Journalism, University of Oxford (2026-06-16). Digital News Report 2026. https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2026/dnr-executive-summary
  9. Pew Research Center (2025-07-22). 'Google users are less likely to click on links when an AI summary appears in the results.' https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
  10. SparkToro (2026). 'In 2026, Less than One Third of Google Searches Still Send a Click.' https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/
  11. Goh, E., Gallo, R., Hom, J. et al. (2024). 'Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial.' JAMA Network Open, 7(10), e2440969. https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2825399
  12. Dahl, M. et al. (2024). 'Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive.' Journal of Legal Analysis, 16, 64-93. https://dho.stanford.edu/wp-content/uploads/Hallucinations_JLA.pdf
  13. Stanford RegLab / Stanford Institute for Human-Centered AI (2024). 'Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools.' https://reglab.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools/
  14. Federal Trade Commission (2025-01-16). 'FTC Finalizes Order with DoNotPay That Prohibits Deceptive "AI Lawyer" Claims.' https://www.ftc.gov/news-events/news/press-releases/2025/02/ftc-finalizes-order-donotpay-prohibits-deceptive-ai-lawyer-claims-imposes-monetary-relief-requires
  15. Federal Trade Commission (2026-07). 'FTC Seeks Public Comment on Policy Statement Addressing AI Accuracy.' https://www.ftc.gov/news-events/news/press-releases/2026/07/ftc-seeks-public-comment-policy-statement-addressing-ai-accuracy
  16. Chen, Y., Li, H., Meng, L., Qiu, X. and Yang, Q. (2026-07-10). 'Directional AI Advice: Experimental Evidence from Healthcare.' arXiv:2607.08706. https://arxiv.org/pdf/2607.08706

Every measured figure is dated to its capture and tagged with an evidence tier. Every cited work is real and locatable. Where an engine could not be captured this round, it is named as uncaptured, not estimated. Small-sample readings are labelled as directional.

Find out where your buyers actually check before they trust you

If clients can't fully judge your work before they hire you, a clinic, a law practice, a repair shop, a financial advisory firm, you're operating inside exactly the gap this study describes, and the evidence says buyers are checking a second source less often than they used to, right at the moment they're deciding whether to trust you. We build a Visibility Corpus for your specific market: a map of where your buyers' attention and trust-checking actually happen today, across search, AI answers, review platforms, and the forums where people ask whether a business is legitimate. From that map we produce your Machine-Readiness Score, a read on where your name currently stands across that terrain, so you can see exactly where the gap between what buyers can verify and what you're able to show them is costing you the decision. You'll know precisely where you stand today and where the next work should go.