Buyer Behavior Science

The Consideration-Set Collapse

Buyers were already comparing a handful of options before any engine got involved. What a single synthesized answer removes is the visible list itself, and no one has yet measured what replaces it.

Original research by Chandranshu Kumar, Founder, Raveneye Global. Published 2026-07-28. · 9 min read

Part of Choice Science in the Insights library.

Abstract

Long before any engine summarized an answer, buyers were already comparing a small handful of options and stopping as soon as one felt good enough. Pew Research Center's 2025 measurement shows that when Google adds a synthesized answer to a search, clicks to any organic result roughly halve and sessions end without further browsing more often. What the evidence base does not yet contain, anywhere, is a direct measurement of how many alternatives a synthesized answer itself names, category by category, which means the shrinking shortlist is still being inferred from what happens around it, not observed at the point where the buyer actually decides.

7, plus or minus 2 Human working-memory capacity for holding discrete items in mind at once, the cognitive ceiling behind why shortlists were always short George A. Miller, Psychological Review, 1956
3.3 to 4.0 brands Mean size of real consumer consideration sets across three studied categories: deodorant, coffee, and frozen dinners Hauser & Wernerfelt, Journal of Consumer Research, 1990
8% vs 15% Click-through rate to an organic Google result when a synthesized summary is present versus when it is absent Pew Research Center, July 2025
approximately 68% Share of Google searches ending without any click to the open web, as of 2026 Similarweb
31% Share of US consumers who now require at least a 4.5-star rating before considering a local business, up from 17% the year before BrightLocal Local Consumer Review Survey, 2026
How the market is evolving

The surface a buyer meets when they search has been getting thinner for longer than most people assume, and the evidence for it spans well beyond the United States. Pew Research Center's 2025 measurement of US search behavior found a majority of adults, 58%, had already encountered a Google AI summary at least once in a single month, even though summaries appeared on only 18% of individual searches at that point. The Reuters Institute's Digital News Report, drawing on its global research base out of the University of Oxford, found AI chatbot use for news itself remained a minority behavior in 2025, just 7% weekly overall, rising to 15% among adults under 25, in the very first year the question was asked. Read together, these two findings describe an adoption curve still in its early, uneven stages rather than a completed shift, which matters for how urgently any individual business should be reacting versus preparing. Zero-click search, queries ending without any click to the open web at all, is reported at roughly 68% of Google searches as of 2026 by one industry tracker, up from about 45% a decade earlier, though that specific figure rests on single-vendor measurement and is tiered accordingly rather than treated as settled fact. Gartner's widely repeated 2024 forecast of a 25% drop in search volume by 2026 remains exactly that, a forecast, not something anyone has yet measured actually happening.

What it does to buyers

The mechanism behind all of this is old, not new. Herbert Simon's satisficing and George Miller's seven-item working-memory limit describe a human mind that has always compared a handful of options and stopped once one was good enough, and field data from 1990 confirms real consideration sets ran three to four brands deep even in categories with dozens of choices on the shelf. Ranked search results were already compressing attention hard on top of that baseline, with the top three positions capturing roughly two-thirds of all clicks according to aggregated Backlinko and Sistrix data. What a synthesized answer changes is not the size of the shortlist so much as its visibility: Pew's 2025 measurement shows that when a summary appears, click-through to any organic result roughly halves and a session is meaningfully more likely to end without further browsing at all. The buyer moves from picking among visible options to deciding whether to accept or push back on the one thing they were handed, and the popular belief that this makes the decision easier or better for them is not well supported by the choice-overload research once you look past its most famous, and least replicated, experiment.

What it means for the attention terrain

Put the layers together and a picture of where attention actually concentrates starts to emerge, even without a direct measurement of the answer layer itself. Attention was already pooling on the top three ranked results, per the aggregated Backlinko and Sistrix data. A rising star-rating threshold, now above 4.5 stars for nearly a third of buyers per BrightLocal, decides who even clears the gate before any ranking or answer gets involved. And a synthesized answer compresses what's left into something closer to a single accept-or-reject moment. None of the three layers has been measured together for any one category yet. That is precisely the terrain a Visibility Corpus reading is built to map: not a guess at where attention should be, but a direct read of where it currently sits across the ranked list, the review gate, and the answers engines are giving for a specific business's actual market, so that effort goes toward the surface that is genuinely reachable rather than the one that sounds most urgent.

The data, in one read

Real Consideration-Set Size, by Category (1990)
Deodorant
3.9
Coffee
4
Frozen Dinners
3.3
establishedAcross three unrelated shelf categories with dozens of options on offer, Hauser and Wernerfelt found real buyers seriously compared only three to four brands each, the baseline this whole study measures the AI-answer layer against. Source: Hauser & Wernerfelt, Journal of Consumer Research, 1990.

The shelf was never that big to begin with

It helps to start with what buyers were doing before any engine wrote them a summary. Herbert Simon's work on bounded rationality, dating back to 1947 and formalized under the name satisficing in 1956, described something simple: people search alternatives only until one clears an acceptability bar, not until they have found the objective best. George Miller's 1956 paper on working memory put a number on part of the reason why, showing that people can reliably hold roughly seven items, plus or minus two, in mind at once. Neither of these findings is about technology. They describe how a human mind handles any comparison task, with or without a screen involved.

Field research on actual consumer behavior confirms the theory holds up outside the laboratory. Hauser and Wernerfelt's 1990 study of real consideration sets found buyers comparing an average of 3.9 deodorant brands, 4.0 coffee brands, and 3.3 frozen dinner brands, in categories where dozens of options sat on the shelf. The gap between what is available and what gets seriously compared has always been wide. That is the baseline this whole study has to be measured against: the shortlist was already short.

The average shopper was never comparing thirty brands. They were comparing three or four, and stopping the moment one felt good enough.

The ranked list already did the narrowing

Before any synthesized answer entered a results page, the ranked list itself was already doing serious compression. Aggregated click data cited from large-scale studies by Backlinko and Sistrix shows the first organic position capturing somewhere between 27.6% and 39.8% of all clicks depending on the study, and the top three results together pulling in 68.7% of clicks on the page cited here. That means well over two-thirds of the attention on an ordinary results page was already concentrated on three listings before generative answers existed at all.

This matters because it sets the right comparison point. The story is not that ranked search once served buyers a wide, evenly-considered field of options and a synthesized answer suddenly narrowed it. The story is that ranked search had already narrowed the field hard, and now a second layer of narrowing sits in front of it.

What changes when the list disappears entirely

The clearest measured evidence of what a synthesized answer does to buyer behavior comes from Pew Research Center's July 2025 study of 68,879 real Google searches by 900 US adults. When a search produced an AI-written summary, users clicked through to a traditional organic result at 8%, compared with 15% when no summary appeared, roughly half the rate. Clicking a link inside the summary itself happened in just 1% of cases. And sessions ended without any further browsing 26% of the time when a summary was present, against 16% when it was not.

The same study found 58% of surveyed US adults had encountered at least one AI summary in a single month, though summaries still appeared on a minority, 18%, of individual searches at that point. Put together, this is not a fringe behavior. It describes a majority of searchers occasionally being handed one answer instead of a list, and behaving differently when that happens: fewer clicks out, more sessions ending on the spot. The mechanism the thesis describes, replacing "pick from a shelf" with "accept or reject one recommendation," is visible directly in this data, even though Pew's study measures the surrounding click behavior rather than the content of the answer itself.

Accept or reject one recommendation is a fundamentally different decision than picking from a shelf, and the click data already shows buyers behaving as if they know it.

The zero-click frontier, read carefully

The broader trend behind this is zero-click search, meaning a query that ends without any click to the open web at all. Similarweb's 2026 tracking put this at roughly 68% of Google searches, up from about 45% a decade earlier. That figure deserves a caveat: it comes from one vendor's measurement rather than an independent, cross-checked census, which is why it is tiered here as emerging rather than established, even though the direction is consistent with independent tracking elsewhere in the industry.

Gartner's often-cited 2024 forecast that traditional search volume would fall 25% by 2026 belongs in a different bucket again. It is a prediction, not a measured outcome, and it is treated that way here. The distinction matters for a business owner deciding how urgently to act: rising zero-click behavior is a real, if imperfectly measured, present-tense trend. A specific double-digit volume drop by a specific year is still an unproven forecast, four years old at the time of this writing.

Does a shorter list actually help the buyer? The evidence says: not obviously

There is a popular assumption sitting underneath a lot of AI-answer optimism: that fewer choices make for a happier, less overwhelmed buyer. The research does not support that assumption cleanly. The famous jam study from 2000, in which a 6-flavor tasting display produced a 30% purchase rate against 3% for a 24-flavor display, is the founding evidence for choice overload and remains widely cited. It has also not reliably replicated in the years since.

A far larger 2010 meta-analysis spanning 50 experiments and roughly 5,000 participants found the average effect of choice overload across all of them was close to zero. The reading is that the choice-overload story is contested at its origin and largely unsupported at scale. That has a direct implication for this thesis: a business should not assume that becoming the sole answer a synthesized response gives a buyer is automatically good for that buyer. It may just mean less information reached them, and no research yet has re-run this question specifically for a generative-answer context rather than a physical or digital shelf.

The pre-filter tightens before any answer engine gets involved

Whatever list eventually reaches a buyer, human-ranked or engine-synthesized, first passes through a filter that has quietly gotten stricter. BrightLocal's 2026 Local Consumer Review Survey of 1,002 US adults found 68% now require at least 4 stars before considering a local business, and 31% require 4.5 stars or higher, up sharply from 17% the year before. Ninety-seven percent read reviews, checking an average of six separate sites before making up their mind.

This is worth naming as its own layer of the collapse, separate from search or generative answers. If a rising share of buyers will not even put a business on their mental shortlist below a 4.5-star threshold, that filter operates upstream of whatever engine eventually surfaces or summarizes the remaining candidates. A business that clears search rankings but sits at 4.2 stars is already excluded from a meaningful and growing share of buyers before the question of AI visibility even arises.

When the algorithm curates the algorithm

A separate but related mechanism showed up in a 2026 study of a major music recommendation system, accepted at the ACM RecSys conference. Researchers found that continuously retrained ranking models fell into feedback loops that suppressed both new releases and previously unheard catalog items, across six different tested interventions in live experiments. Architectural fixes partially restored some variety but did not create genuinely new discovery.

This finding is about a recommendation system, not a text-answer engine, and the evidence does not yet extend it directly to the AI-answer layer covered by the rest of this study. It is included here because the mechanism, a system that continuously trains on its own prior output and narrows what it surfaces over time, regardless of what any individual user wants, is exactly the kind of dynamic worth watching as generative-answer engines mature and begin training on their own usage patterns. Treat it as emerging evidence and a plausible, not proven, extension of the thesis.

Where the evidence stops

The single largest gap in this research is also the most important one: no published, retrievable study measures the size or composition of the consideration set a synthesized answer itself presents. Nobody has published a count of how many alternatives ChatGPT, Perplexity, or an AI Overview typically names for a given query, broken out by category. Every figure in this study describes what happens around that moment, the clicks that don't happen, the sessions that end early, not the moment itself.

The choice-overload literature has never been re-run in a generative-answer setting, so it is unknown whether the near-zero average effect found in physical and digital shelf experiments holds when the shelf in question is a single written answer. There is no independent, cross-engine census of the AI-answer layer comparable to what Pew and the Reuters Institute have built for one slice of one search engine. And there is no longitudinal data tracking the same buyers before and after they adopted generative answers within one category, nor any published comparison of how this plays out differently for a low-stakes grocery reorder versus a high-stakes purchase like a home renovation or a piece of business software. This is a genuine frontier, not a gap that will close on its own, and any business making decisions in this space should treat the measurement itself, not just the tactics, as unfinished work.

The evidence, in numbers

Key findings, dated and sourced

  • Herbert Simon introduced satisficing, the idea that people search alternatives only until an acceptability threshold is met rather than until the optimum is found, as the core mechanism of bounded rationality.

    established Herbert A. Simon, Administrative Behavior; term coined 1956, concept originated 1947

  • Human short-term processing capacity for holding discrete items in mind is limited to roughly seven, plus or minus two, a standard building block for why comparison sets stay small even without any algorithm involved.

    established George A. Miller, Harvard University, "The Magical Number Seven, Plus or Minus Two," Psychological Review, Vol. 63, Issue 2, pp. 81-97 (1956)

  • Field data on real consumer consideration sets found mean sizes far smaller than the number of brands actually on the market: 3.9 for deodorant, 4.0 for coffee, 3.3 for frozen dinners.

    established Journal of Consumer Research, Hauser & Wernerfelt, "An Evaluation Cost Model of Consideration Sets" (1990)

  • The original jam study found a 6-flavor tasting display produced a far higher purchase rate than a 24-flavor display, the founding evidence for choice overload, but this specific finding has not reliably replicated since.

    contested Columbia University / Stanford University, Iyengar & Lepper, "When Choice is Demotivating," Journal of Personality and Social Psychology; 30% purchase rate at 6 options versus 3% at 24 options (2000)

  • A meta-analysis of 50 experiments and roughly 5,000 participants found the average effect of choice overload across studies was close to zero, challenging the popular narrative built on the jam study.

    established Journal of Consumer Research, Scheibehenne, Greifeneder & Todd, "Can There Ever Be Too Many Options? A Meta-Analytic Review" (2010)

  • When a Google search shows an AI-written summary, users click through to a traditional organic result at roughly half the rate of searches with no summary present, rarely click inside the summary itself, and end their session without further browsing more often.

    established Pew Research Center, "Do people click on links in Google AI summaries?", based on 68,879 searches from 900 US adults, collected April 7-17, 2025: 8% CTR to organic results with a summary present versus 15% without; 1% click within the summary; 26% of visits end the session with a summary present versus 16% without (2025-07-22)

  • A majority of US adults encountered a Google AI summary at least once in a single month, though summaries still appeared on a minority of individual searches.

    established Pew Research Center, "Do people click on links in Google AI summaries?": 58% of surveyed users hit at least one AI summary in March 2025; 18% of all searches in the study produced one (2025-07-22)

  • Zero-click search, queries that end without any click to the open web, has risen sharply as generative summaries have rolled out across a growing share of Google queries.

    emerging Similarweb, "Zero-Click Marketing: What the 2026 Data Means"; approximately 65-69% of Google searches end without a click to the open web as of 2026 (the source's own headline figure is 68%), up from roughly 45% a decade earlier. Two more specific sub-figures in the original research pack could not be corroborated on the cited page and have been removed; only the confirmed headline figure is retained here, at a downgraded tier reflecting single-vendor sourcing

  • Gartner formally predicted that traditional search engine volume would fall as generative AI chatbots and virtual agents substitute for search queries.

    contested Gartner, Inc., Gartner press release, corroborated via a secondary source quoting the same forecast after the primary URL returned an access error on direct verification; a forecast, not a measured outcome, so it is treated as such throughout (2024-02-19)

  • Organic click-through on a ranked results page is heavily concentrated at the top: the first position receives roughly 27.6-39.8% of clicks depending on the study, and the top 3 results together capture the clear majority of all clicks.

    established Backlinko / Sistrix, aggregated via First Page Sage, Aggregated CTR data citing a 4-million-result Backlinko study and an 80-million-keyword Sistrix study; the top-3 concentration figure confirmed on the cited page at 68.7%, correcting a lower figure in earlier drafts (2023-2025)

  • Use of AI chatbots for news remains a minority behavior overall, concentrated among younger users, in the first year this question was measured.

    established Reuters Institute for the Study of Journalism, University of Oxford, Digital News Report 2025: 7% weekly use overall; 15% among under-25s (2025-06)

  • Consumer star-rating thresholds for even considering a local business have tightened sharply year over year, and most consumers now check several review sites before deciding.

    emerging BrightLocal, Local Consumer Review Survey 2026, 1,002 US adults: 68% require at least 4 stars; 31% require at least 4.5 stars, up from 17% the prior year; 97% read reviews; average of 6 sites checked

  • A live industry study of a major music recommendation system found continuously retrained ranking models fall into feedback loops that suppress both new releases and previously unheard catalog items; interventions partially restored variety but did not create new discovery.

    emerging YouTube / Google research team, "Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music," arXiv:2607.23749; documented feedback-loop suppression across six tested interventions in live A/B tests, accepted at ACM RecSys 2026 (2026-07)

Learning outcomes

What this study teaches

  1. Your buyers were never comparing every option on the market. Established research puts the real number at three to seven items even in ordinary, unmediated shopping, so the real question is not whether comparison is shrinking but what happens when the visible list disappears entirely.
  2. A ranking on a results page and a mention inside a synthesized answer are two different forms of visibility, measured differently and won differently. Pew's data shows the presence of a summary changes click behavior by roughly half, which means winning the old kind of visibility does not guarantee the new kind.
  3. The research literature does not support the comfortable assumption that fewer choices make for a happier or better-served buyer. Being the one answer an engine gives is a responsibility to earn honestly, not a shortcut to assume is automatically good for the person on the other end.
  4. Review thresholds are rising fast and now function as a pre-filter that decides whether a business is even eligible to appear in whatever shorter list a searcher, or an engine, ends up building. That filter is worth taking seriously well before any answer engine enters the picture.
  5. No one, including the researchers who study this closely, has yet measured what a synthesized answer actually names by category. That gap is real, and it means there is no proven playbook to copy yet, which is exactly why measuring your own category's attention terrain now is worth more than waiting for someone else to publish the definitive study.

Honest limits

What this does not yet settle

  • No published, retrievable study directly measures the size or composition of the consideration set a synthesized answer itself presents, meaning how many alternatives an engine typically names per query, category by category. Every figure in this study describes what happens around that moment, in clicks and sessions, not the moment itself.
  • The choice-overload literature, the contested jam study set against the much larger meta-analysis that found no reliable effect, has never been re-run in a generative-answer context. Whether a collapsed shortlist helps or hurts the buyer when the shelf is a single written answer instead of a longer list is genuinely unknown.
  • There is no independent, census-grade measurement of the AI-answer layer's reach or behavior across engines, comparable to what Pew and the Reuters Institute have done for one slice of one search engine. The layer is a measurement frontier, not a solved census.
  • No longitudinal data compares the same buyer population's consideration-set size or comparison behavior before and after adopting generative answers within a single category. The available evidence is cross-sectional and platform-specific.
  • Category-level variation, whether a low-stakes reorder purchase behaves differently from a high-stakes considered purchase like a home renovation or a piece of business software, is not addressed anywhere in the current evidence base.

This is a synthesis of dated, attributed evidence, not a census. The AI-answer layer in particular has no independent, Nielsen-grade measurement yet, so readings of it are directional and named as a frontier, never presented as settled.

Straight answers

Frequently asked questions

Does a synthesized search summary actually reduce clicks to other websites?

Yes, and it has been measured directly. Pew Research Center studied 68,879 real Google searches by 900 US adults in April 2025 and found that when a search produced an AI summary, users clicked through to a traditional organic result only 8% of the time, versus 15% when no summary appeared, roughly half the rate. Sessions also ended without any further browsing more often when a summary was present, 26% of the time versus 16% without one.

Were buyers really comparing many options before AI summaries existed?

No, the shortlist was already short. Herbert Simon's satisficing theory and George Miller's finding that people hold roughly seven items, plus or minus two, in working memory both describe a mind that stops comparing once one option feels good enough. Field research backs this up: Hauser and Wernerfelt's 1990 study found real buyers seriously compared an average of only 3.3 to 4.0 brands per category, even in categories with dozens of options on the shelf.

How many alternatives does an AI answer like ChatGPT or an AI Overview actually name for a given query?

That specific number has not been published anywhere. This is the single largest gap identified in the research: no retrievable study measures the size or composition of the consideration set a synthesized answer itself presents, category by category. Every available figure describes what happens around that moment, the clicks that don't happen and the sessions that end early, not the content of the answer itself.

Does giving a buyer fewer choices actually make them happier or more likely to buy?

The evidence does not support that assumption. The famous 2000 jam study, which found a 6-flavor display outsold a 24-flavor display, is the founding case for choice overload but has not reliably replicated since. A much larger 2010 meta-analysis of 50 experiments and roughly 5,000 participants found the average choice-overload effect across all of them was close to zero. No research has yet re-run this question specifically for a generative-answer setting, so it remains genuinely unknown whether being the one answer an engine gives a buyer helps or simply means less information reached them.

Besides AI summaries, what else is narrowing the list a buyer sees before they choose a business?

Two other layers matter and both predate or sit alongside AI answers. Ranked search results were already concentrating attention hard, with the top three organic positions capturing 68.7% of clicks on a page, per aggregated Backlinko and Sistrix data. Separately, review-star thresholds have tightened sharply: BrightLocal's 2026 survey found 31% of US consumers now require at least a 4.5-star rating before even considering a local business, up from 17% the year before, a filter that operates before any ranking or AI answer gets involved.

Provenance

References

  1. Wikipedia, "Satisficing" (Herbert A. Simon) https://en.wikipedia.org/wiki/Satisficing
  2. Miller, G. A., "The Magical Number Seven, Plus or Minus Two," Psychological Review, 1956 https://en.wikipedia.org/wiki/The_Magical_Number_Seven,_Plus_or_Minus_Two
  3. Hauser & Wernerfelt, "An Evaluation Cost Model of Consideration Sets," Journal of Consumer Research, 1990 https://discovery.ucl.ac.uk/id/eprint/10218717/1/akchen_mitrofanov_consideration_sets.pdf
  4. Iyengar & Lepper, "When Choice is Demotivating," 2000, and Scheibehenne, Greifeneder & Todd, "Can There Ever Be Too Many Options? A Meta-Analytic Review," 2010 (both discussed via secondary replication-literature source) https://atticusli.com/replication-crisis/choice-overload-jam-study/
  5. Pew Research Center, "Do people click on links in Google AI summaries?", July 2025 https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
  6. Similarweb, "Zero-Click Marketing: What the 2026 Data Means" https://www.similarweb.com/blog/marketing/geo/zero-click-marketing/
  7. Gartner, Inc., press release, February 2024 https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents
  8. First Page Sage, aggregated Google CTR-by-position data citing Backlinko and Sistrix https://firstpagesage.com/reports/google-click-through-rates-ctrs-by-ranking-position/
  9. Reuters Institute for the Study of Journalism, University of Oxford, Digital News Report 2025 https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2025/dnr-executive-summary
  10. BrightLocal, Local Consumer Review Survey 2026 https://www.brightlocal.com/research/local-consumer-review-survey/
  11. "Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music," arXiv:2607.23749, 2026 https://arxiv.org/abs/2607.23749

Every measured figure is dated to its capture and tagged with an evidence tier. Every cited work is real and locatable. Where an engine could not be captured this round, it is named as uncaptured, not estimated. Small-sample readings are labelled as directional.

Find out what your shortlist actually looks like

Every figure in this study describes the world in general. Your category, your metro, and your buyers are a specific and different question, and right now most businesses have no idea whether they clear the smaller list an engine or a searcher actually builds before deciding. That is what a Visibility Corpus reading does: it maps where your market's attention sits today across ranked results, review platforms, and the answers engines are giving, and turns that into a Machine-Readiness Score you can act on. It tells you where you stand and what is realistically movable from there.