MSME & Global Commerce · emerging evidence

The Category League Table: Which MSME Verticals Actually Get Named by AI, and Which Don't

Last reviewed 2026-08-09. Written by Chandranshu Kumar, Founder, Raveneye Global. · 8 min read

No public, ongoing dataset yet functions as an MSME vertical AI answer audit for India, showing category by category how often Indian AI answer engines name a real MSME versus defaulting to generic advice or a marketplace link. What exists is aggregate global citation data from Ahrefs, the peer-reviewed Generative Engine Optimization (GEO) paper, and a handful of small, non-peer-reviewed industry tests run in the United States, the United Kingdom, and Germany. Those tests already show naming behavior varies sharply by category, driven mainly by whether a credentialed directory layer already exists for that vertical. India adds a variable none of those studies tested: a government-backed open commerce protocol, ONDC, alongside dominant vertical marketplaces such as IndiaMART, sitting on top of a Udyam-registered base of more than 92 million enterprises. This piece reads the published evidence, proposes the method a real Indian category league table would need, and states clearly which conclusions the evidence supports and which remain open questions.

The table that does not exist yet

No public, continuously updated dataset shows, category by category, how often India's AI answer engines name an actual micro, small, or medium enterprise rather than defaulting to generic buying advice or a link to a marketplace. What exists instead is aggregate citation data that is mostly global and domain-level, plus a small number of non-peer-reviewed industry studies testing a handful of categories in a handful of Western cities. Nothing purpose-built for India's MSME base, which the government's own real-time Udyam dashboard put at 9,24,39,224 registered enterprises as of 9 August 2026, exists in public form.

That gap matters because the answer to "should an MSME worry about AI answers" is not one number. A property lawyer and an accountant in the same city are not competing against the same aggregator gravity, and neither is competing on the same terms as a supplier already listed among IndiaMART's roughly 88 lakh registered suppliers. Treating AI visibility as one undifferentiated risk, or one undifferentiated opportunity, understates the variation the closest available evidence already shows.

This piece does two things a research-grade reading of an open question should do. It reads what has actually been published, most of it global rather than Indian, and states plainly where that evidence does and does not reach. And it proposes the method a genuine Indian category league table would need, rather than presenting an invented one.

What published citation data already shows: the aggregator head start

Ahrefs' Brand Radar analysis, built from roughly 76.7 million Google AI Overviews and close to a million prompts each in ChatGPT and Perplexity, found that Wikipedia is the single most cited domain across all three systems, drawing 16.3% of citations in ChatGPT, 12.5% in Perplexity, and 8.4% in AI Overviews. YouTube leads in Perplexity at 16.1% and AI Overviews at 9.5%, while Reddit and Quora rank high only in AI Overviews. The consistent finding across all three engines is that independent brand and small-business websites do not appear in the top ten cited domains for any of them; the citation layer is dominated by aggregators, user-generated platforms, and news outlets.

A separate Ahrefs study of 75,000 brands, filtered to domains with a Domain Rating above 40 and keywords with meaningful search volume, measured what actually correlates with being named. The strongest correlate was YouTube mentions, a Spearman coefficient of 0.737 with both ChatGPT and AI Overviews visibility. Broad brand mentions across the web correlated almost as strongly, 0.656 to 0.709 depending on the engine. Classic SEO metrics performed far worse: backlinks and URL Rating showed weak correlations, and Domain Rating only a moderate one. The practical reading is that being named now depends more on being mentioned across a network of independent, trusted sources than on ranking mechanics alone, and that is a structural head start for any aggregator or marketplace whose name already appears in thousands of reviews, directory listings, and social posts.

The peer-reviewed counterweight is the Generative Engine Optimization (GEO) paper published at KDD 2024, which tested specific content interventions rather than domain-level footprint. It found that adding cited statistics, direct quotations, and references to authoritative sources measurably raised a source's visibility inside generated answers across the systems it tested. That result matters for a small firm precisely because it does not depend on scale: a single well-structured, well-corroborated page can earn visibility the same lever an aggregator earns it with, even if the aggregator still starts with a wider footprint.

The category effect, tested outside India

The clearest evidence that naming behavior varies by category, not just by engine, comes from two independent, non-peer-reviewed industry studies run in 2025 and 2026. Neither covers India, and neither should be read as settled science, but both are transparent about method and both point the same direction: category structure predicts outcomes more than any single firm's quality does.

Property lawyers, accountants, and dentists diverge under the same test

An April 2026 study by independent researcher Simon Moser tested three professional-services categories, property lawyers, accountants, and dentists, across six cities in the United States, the United Kingdom, and Germany, running identical prompts through ChatGPT 5.3, Gemini 2.5 Pro, and Claude Sonnet 4.6. Across the full test, ChatGPT named an individual local business in about 1.2% of category instances, against roughly 35.9% for Google's local three-pack in the same categories, with Gemini reaching about 11% and Perplexity about 7.4%. The category breakdown is the more useful finding: property lawyers showed the strongest alignment between AI recommendations and traditional search results, because established legal directories already supply structured data both systems recognize. Accountants showed the opposite, zero confirmed overlap between AI recommendations and traditional search rankings across all six cities. Dentists fell in between, with neighborhood-level naming strategies proving effective in at least one market tested.

A 267,280-citation scan shows where niche directories, not broad platforms, do the gatekeeping

A separate case study by Local Dominator founder Eldar Cohen scanned 267,280 AI citation mentions across ChatGPT, Gemini, Perplexity, Claude, Grok, and Google AI Mode. Broad platforms dominated overall, Yelp at 71,512 citations and Google at 55,977, followed by Reddit and Facebook. But the category-specific layer told a different story: in home services, directories such as Angi (17,360 citations), HomeAdvisor (6,435), and Thumbtack (5,939) captured an outsized share, and in B2B categories, the Better Business Bureau (10,713) and Clutch (6,058) did the same. The implication is that two separate mechanisms are at work, a broad-platform effect that disadvantages every small business roughly equally, and a narrower, category-specific directory effect that only matters if that category happens to have built a dominant vertical directory.

Where India's aggregator layer differs from the studies above

None of the studies above tested India, and India's aggregator layer is not a simple import of the Western pattern. It combines dominant single-category marketplaces with a government-backed open protocol built explicitly to avoid a single platform owning the layer. IndiaMART describes itself as India's largest online B2B marketplace, reporting 23.4 crore buyers, 88 lakh suppliers, and 13.2 crore product and service listings on its own corporate site, a scale that makes it the default aggregator for a large share of India's industrial and wholesale MSME supply base.

Running alongside it is the Open Network for Digital Commerce (ONDC), a protocol rather than a platform, on which 306 network participants were live across 616-plus cities and 26 active domains as of its own published network figures, with 7.64 lakh-plus sellers and service providers and more than 16 million total orders recorded in a single month (May 2025). ONDC was built precisely to prevent any one buyer or seller app from becoming the sole gatekeeper the way a closed marketplace can. Whether that open-protocol design changes how AI answer engines treat an ONDC-listed seller, compared to a seller on a closed marketplace or a seller with no aggregator presence at all, is not something any published study has yet tested. It is a genuinely open question, and a Indian-specific variable the method below needs to account for rather than assume an answer to.

A proposed method for building the missing table

Building a real Indian category league table is not a technical problem so much as an undone one. A credible version would need a category taxonomy anchored to an existing standard, such as Udyam or NIC sector codes, rather than an invented list, so results are comparable to India's own MSME classification. It would need a consistent query set per category, mirroring the four-prompt structure the PolyGrowth study used (generic, qualified, neighborhood-level, and recommendation-framed), run across at minimum the three engines the closest global studies already track: ChatGPT, Gemini, and Perplexity.

Each response would need to be scored into one of a small number of clear categories, a named individual business, generic buying advice with no business named, a link or reference to a marketplace or directory rather than an individual seller, or no usable answer at all, and that scoring would need to repeat on a fixed cadence, because model outputs drift and a single snapshot cannot support a claim about a category's standing over time. The output of that method is a per-category naming rate relative to other categories, not a single blended national score, because the evidence above already shows a blended score would hide the exact variation that matters to an individual business owner deciding where to invest attention.

What the evidence already implies, and what it does not

Two things can be said with reasonable confidence from what is already published, without inventing an Indian-specific number. First, categories that already have a credentialed, structured directory layer, the way property lawyers do in the PolyGrowth data and home services and B2B categories do in the Local Dominator data, are likely to show measurably different naming behavior than categories without one, largely independent of any individual firm's quality. Second, the mechanics that correlate with being named, broad mentions across a network of trusted third parties and structured, cited content, structurally favor whichever entity already has the wider, more corroborated footprint, which in most categories today is the marketplace or directory rather than any single MSME.

What remains genuinely unknown is the exact naming rate for any specific Indian vertical, whether ONDC's open-protocol design shifts the pattern in a way a closed marketplace does not, and whether queries in Hindi or another Indian language change engine behavior relative to the English-language tests summarized here. Those are the questions the proposed method above is built to answer. Until someone runs it for India, the responsible claim is the one this article has made: category structure is a real and evidenced predictor of AI-naming behavior, and no one has yet measured where India's own MSME categories fall on that spectrum.

The evidence

Key findings, with their sources

  • Wikipedia was the single most cited domain across ChatGPT (16.3%), Perplexity (12.5%), and Google AI Overviews (8.4%); independent brand or small-business sites did not appear in the top-10 cited domains for any of the three engines.

    established Ahrefs, "The 10 Most Mentioned Domains for ChatGPT, Perplexity, and AI Overviews" (Brand Radar analysis of ~76.7M AI Overviews and ~957K/953.5K ChatGPT/Perplexity prompts), 2026, ahrefs.com.

  • Across 75,000 brands analyzed, YouTube mentions correlated most strongly with AI visibility (Spearman 0.737 with ChatGPT and AI Overviews), broad brand web mentions correlated 0.656-0.709, and classic SEO metrics like backlinks and URL Rating correlated weakly.

    established Ahrefs, "Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews", 2026, ahrefs.com.

  • Adding cited statistics, quotations, and authoritative sources measurably raised a source's visibility inside generated answers in tested engines.

    established Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed).

  • India's Udyam dashboard recorded 9,24,39,224 (over 92 million) registered MSMEs in real time as of 9 August 2026, split into micro, small, and medium tiers.

    established Ministry of MSME, Government of India, Udyam Registration dashboard, dashboard.msme.gov.in (live figure, accessed 2026-08-09).

  • IndiaMART reports 23.4 crore buyers, 88 lakh suppliers, and 13.2 crore product and service listings, describing itself as India's largest online B2B marketplace.

    established IndiaMART InterMESH, corporate website, corporate.indiamart.com.

  • ONDC reported 306 live network participants across 616-plus cities and 26 active domains, with 7.64 lakh-plus sellers/service providers and more than 16 million total orders in May 2025.

    established Open Network for Digital Commerce (ONDC), official network statistics, ondc.org.

  • In a small independent test across property lawyers, accountants, and dentists in six US/UK/German cities, ChatGPT named an individual local business in about 1.2% of category instances versus roughly 35.9% for Google's local three-pack, with results varying sharply by category (strongest alignment for lawyers, zero cross-validation for accountants).

    contested Simon Moser, PolyGrowth, "Local Business AI Recommendations Study", April 2026, polygrowth.io (independent, non-peer-reviewed).

  • A scan of 267,280 AI citation mentions across six engines found niche vertical directories (Angi, HomeAdvisor, Thumbtack for home services; BBB, Clutch for B2B) captured an outsized share of category-specific citations even though broad platforms (Yelp, Google, Reddit) dominated overall.

    contested Eldar Cohen, Local Dominator, "AI Local SEO Citations Report" case study, 2026, localdominator.co (independent, non-peer-reviewed).

  • Users clicked a traditional search result in about 8% of searches when an AI summary was present, versus 15% without one, and clicked a link inside the summary itself only about 1% of the time.

    established Pew Research Center, "Do people click on links in Google AI summaries?", July 2025, pewresearch.org (browsing panel).

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedAggregate global citation and correlation data (Ahrefs Brand Radar); the GEO paper's content-lever findings; India's primary-source scale figures for Udyam, IndiaMART, and ONDC; the AI-summary click-through effect.Ahrefs (2026); Aggarwal et al., KDD 2024; Ministry of MSME Udyam dashboard; IndiaMART corporate site; ONDC network statistics; Pew Research Center (2025).
emergingThe thesis that category structure, specifically whether a vertical already has a credentialed directory layer, predicts AI-naming behavior more than firm quality, and that this pattern should extend to Indian MSME verticals.Extension of the PolyGrowth and Local Dominator category-level findings to a market and language context neither study tested.
contestedThe specific naming-rate percentages (1.2% ChatGPT vs 35.9% Google three-pack) and the niche-directory citation counts.Small, independently run, non-peer-reviewed industry studies (six cities, three categories; 267,280 citations from unspecified campaign sources) confined to the United States, United Kingdom, and Germany.

Reference

Glossary

Category league table
A comparative ranking, by business vertical rather than by individual firm, of how often AI answer engines name a real business in that category versus defaulting to generic advice or an aggregator link. No public version exists for India yet.
Named response vs. generic response
The scoring distinction a category audit needs: whether an AI answer identifies a specific business by name, offers generic buying advice with no business named, or points to a marketplace or directory rather than a seller.
Aggregator dominance
The tendency of AI answer engines to cite platforms that are already mentioned broadly across the web, reviews, directories, and social posts, rather than an individual business with a narrower footprint. Evidenced in Ahrefs' domain-level and correlation data.
GEO (Generative Engine Optimization)
The academic term, introduced in a 2024 KDD paper, for structuring content, citations, statistics, and authoritative sourcing so that generative answer engines are more likely to cite it.
ONDC (Open Network for Digital Commerce)
A government-backed, protocol-based (not platform-based) Indian commerce network that lets many independent buyer and seller apps interoperate, distinguishing India's aggregator layer from a single dominant marketplace.

Straight answers

Frequently asked questions

Does a public ranking exist showing which Indian MSME categories AI names businesses in?

No. What is published is mostly global, domain-level citation data plus a small number of non-peer-reviewed studies testing a handful of categories in the United States, United Kingdom, and Germany. Nothing purpose-built for Indian MSME verticals has been published in a form that lets categories be compared against each other.

Which categories are more likely to get an individual business named by AI?

The closest available evidence suggests categories with an existing credentialed, structured directory layer behave differently than categories without one. In a small non-peer-reviewed test, property lawyers showed strong alignment between AI and traditional search because established legal directories already supply structured data; accountants showed no overlap at all. That pattern has not been tested for Indian verticals specifically.

Why do marketplaces and directories dominate AI answers in some categories more than others?

Two mechanisms appear to be at work in published research. A broad-platform effect favors sites like Wikipedia, YouTube, and Reddit across nearly every category. A separate, narrower effect favors category-specific directories, such as home-services or B2B directories, but only in the categories where that kind of directory already exists and is heavily cited.

Will ONDC change this pattern for India?

That is an open question, not a settled one. ONDC is a protocol rather than a single platform, which is structurally different from the closed marketplaces the published studies tested. No study has yet measured whether being listed through an ONDC-connected seller app changes an individual MSME's odds of being named in an AI answer.

How would you actually measure this for India?

Define categories against an existing standard such as Udyam or NIC codes, run a consistent set of prompts per category across ChatGPT, Gemini, and Perplexity at minimum, score each answer as naming a business, giving generic advice, or pointing to a marketplace, and repeat the test on a fixed cadence since model behavior drifts over time.

What can an individual MSME do given this uncertainty?

Start from where the business actually stands today rather than a generic industry claim. A Machine-Readiness Score reads a business against classic search, the local map pack, AI answers, and reputation, so the read is specific to that business and its category rather than borrowed from a study run somewhere else.

Provenance

Sources

  1. Ahrefs, "The 10 Most Mentioned Domains for ChatGPT, Perplexity, and AI Overviews" (Brand Radar analysis, established)ahrefs.com
  2. Ahrefs, "Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews" (75,000-brand correlation study, established)ahrefs.com
  3. Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed, established)arxiv.org
  4. Ministry of MSME, Government of India, Udyam Registration dashboard (live official figures, established)dashboard.msme.gov.in
  5. IndiaMART InterMESH, corporate statistics page (established, primary source)corporate.indiamart.com
  6. Open Network for Digital Commerce (ONDC), official network statistics (established, primary source)ondc.org
  7. Simon Moser, PolyGrowth, "Local Business AI Recommendations Study", April 2026 (contested, independent industry study)polygrowth.io
  8. Eldar Cohen, Local Dominator, "AI Local SEO Citations Report" case study, 2026 (contested, independent industry study)localdominator.co
  9. Pew Research Center, "Do people click on links in Google AI summaries?", July 2025 (established)pewresearch.org

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

The measurement behind this

Where a category sits on this spectrum, close to a credentialed directory layer or far from one, appears to matter more to a business's AI visibility than most generic advice accounts for, and no published study has measured that reading for Indian MSME verticals yet. What can be measured today is where a specific business stands, not a borrowed category average. Raveneye's Machine-Readiness Score reads a business across classic search, the local map pack, AI answers, and reputation, applying the same evidence-first method used in this analysis to one business at a time.

diagnostic Surface Intelligence Audit A measured read of where a business stands across the surfaces buyers now use to find and choose a supplier, set against the competitors and aggregators already appearing ahead of it in that category. See how it works

A free, specialist-reviewed reading of where a business stands across search and AI answers today. No guaranteed number, and no obligation.