MSME & Global Commerce · emerging evidence
GEO Isn't SEO for Bharat: Do Aggregators Eat the Answer Before the MSME Gets a Turn?
No public study has yet measured whether ChatGPT, Perplexity, or Google's AI Overviews cite aggregators like IndiaMART, JustDial, and Sulekha more often than the individual MSMEs listed inside them, so any specific percentage claiming to answer that would be invented. What does exist is adjacent, real evidence pointing the same direction: the 2024 peer-reviewed paper that introduced Generative Engine Optimization (GEO) found that citation inside a generated answer, not rank in a list, is now the object being competed for, and that citable authority and corroboration are what move it. Separately, published citation-concentration data shows a small number of very large, heavily linked domains capturing outsized citation share in general AI answers. IndiaMART and JustDial sit on exactly that kind of scale in India's B2B and local-search categories. Whether that pattern repeats for MSME-specific queries is a hypothesis consistent with the evidence, not a demonstrated result. This piece lays out the evidence, why aggregators would structurally have the advantage if the pattern holds, why ONDC's protocol design does not settle the question, and the method that would actually need to run to find out.
The question Bharat's MSMEs cannot yet see the answer to
Ask ChatGPT or Perplexity for a steel pipe supplier in Ludhiana or a reliable AC repair service in Chennai, and the assistant has to decide, in a fraction of a second, which sources to trust and which few names to put in the answer. The practical question behind this article is whether that decision structurally favors a small number of large aggregators, IndiaMART for B2B manufacturing and trade, JustDial and Sulekha for local services, over the individual business a buyer actually needs to reach. It is the Bharat-specific version of a debate playing out globally as generative answer engines replace the ranked list with a synthesized recommendation.
The stakes sit on top of a population that has the least capacity to respond to a citation disadvantage it cannot see. Over 7.83 crore enterprises, more than 78 million, had registered on India's Udyam Registration Portal and Udyam Assist Platform as of 28 February 2026, according to the Ministry of MSME's reply in the Rajya Sabha. Yet the same government's own survey partner found that internet usage for actual entrepreneurial activity, placing orders, taking UPI payments, transacting online, ran at only about 26.7 percent among unincorporated enterprises in 2023-24, and that a mere 81 establishments per 1,000 used the internet to make sales at all. That is the base population any citation-concentration effect would land on, whether or not anyone has yet measured the effect directly.
What Generative Engine Optimization actually measures
The relevant research anchor is the 2024 paper that coined the term: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande's "GEO: Generative Engine Optimization," accepted to KDD 2024. The paper's central move is a reframing. Classic SEO optimizes for rank in a list a human scans. GEO treats the unit of competition as a citation inside a synthesized answer, a fundamentally different target, since an engine can answer a query while naming only a handful of sources out of the thousands that rank for it.
Using a benchmark of diverse queries paired with real web sources (GEO-bench), the authors tested which content-level changes moved a source's visibility inside generated answers, and found that adding citations, statistics, and quotations from authoritative sources could lift visibility by up to 40 percent in the engines they tested, though the effect varied meaningfully by domain, the paper is explicit that no single tactic works uniformly everywhere. The paper was built and tested on general web queries, not India-specific B2B or local-service search, which matters: nothing in it directly measures Indian MSME categories, and applying it here is an extrapolation, not a replication.
What the paper does establish, and what carries over regardless of geography, is the shape of the game: a source wins a place in the answer by being citable, meaning structured, corroborated, and backed by the kind of authority signals an engine can verify quickly. That is a different, and in some ways higher, bar than simply existing online.
The structural case that aggregators start ahead
If citability is the currency, then scale, content velocity, and inbound-link density become direct advantages, and those are exactly the assets a large marketplace accumulates by default and an individual MSME's often-thin website usually lacks. Published data on which domains actually get cited in AI answers, while general-web rather than category-specific, shows how concentrated that advantage already is elsewhere. Ahrefs' Brand Radar analysis of roughly 76.7 million Google AI Overviews and about 950,000 prompts each on ChatGPT and Perplexity for June 2025 found Wikipedia cited in 16.3 percent of the ChatGPT answers sampled and 12.5 percent of Perplexity answers, with YouTube cited in 16.1 percent of Perplexity answers and 9.5 percent of AI Overviews, a handful of very large domains capturing a disproportionate share of citations across unrelated queries.
That pattern is not India-specific and it is not about B2B or local-service commerce, so it cannot be read as proof about IndiaMART, JustDial, or Sulekha. But it is the same structural signature, constant fresh content, heavy internal and external linking, high domain trust, that a national marketplace or local-search aggregator would be expected to carry into its own category if the pattern repeats there. The open question is whether it does.
IndiaMART, JustDial, and Sulekha, by the numbers
The scale differential, at least, is not in question. IndiaMART's FY2024-25 annual report states 211 million registered buyers and 8.4 million registered suppliers on its B2B marketplace. Justdial reports a database of approximately 56.1 million business listings and 192.9 million quarterly unique users across web, mobile, app, and voice as of 30 June 2026, up from 48.8 million listings and 191.3 million quarterly unique visitors disclosed in its April 2025 corporate presentation, a visible growth trend in both listings and reach. Sulekha describes itself as a leading digital platform for local service businesses across categories including home services, education, property, and events, active in major Indian cities. Unlike IndiaMART and JustDial, both publicly listed companies that disclose audited scale metrics in regulatory filings, Sulekha is privately held and does not publish comparable numbers, which means its citation exposure is harder for an outside researcher to size than the other two, an asymmetry that is itself worth noting rather than papering over.
ONDC was built to counter this, one layer down
India's own policy answer to a related concentration problem already exists, and it is worth reading precisely because it does not settle the question this article is asking. The Open Network for Digital Commerce (ONDC) is explicitly designed, per a Ministry of Commerce & Industry reply in the Lok Sabha, so that "seller-side apps make their full catalogues discoverable to all buyer-side apps, while buyer-side apps disclose key parameters used for sorting or listing search results, enabling sellers to understand and improve their ranking," with the stated goal that sellers are visible "regardless of size, scale or digital sophistication." As of 9 December 2025, more than 1.16 lakh retail sellers were live on the network across over 630 cities and towns, per the same government reply.
The network has since scaled further. Per a Department for Promotion of Industry and Internal Trade written reply in the Rajya Sabha on 7 August 2026, ONDC grew from 0.2 million transactions in FY23 to 218 million in FY26, crossing 500 million cumulative transactions in July 2026, with more than 2 lakh retail merchants and over 10 lakh mobility and logistics service providers active on the network by that point.
That is a genuine, code-level commitment to non-discriminatory discovery, and it directly addresses concentration risk inside apps built on the ONDC protocol. But it is a different layer from the one this article is about. ONDC's equal-visibility rules bind buyer-side and seller-side apps that speak the ONDC protocol. ChatGPT, Perplexity, and Google's AI Overviews are not ONDC buyer apps. Nothing in the public ONDC specification requires or even addresses how a general-purpose AI assistant should weigh an ONDC-registered seller against an IndiaMART listing or a JustDial profile when it synthesizes a natural-language answer to an open question. ONDC solving discovery fairness inside its own network, if it does, is not evidence about what happens in the separate, ungoverned layer where a generative engine decides whom to cite.
What a real test would look like
Because no public dataset currently applies a citation-share method to Indian MSME discovery queries, the next step is to describe the test rather than assert a result for it. A defensible version would follow the same logic as GEO-bench: assemble a representative sample of buyer-style queries across MSME-relevant verticals and cities, the kind an actual buyer or consumer would type, such as a category plus a location, rather than a brand name. Run each query across multiple engines, ChatGPT, Perplexity, Google's AI Overviews, and Gemini, since the Ahrefs data above shows citation behavior already differs meaningfully between them. Code every citation that appears by domain type: national aggregator (IndiaMART, JustDial, Sulekha, and comparable platforms), an ONDC-linked buyer-app storefront, an individual business's own website or Google Business Profile, or another category entirely. Compute the share of citations landing on each type, and repeat the run over time, since answer-engine behavior changes as the underlying models and their retrieval layers are updated.
That is a method, not a finding. Until it, or something like it, is actually run against Indian queries, any specific percentage claiming to show how often aggregators out-cite individual MSMEs in AI answers should be treated as unverified, regardless of how confidently it is stated.
What the evidence supports
Two things can be true at once. The general pattern in published citation data, a small number of large, heavily linked domains capturing outsized citation share, is real and independently documented, and GEO's own findings show that citable, corroborated authority is what moves visibility inside a generated answer, not simply being online. IndiaMART and JustDial, and by its own description Sulekha, carry exactly the kind of scale and content velocity that pattern rewards elsewhere, and the population they aggregate, more than 78 million registered MSMEs, most still building basic transactional internet habits per ICRIER's own 2025 survey, is the population least equipped to counter a citation disadvantage it cannot measure.
What is not established is the number itself. ONDC shows that India's policy apparatus has already built a serious, code-enforced answer to a related fairness problem, but at the marketplace-protocol layer, not the generative-answer layer this article is about, and the two should not be conflated. Until the query-sampling method above is actually run against Indian MSME categories, aggregators structurally out-citing individual businesses in AI answers remains a hypothesis consistent with the surrounding evidence, not a measured fact, and the only way to know where a specific business stands today is to check it directly rather than assume the pattern from adjacent data.
The evidence
Key findings, with their sources
-
GEO reframes the unit of optimization from search rank to citation inside a generated answer; tested content strategies (adding citations, statistics, and quotations) lifted visibility by up to 40% in the engines tested, with effectiveness varying by domain.
established Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed).
-
Across roughly 76.7 million Google AI Overviews and about 950,000 prompts each on ChatGPT and Perplexity (June 2025), Wikipedia was cited in 16.3% of sampled ChatGPT answers and 12.5% of Perplexity answers; YouTube was cited in 16.1% of Perplexity answers and 9.5% of AI Overviews.
contested Ahrefs, "The 10 Most Mentioned Domains for ChatGPT, Perplexity, and AI Overviews" (Brand Radar data, June 2025).
-
IndiaMART reported 211 million registered buyers and 8.4 million registered suppliers on its B2B marketplace as of FY2024-25.
established IndiaMART InterMesh, Annual Report FY2024-25, investor.indiamart.com.
-
Justdial reported a database of approximately 56.1 million business listings and 192.9 million quarterly unique users across web, mobile, app, and voice as of 30 June 2026, up from 48.8 million listings and 191.3 million quarterly unique visitors in its April 2025 corporate presentation.
established Justdial Limited, Company Overview (investor relations, justdial.com) and Corporate Presentation, April 2025.
-
Over 7.83 crore (78.3 million) enterprises had registered on India's Udyam Registration Portal and Udyam Assist Platform as of 28 February 2026.
established Ministry of Micro, Small & Medium Enterprises, Government of India, Rajya Sabha written reply, PIB Delhi, 30 March 2026.
-
Internet usage for entrepreneurial activity (online transactions, order placement, UPI payments) was about 26.7% among unincorporated enterprises in 2023-24; only 81 establishments per 1,000 used the internet to make sales.
established ICRIER, "Annual Survey of Micro, Small and Medium Enterprises (MSMEs) in India: The Role of Digitalisation in Enterprise Development", March 2025.
-
ONDC's protocol requires seller-side apps to make full catalogues discoverable to all buyer-side apps and buyer-side apps to disclose ranking parameters, with more than 1.16 lakh retail sellers live across over 630 cities and towns as of 9 December 2025.
established Ministry of Commerce & Industry, Government of India, Lok Sabha written reply, PIB Delhi, 16 December 2025.
-
ONDC transaction volume grew from 0.2 million in FY23 to 218 million in FY26, crossing 500 million cumulative transactions in July 2026, with more than 2 lakh retail merchants and over 10 lakh mobility/logistics service providers active on the network.
emerging Department for Promotion of Industry and Internal Trade, Rajya Sabha written reply, 7 August 2026, reported via bricscompetition.org.
-
No public dataset currently measures the share of AI-answer citations that go to Indian aggregators (IndiaMART, JustDial, Sulekha) versus individual MSME websites; any specific percentage for this comparison would be unverified.
contested Absence of a published Indian citation-share study; method proposed in this article.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | GEO's citation-not-rank reframing; India's Udyam registration base and low e-commerce/internet-for-sales usage rates; IndiaMART, Justdial, and ONDC scale metrics from official filings and government replies. | Aggarwal et al., KDD 2024 (arXiv:2311.09735); PIB releases (MSME, Commerce & Industry ministries); ICRIER MSME Survey 2025; IndiaMART Annual Report FY2024-25; Justdial investor disclosures. |
| emerging | Extending GEO's general-web citation-concentration pattern to India-specific B2B and local-service queries; treating ONDC's protocol-level fairness design as a distinct governance layer from generative-engine citation. | Reasoned extrapolation from Ahrefs Brand Radar data and the GEO paper; no India-specific or MSME-category study yet published; DPIIT transaction figures reported via secondary news coverage rather than a primary PIB release. |
| contested | Any specific claim about how often aggregators out-cite individual MSMEs in Indian AI answers. | No public dataset exists; this article proposes a method rather than reporting a result, and flags this explicitly as an open question. |
Reference
Glossary
- Generative Engine Optimization (GEO)
- The practice, named in a 2024 peer-reviewed paper, of optimizing content so it gets cited inside a generative engine's synthesized answer, rather than optimizing for rank in a traditional results list.
- The proportion of citations inside a generative engine's synthesized answers that go to a given domain or type of domain (for example, aggregator versus individual business site), for a defined set of queries.
- Aggregator
- A platform, such as IndiaMART, JustDial, or Sulekha, that lists many individual businesses under one domain, accumulating scale and content volume no single listed business controls on its own.
- ONDC (Open Network for Digital Commerce)
- A government-backed, protocol-based digital commerce network designed so seller catalogues are discoverable across participating buyer apps under disclosed, non-discriminatory ranking rules, distinct from how general-purpose AI assistants decide what to cite.
- GEO-bench
- The benchmark of diverse queries paired with real web sources built by the GEO paper's authors to measure which content changes move visibility inside generated answers; built for general web queries, not Indian MSME categories specifically.
Straight answers
Frequently asked questions
Do IndiaMART, JustDial, or Sulekha actually get cited more often than individual MSMEs in ChatGPT or Perplexity answers?
No public study has measured this specifically for Indian MSME queries, so there is no verified percentage to report. Published evidence from general AI-answer citation data shows large, heavily linked domains capturing outsized citation share elsewhere, and GEO research shows citable authority is what earns a place in an answer, both consistent with aggregators having an advantage, but neither one proves it for this specific comparison.
What is Generative Engine Optimization (GEO) and how is it different from SEO?
GEO, introduced in a 2024 peer-reviewed paper (arXiv:2311.09735), treats the goal as being cited inside a generative engine's synthesized answer rather than ranking in a traditional list of links. The paper found that adding citations, statistics, and quotations from authoritative sources could raise a source's visibility in generated answers by up to 40% in tested engines, with effectiveness varying by domain.
Does ONDC solve the problem of AI engines favoring aggregators?
Not directly. ONDC's protocol requires seller catalogues to be discoverable across participating buyer apps under disclosed ranking rules, a real fairness commitment inside its own network of more than 1.16 lakh sellers. But ChatGPT, Perplexity, and similar assistants are not ONDC buyer apps, and nothing in the ONDC specification governs how they decide whom to cite. ONDC addresses a different, marketplace-protocol layer of the discovery problem.
How large are IndiaMART and JustDial compared to an individual MSME website?
IndiaMART reported 211 million registered buyers and 8.4 million registered suppliers in its FY2024-25 annual report. Justdial reported roughly 56.1 million listings and 192.9 million quarterly unique users as of mid-2026. That scale gap is well documented; whether it translates into a proportional citation advantage inside AI answers is the open question this article describes rather than answers.
How could someone actually test whether aggregators out-cite individual businesses in AI answers?
By sampling representative buyer-style queries across MSME verticals and cities, running them across multiple engines, coding each citation by domain type (aggregator, ONDC-linked storefront, individual business site, other), and computing the citation share for each type over time. This mirrors the method the GEO paper used to build its own benchmark, applied to Indian queries it was not originally built to cover. No public dataset has run this yet.
What can an individual MSME do without waiting for that study to exist?
Check its own position directly rather than assume the pattern from adjacent data. A Machine-Readiness Score gives a specialist-reviewed read of where a specific business currently stands across search and AI answers, which is a narrower, more useful question than the unmeasured category-wide one.
Provenance
Sources
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed, established)arxiv.org
- Ahrefs, "The 10 Most Mentioned Domains for ChatGPT, Perplexity, and AI Overviews" (Brand Radar data, June 2025, contested/industry tier)ahrefs.com
- IndiaMART InterMesh, Annual Report FY2024-25 (established)investor.indiamart.com
- Justdial Limited, Company Overview, investor relations (established)justdial.com
- Justdial Limited, Corporate Presentation, April 2025 (established)justdial.com
- Sulekha.com, About Us (self-description, contested tier)sulekha.com
- Ministry of Micro, Small & Medium Enterprises, Government of India, "Over 7.83 crore enterprises registered on Udyam Registration Portal (URP)", PIB Delhi, 30 March 2026 (established)pib.gov.in
- Ministry of Micro, Small & Medium Enterprises, Government of India, "Registrations of Informal Micro Enterprises on Udyam Assist Platform cross 1.50 crore", PIB Delhi, 4 March 2024 (established)pib.gov.in
- ICRIER, "Annual Survey of Micro, Small and Medium Enterprises (MSMEs) in India: The Role of Digitalisation in Enterprise Development", Goyal, Puri & Khanna, March 2025 (established)icrier.org
- Ministry of Commerce & Industry, Government of India, "ONDC Enables Fair, Transparent and Inclusive E-Commerce...", PIB Delhi, 16 December 2025 (established)pib.gov.in
- BRICS Competition Law and Policy Centre, "India's Government-Backed ONDC Surpasses 500 Million Transactions", reporting a DPIIT Rajya Sabha written reply of 7 August 2026 (emerging, secondary-source tier)bricscompetition.org
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.