MSME & Global Commerce · emerging evidence
The Vernacular Answer Gap: Why 900 Million Indians Search in Indic Languages While AI Engines Answer in English
India now has more than 900 million internet users, and by the IAMAI-Kantar count nearly all of them, about 98 percent, consume content in Indic languages rather than English. Yet the generative engines that increasingly answer their questions, ChatGPT, Google's Gemini and AI Mode, and Perplexity, are trained and evaluated most heavily in English, which makes vernacular AI search visibility in India an open and measurable question. Peer-reviewed benchmarks such as IndicGenBench and factual-accuracy studies across 19 Indic languages find a consistent gap: the same question answered in an Indic language is more likely to be thinner, mistranslated, or hallucinated than its English twin. This is the vernacular answer gap: Indic AI works, just less well than its English counterpart, and the difference lands hardest on the small businesses whose customers ask in their own language. This piece assembles the public evidence, separates what is measured from what is not, and explains why machine legibility in the right language is now a visibility question, not a translation footnote.
Two facts that do not sit comfortably together
The first fact is about demand. India's internet population has crossed 900 million, and it is overwhelmingly a vernacular population. The second fact is about supply. The systems that now read the web and hand back a single synthesized answer are, by their own builders' admission and by independent measurement, at their strongest in English. Put those two facts in the same room and a gap appears between the language a market searches in and the language a machine answers best in.
That gap matters more than it used to, because the interface changed. When search returned ten blue links, a weak Indic-language result still sat on a page a person could scan, and the person could switch to English, reword the query, or scroll. When an answer engine returns one paragraph and names one or two businesses inside it, the quality of that single Indic-language answer decides what the searcher sees. There is less room to recover from a thin or mistranslated response, because the response is the destination, not a list of options.
No public dataset yet measures the specific case at issue, how ChatGPT, Gemini, and Perplexity answer local-business and MSME-relevant queries across India's scheduled languages. That test does not exist in the open literature. What does exist is a body of peer-reviewed and government evidence about the two ends of the gap: how India searches, and how well these models perform in Indic languages generally. This article reads across that evidence and is explicit about where measurement stops and inference begins.
How India actually searches
The demand side is well documented. The IAMAI-Kantar Internet in India 2024 report counted 886 million active internet users in 2024, growing about 8 percent year on year, and projected the base past 900 million in 2025. Rural India, at 488 million users, is now the larger half at 55 percent of the total, and it is growing at roughly double the urban rate. This is not a metro, English-first internet. It is a national one.
Language use follows from that. By the same report, nearly all internet users, about 98 percent, access content in Indic languages, and 57 percent of even urban users prefer content in a regional language. The census numbers explain why. In the 2011 Census, roughly 128 million Indians, about 10 percent, reported any ability in English, and only a fraction of a percent spoke it as a first language, against 528 million native Hindi speakers and hundreds of millions more across Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam, Odia, Punjabi, and the rest of the 22 languages in the Constitution's Eighth Schedule.
The policy layer has moved to match this. Bhashini, the government's National Language Technology Mission under the Ministry of Electronics and Information Technology, was launched in 2022 to build translation and speech models as public infrastructure across all 22 scheduled languages, and has published more than a thousand models on its platform. The direction of travel is unambiguous: the Indian internet is becoming more vernacular, not less, and the state is investing to make it so.
What the models actually do in Indic languages
The supply side is where the gap is measured. In 2024, a Google Research team published IndicGenBench at the ACL conference, a human-curated benchmark spanning 29 Indic languages across 13 scripts and 4 language families, covering tasks like cross-lingual summarization, translation, and question answering. Its headline finding was blunt: even the strongest model tested showed a significant performance gap in every Indic language compared with English, and the authors concluded that further work is needed before multilingual models can be called inclusive.
A 2025 study from a group at the Pune Institute of Computer Technology and IIT Madras, pointedly titled Better To Ask in English?, put the question to newer models. Using the IndicQuest dataset of 4,000 fact-based question-answer pairs across English and 19 Indic languages, the authors evaluated GPT-4o, Gemma-2, and Llama-3.1 on region-specific questions about Indian history, geography, politics, literature, and economics. Their conclusion is the one that matters for this piece: English consistently outperformed the Indic languages across the models and metrics tested, even on questions rooted in Indian context, and the models hallucinated more often when answering in low-resource Indic languages.
The gap is not uniform, and that is the point
The same study is worth reading closely, because the degradation is uneven. The smaller open models held up reasonably in Hindi and Marathi, which are comparatively well resourced, but fell sharply in the lowest-resource languages, with Odia, Urdu, Sindhi, and Manipuri scoring lowest on factual accuracy. History and geography, the domains that lean most on region-specific local knowledge, were the hardest of all.
That unevenness is the practical shape of the problem. A business in a Hindi-speaking market sits closer to the models' strong zone than a business serving Odia, Assamese, or Manipuri customers, where the answer a machine returns is more likely to be wrong or fabricated. The vernacular answer gap is not one gap; it is a gradient, widest exactly where digital resources are already thinnest, which mirrors the older pattern where the businesses least represented online are the ones a machine reads least reliably.
The visible symptom: AI answers in Hindi
The benchmark findings show up in shipping products. When Google brought AI Overviews to Hindi in India in 2024, TechCrunch documented the quality problems directly. Asked in Hindi what foods to eat in summer, one overview returned chiknai wali cheezien, literally sticky or greasy things, where the English answer said oily. Asked about YouTube's leadership, the Hindi answer garbled the tense and stated that Neal Mohan was Google's CEO until a date, while the English version correctly said he is the CEO as of that date. Small rewordings of the same Hindi question produced an answer in one form and nothing in another.
The reporting also named the mechanism, and it is the same one the academic work points to. As TechCrunch put it, results for similar questions in English were much better than the Hindi ones, partly because there are more and better sources available in that language. The engine is only as good as the corroborated material it can read, and for most topics there is far more of that material in English than in any Indic language. The model does not fail because it dislikes Hindi; it fails because the web it reads is thin in Hindi on that topic.
The gap is being worked on, not closed
The trajectory is one of investment, not resolution. Google extended its conversational AI Mode to India in June 2025 and added Hindi to it in September 2025, alongside a handful of other languages, using a customized version of its Gemini model. Bhashini keeps publishing models, and Indian labs keep releasing Indic benchmarks and datasets. All of that is real progress. None of it has yet erased the measured gap, and a business cannot wait for a future in which it might. It has to be found and chosen in the language landscape that exists now, where the Indic-language answer is improving but still trails its English equivalent.
Why this lands hardest on MSMEs
A large enterprise can compensate for a weak Indic-language web. It can publish authoritative content in several languages, earn coverage that gets cited, and maintain the consistent, structured identity that answer engines reward. Most micro, small, and medium enterprises cannot. Their online presence is often a single listing, a social profile, and a handful of reviews, frequently in a mix of languages and rarely structured for a machine to parse. When a customer asks an engine in Marathi or Tamil for a recommendation, the business that gets named is the one the machine can read and corroborate in that language, and that is usually not the small local firm.
This is where the vernacular answer gap becomes a machine-readiness problem rather than a linguistics problem. The question is not whether an MSME's customers speak the language; it is whether the business's identity, offering, and credibility exist in a form that an engine can find and trust when it answers in that language. The corroboration an answer engine looks for, consistent details across sources, third-party mentions, structured data, reviews, is scarcer in Indic-language form for small firms, so even a firm whose customers are entirely vernacular can be invisible to the vernacular answer.
The stakes are rising because commerce itself is going vernacular. India's Open Network for Digital Commerce, ONDC, built on the open Beckn protocol to bring kirana stores and small sellers onto a shared network, had onboarded hundreds of thousands of sellers and was processing more than twelve million transactions a month by early 2025, having crossed 200 million cumulative transactions. As discovery for these sellers moves toward voice and vernacular interfaces, the ability to be surfaced correctly in an Indic-language answer stops being a marketing nicety and becomes a condition of being found at all.
There is a compounding effect worth naming. A small firm that is thinly represented in its customers' language generates little corroborated Indic-language material about itself, which gives the engines less to read, which makes them likelier to omit or misdescribe it, which in turn produces fewer vernacular mentions of the firm. The under-represented stay under-represented. Breaking that loop does not require a large budget so much as deliberate structure: a clear, consistent business identity, accurate details repeated across the places an engine checks, and content in the languages the firm's customers actually use rather than only in English.
What we can claim, and what we cannot
Three things here are well established. India searches overwhelmingly in Indic languages, by the IAMAI-Kantar and census records. Generative models perform measurably worse in Indic languages than in English, by IndicGenBench and the IndicQuest factual-accuracy study. And that weakness is visible in shipped products, by the AI Overviews reporting. None of those three is speculative.
One thing is not yet established. There is no public, audited dataset that measures how ChatGPT, Gemini, and Perplexity specifically answer local-business and MSME-relevant queries across India's scheduled languages, or how often a small firm that should be recommended is omitted or misnamed in a vernacular answer. That is an open question. It could be measured, by running a matched set of buyer-intent queries in English and in each target Indic language across the major engines and scoring the answers for accuracy, completeness, and correct business attribution, but until someone runs that study at scale, the per-engine MSME figure does not exist and no one should quote one.
What the surrounding evidence supports is a directional conclusion, not a precise number: the general Indic-language performance gap almost certainly extends into local and commercial queries, and the effect is worst in the least-resourced languages. That is enough to act on without inventing a statistic. The right response is to measure a specific business against the specific languages and surfaces its customers use, rather than to assume either that vernacular AI already works or that it is hopeless.
The evidence
Key findings, with their sources
-
India's internet base reached 886 million active users in 2024, growing about 8% year on year, and is projected past 900 million in 2025; nearly all users (about 98%) access content in Indic languages.
established IAMAI-Kantar, Internet in India 2024 report, summarized via IBEF, 2025, ibef.org.
-
Rural India, at 488 million users, is now 55% of the internet population and growing at roughly double the urban rate; 57% of even urban users prefer regional-language content.
established IAMAI-Kantar, Internet in India 2024 report, via IBEF, 2025, ibef.org.
-
In the 2011 Census, about 128 million Indians (roughly 10%) reported any ability in English and only a fraction of a percent as a first language, against 528 million native Hindi speakers.
established Census of India 2011, via "Multilingualism in India", Wikipedia (established, primary census data).
-
The Constitution's Eighth Schedule lists 22 scheduled languages, and the government's Bhashini mission (launched 2022) is building translation and speech models across all 22, with more than 1,000 models published.
established Eighth Schedule to the Constitution of India; Bhashini / National Language Technology Mission, MeitY, en.wikipedia.org and bhashini.gov.in.
-
IndicGenBench, a Google Research benchmark across 29 Indic languages, 13 scripts, and 4 language families, found a significant performance gap in every Indic language compared with English.
established Singh et al., "IndicGenBench", Proceedings of ACL 2024, aclanthology.org (peer-reviewed).
-
Evaluating GPT-4o, Gemma-2, and Llama-3.1 on 4,000 fact-based question-answer pairs across English and 19 Indic languages, English consistently outperformed the Indic languages, with more hallucination in low-resource languages.
established Rohera et al., "Better To Ask in English?", arXiv:2504.20022, 2025 (academic preprint).
-
In that study the degradation was uneven: comparatively well-resourced Hindi and Marathi held up better, while Odia, Urdu, Sindhi, and Manipuri scored lowest on factual accuracy, and history and geography were the hardest domains.
established Rohera et al., "Better To Ask in English?", arXiv:2504.20022, 2025.
-
When Google brought AI Overviews to Hindi, documented failures included rendering "oily" as "greasy or sticky things" and garbling a factual statement's tense; the reason cited was that more and better sources exist in English.
emerging Ivan Mehta, "Google's AI Overviews in Hindi need a quality upgrade", TechCrunch, 27 Aug 2024, techcrunch.com.
-
Google extended AI Mode to India in June 2025 and added Hindi to it in September 2025 using a customized Gemini model, evidence of active investment rather than a closed gap.
established "Google's AI Mode adds 5 new languages including Hindi", TechCrunch, 8 Sep 2025, techcrunch.com.
-
ONDC, India's open commerce network built on the Beckn protocol, had onboarded hundreds of thousands of sellers and was processing over 12 million transactions a month by early 2025, having crossed 200 million cumulative transactions.
emerging ONDC network updates, reported via YourStory and ondc.org, 2025.
-
No public, audited dataset yet measures how ChatGPT, Gemini, or Perplexity answer local-business and MSME-relevant queries across India's scheduled languages; the per-engine figure for that specific gap does not currently exist.
contested Author's assessment of the published literature as of August 2026 (open question, not a measured result).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | India's vernacular-first internet; the measured Indic-versus-English performance gap in generative models; the source-availability mechanism behind it. | IAMAI-Kantar / IBEF (2024-2025); Census of India 2011; IndicGenBench (ACL 2024); Rohera et al. (arXiv:2504.20022, 2025); TechCrunch (2024). |
| emerging | The framing of vernacular machine legibility as a distinct MSME visibility problem, and the shift of small-seller discovery toward voice and vernacular interfaces on networks like ONDC. | Extension of the machine-legibility thesis to Indic languages; ONDC scale reporting; Google AI Mode and Bhashini rollouts, which are progressing but not complete. |
| contested | Any precise per-engine claim about how often ChatGPT, Gemini, or Perplexity misnames or omits a specific MSME in a vernacular answer. | No audited public study measures this yet; the direction is inferable from the established evidence, but the magnitude is not, and no number should be quoted. |
Reference
Glossary
- Vernacular answer gap
- The gap between the Indic languages most Indians search in and the language, English, in which generative engines answer most reliably, so the same query yields a weaker answer in the vernacular than in English.
- Indic languages
- The languages of India, including the 22 in the Constitution's Eighth Schedule such as Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam, Odia, and Punjabi, in which the great majority of Indian internet users consume content.
- Low-resource language
- A language with comparatively little high-quality digitized text and training data. Models tend to be less accurate and more prone to hallucination in low-resource languages, a pattern documented for many Indic languages.
- Machine legibility
- The degree to which a business's identity, offering, and credibility are structured, consistent, and corroborated in forms an answer engine can parse and cite, here specifically in the language the customer uses.
- Answer engine
- A system such as ChatGPT, Google's Gemini and AI Mode, or Perplexity that reads the web and returns one synthesized answer, often naming a few businesses, rather than a ranked list of links.
- Bhashini
- India's government National Language Technology Mission, launched in 2022 under MeitY, building translation and speech models as public infrastructure across all 22 scheduled languages.
Straight answers
Frequently asked questions
Do AI search engines work worse in Indian languages than in English?
The published evidence says yes, on average. The IndicGenBench benchmark found a significant performance gap in every Indic language against English, and a 2025 study across 19 Indic languages found models answered the same factual questions more accurately in English and hallucinated more in low-resource Indic languages. The gap is real but uneven, worst in the least-resourced languages.
Why do generative engines default to English or mistranslate for Indic queries?
Mostly because the corroborated web they read is far larger in English than in any Indic language. When TechCrunch documented Google's Hindi AI Overviews failing, the cited reason was that more and better sources exist in English. The model is limited by the material available in each language, not by any dislike of the language itself.
Was an independent multi-engine test run across ChatGPT, Gemini, and Perplexity?
No. This article deliberately does not claim a proprietary multi-engine test, because no audited public dataset measures how those engines answer MSME-relevant queries across India's scheduled languages. It synthesizes the published research on how India searches and how models perform in Indic languages, and it describes how such a test could be run rather than inventing its results.
Why does this matter more for small businesses than large ones?
Large firms can publish authoritative multilingual content and earn the third-party corroboration answer engines reward. Most MSMEs have a thin, mixed-language presence that machines struggle to parse. So even when their customers are entirely vernacular, small firms can be omitted from the vernacular answer that decides who gets recommended.
Is the gap closing?
It is being worked on. Google added Hindi to its AI Mode in 2025, and the government's Bhashini mission keeps publishing Indic-language models. That is genuine progress, but the measured gap has not closed, so a business needs to be found in the language landscape that exists today rather than waiting for a future one.
What should an Indian MSME actually do about this?
Measure, do not assume. An MSME can check how it appears across classic search, the local map, and AI answers in the specific languages its customers use, then address the structure, consistency, and corroboration that let an engine read and recommend it. A Machine-Readiness Score is a starting read of exactly that.
Provenance
Sources
- IAMAI-Kantar, "Internet in India 2024" report, summarized: "India's internet users to exceed 900 million in 2025, driven by Indic languages", IBEF, 2025 (established)ibef.org
- IAMAI, "Internet in India 2024" (Kantar-IAMAI ICUBE), full report PDF (established)iamai.in
- "Multilingualism in India" (Census of India 2011 language data: English and Hindi speaker counts), Wikipedia (established, primary census)en.wikipedia.org
- "Eighth Schedule to the Constitution of India" (22 scheduled languages), Wikipedia (established)en.wikipedia.org
- "Bhashini" (National Language Technology Mission, MeitY), Wikipedia (established)en.wikipedia.org
- Singh et al., "IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages", ACL 2024 (peer-reviewed, established)aclanthology.org
- Rohera et al., "Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages", arXiv:2504.20022, 2025 (academic preprint, established)arxiv.org
- Ivan Mehta, "Google's AI Overviews in Hindi need a quality upgrade", TechCrunch, 27 Aug 2024 (emerging, press)techcrunch.com
- "Google's AI Mode adds 5 new languages including Hindi, Japanese, and Korean", TechCrunch, 8 Sep 2025 (established, product announcement)techcrunch.com
- ONDC (Open Network for Digital Commerce) network scale updates, 2025 (emerging, network reporting)ondc.org
- "ONDC crosses record 200M transactions", YourStory, 2025 (emerging, press)yourstory.com
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.