MSME & Global Commerce · emerging evidence

Which Indian Language Does ChatGPT Actually Trust? A Reading of the Hallucination Data

Last reviewed 2026-08-09. Written by Chandranshu Kumar, Founder, Raveneye Global. · 8 min read

A peer-reviewed 2025 benchmark on AI hallucination in the Hindi language tested six large language models, including the GPT-4o model that has powered ChatGPT, on conversational replies in Hindi, Farsi, and Mandarin. Across nearly every model and dataset tested, Hindi and Farsi responses scored far lower on semantic alignment with the correct reply than Mandarin did, a pattern the researchers attribute to less Hindi and Farsi training data. On one dataset, GPT-4o's semantic-alignment score in Hindi was roughly a third of its score in Mandarin. The study did not test English directly, so "worse than English" is an inference, not a measured result, and a separate, larger benchmark across 30 languages found no overall link between a language's resource level and its hallucination rate. Read together, the conclusion is narrower than the headline suggests: in tested conversational settings, Hindi answers carry a measurably higher risk of semantic error, but that risk has not been shown to be a fixed law of every AI task in every setting.

What the benchmark actually tested

In 2025, a team of researchers led by Amit Das published "Investigating Hallucination in Conversations for Low Resource Languages," a study that ran six large language models, GPT-3.5, GPT-4o, Llama-3.1, Gemma-2, Qwen-3, and DeepSeek-R1, through two English-language conversational datasets, BlendedSkillTalk and DailyDialog, after translating each conversation into Hindi, Farsi, and Mandarin. Native speakers checked every translation before the models were asked to generate the next turn of the conversation.

The models' replies were then scored three ways. ROUGE-L measures raw word-overlap with the reference reply, a surface-level check. FactCC checks for gross factual inconsistency. The score that carries the real diagnostic weight is NLI, or natural language inference, which grades whether a model's reply is semantically aligned with what the conversation actually called for, catching the kind of hallucination where an answer is fluent, plausible, and unrelated to what was asked.

This is a narrower and more specific claim than "AI is bad at Hindi." It is a measurement of how often six named models, on two specific datasets, produced a reply that a semantic-alignment check flagged as disconnected from the conversation, in three specific languages. That precision is what makes the finding usable rather than a vibe.

The number: a wide semantic-alignment gap on the cleaner dataset

On the BlendedSkillTalk dataset, the gap between Hindi and Mandarin was consistent across every one of the six models tested. Hindi NLI scores ranged from 29.20 to 42.46; Mandarin NLI scores ranged from 80.53 to 95.41. For GPT-4o specifically, the OpenAI model that has powered ChatGPT, the Hindi NLI score was 29.20 against a Mandarin score of 88.54, roughly a third as high. FactCC, by contrast, stayed above 89 percent for every model in every language, meaning the more traditional factual-consistency check barely distinguished between them. The semantic-alignment metric is where the real signal appeared.

The paper's own reading of the pattern is direct: "the notably low hallucination rate observed in Mandarin can be attributed to the availability of large volumes of high-quality training data for this language. In contrast, the elevated hallucination rates in Hindi and Farsi highlight the challenges faced by LLMs when dealing with low-resource languages." Note that Farsi, not Hindi alone, shares this pattern in the study; the researchers group the two together throughout, not as a Hindi-specific finding.

The second dataset, DailyDialog, showed the same broad direction but with more model-to-model noise. GPT-4o again scored far higher on Mandarin (93.56) than Hindi (8.09). But GPT-3.5 broke the pattern on this dataset, scoring higher on Hindi (55.07) than on Mandarin (26.83). A careful reading of the paper does not smooth this away: the Hindi-versus-Mandarin gap is real and repeats across most model-dataset combinations, but it is not a clean, uniform law that holds without exception in every cell of the table.

Why less training data produces this specific failure mode

The mechanism the researchers propose is not exotic. A model that has seen comparatively little high-quality Hindi conversational text has to extrapolate more when generating a Hindi reply, and extrapolation is where hallucination lives. The paper's worked examples show what this looks like in practice: a Hindi speaker asks "I hope so, how old are your kids?" and GPT-3.5 replies "I would be happy to help answer your question," a fluent sentence that answers nothing. A Hindi speaker mentions two young children and Gemma-2, on a different prompt, launches into an unrelated safety disclaimer about controlled substances. These are not garbled or ungrammatical outputs. They are confident, well-formed sentences that simply are not connected to what was said, which is exactly the failure mode NLI is built to catch and ROUGE-L is not.

By contrast, the paper reports that Mandarin hallucinations, when they occurred, were typically partial: a model would preserve the general topic but invent a specific, unprompted detail, such as GPT-3.5 mentioning "this dish" in a Mandarin conversation that never referenced food. The researchers attribute this difference in severity, not just frequency, to Qwen and other models having been more heavily pretrained on Mandarin-language data specifically, a language-aware advantage rather than a property of Mandarin itself.

The complication: resource level is not a universal predictor

A separate 2025 study, "How Much Do LLMs Hallucinate across Languages?" by Saad Obaid ul Islam, Anne Lauscher, and Goran Glavaš, tested a different task, open-domain long-form question answering rather than conversational reply generation, across 30 languages and eleven model families. Its finding complicates any attempt to generalize the Hindi result into a universal rule: across all 30 languages, the researchers "find no correlation between the hallucination rates and measures of language resourceness." In their data, the lowest hallucination rate belonged to Sindhi, a lower-resource language, at 5.83 percent of tokens; the highest belonged to Hebrew, a comparatively higher-resource language, at 16.81 percent.

The same study found that models with broader declared language support hallucinated more on average, not less (a statistically significant correlation, r = 0.88), and that larger models hallucinated significantly less than smaller ones regardless of language. Neither of those findings is what a simple "low-resource languages hallucinate more" story would predict.

What this means for how much weight the Hindi finding can carry

The two studies are not testing the same thing, and that is the point. Conversational reply generation, tested by Das and colleagues, rewards a model for staying tightly grounded in the immediate context of what a specific person just said, a task where thinner training data in a specific language shows up as clear semantic drift. Open-domain factual question answering, tested by Islam, Lauscher, and Glavaš, draws on a different mix of pretraining, retrieval-style knowledge, and model architecture, where resource level turned out not to be the dominant factor. Both results can be true at once. The synthesis is that Hindi's higher hallucination risk is well evidenced in the specific setting that was tested, conversational, dialogue-style generation, and should not be quietly extended into a blanket claim about every kind of question an Indian buyer might ask an AI engine.

Why this is not an abstract problem for Indian MSMEs

The scale of vernacular usage in India is not in question, whatever the hallucination literature ultimately settles on. The Internet in India 2024 report, jointly produced by the Internet and Mobile Association of India and Kantar, put active internet users in the country at 886 million, with 98 percent of them consuming content in an Indic language and 57 percent of urban internet users actively preferring regional-language content over English. The same report found that nine in ten internet users had already interacted with an app carrying embedded AI capabilities. Vernacular is not an edge case in the Indian market; it is closer to the default.

Google has been building toward exactly this reality inside Search itself. AI Mode, its generative search layer built on a custom version of Gemini 2.5, rolled out in India in English in mid-2025 and then, according to Google's own product announcement, expanded into Hindi starting September 8, 2025, with additional Indian languages following. That is a direct, verifiable signal that a meaningfully large share of the AI-engine answers an Indian buyer sees when researching a local business are already being generated in Hindi rather than translated from English, and the share is set to grow.

Put the two facts together and the exposure becomes concrete. An MSME with an English-only web presence is already invisible to a Hindi-language AI Mode query, regardless of any hallucination question. But even an MSME with a well-built vernacular presence is not automatically safe from the second risk: the tested evidence above suggests that when a model does generate a Hindi-language answer, in the conversational settings studied so far, it carries a measurably higher chance of semantic drift than the same model would show in a higher-resource language.

What is already being built to close the gap

India's public and private sectors are treating this as a known, active infrastructure problem rather than a settled fact of nature. Bhashini, the National Language Translation Mission run by the Ministry of Electronics and Information Technology, exists specifically to build shared AI language infrastructure across Indian languages so that individual products do not each have to solve the low-resource problem alone. In April 2025, the Government of India selected Sarvam, working with the AI4Bharat research group at IIT Madras, to build the country's sovereign large language model under the IndiaAI Mission, with an explicit mandate to be "fluent in Indian languages" rather than fluent in English with Indian languages as an afterthought. Sarvam's existing SarvamM model, trained across ten Indian languages, is one concrete instance of the approach the original hallucination study itself recommends: "models specifically pretrained or fine-tuned on native corpora... show reduced hallucination, highlighting the importance of language-aware pretraining strategies."

None of this means the gap is closed. It means the gap is recognized, funded, and actively being worked by people with strong incentives to close it, which is a different, more optimistic claim than either "Hindi AI will always hallucinate more" or "the problem is already solved."

What this evidence does and does not show

No public dataset yet measures, directly and specifically, how often an AI engine states something false about a named small business when a buyer asks in Hindi versus in English. That is a real, open measurement gap, not a number this piece is going to fabricate or round up to fill. What can be said, and can be traced to a working URL, is narrower and still useful: a peer-reviewed 2025 benchmark found a consistent, sizable semantic-alignment gap between Hindi and Mandarin conversational replies across nearly every large language model it tested, attributed by the researchers to thinner Hindi training data; a separate, larger benchmark found that this pattern does not generalize cleanly across all 30 languages and all task types; and India's own AI infrastructure investment is being built around closing exactly this kind of gap.

For a business owner, the practical posture that follows is not panic and is not indifference. It is verification. The presence of an answer from an AI engine about a given business, in Hindi or in English, is not proof that the answer is accurate. Given what the published evidence actually shows about Hindi-language generation specifically, that gap between presence and accuracy is worth checking directly rather than assuming away in either direction.

The evidence

Key findings, with their sources

  • Across all six models tested on the BlendedSkillTalk dataset, Hindi NLI (semantic-alignment) scores ranged 29.20-42.46, versus 80.53-95.41 for Mandarin.

    established Das, A. et al., "Investigating Hallucination in Conversations for Low Resource Languages", arXiv:2507.22720, Table 2.

  • GPT-4o's Hindi NLI score on BlendedSkillTalk was 29.20 versus 88.54 in Mandarin, roughly a third as high.

    established Das, A. et al., arXiv:2507.22720, Table 2.

  • FactCC (factual-consistency) scores stayed above 89% for every model in every language tested, meaning the semantic-alignment (NLI) metric, not FactCC, is what actually distinguished Hindi from Mandarin.

    established Das, A. et al., arXiv:2507.22720, Tables 2-3.

  • On the DailyDialog dataset, the Hindi-versus-Mandarin gap held for most models but reversed for GPT-3.5 (Hindi NLI 55.07 vs. Mandarin NLI 26.83), showing the pattern is consistent but not universal within the same study.

    established Das, A. et al., arXiv:2507.22720, Table 3.

  • A separate 30-language benchmark using open-domain question answering found no correlation between a language's resource level and its hallucination rate; the lowest rate was in Sindhi (5.83% of tokens) and the highest in Hebrew (16.81%).

    established Islam, S.O. et al., "How Much Do LLMs Hallucinate across Languages?", arXiv:2502.12769.

  • In that same 30-language study, LLMs with broader declared language support showed significantly higher average hallucination rates (r = 0.88, p = 0.049), and larger models hallucinated significantly less than smaller ones.

    established Islam, S.O. et al., arXiv:2502.12769.

  • 886 million Indians were active internet users in 2024, with 98% consuming content in an Indic language and 57% of urban users preferring regional-language content.

    established Internet and Mobile Association of India (IAMAI) and Kantar, "Internet in India Report 2024".

  • Nine in ten Indian internet users had already interacted with an app carrying embedded AI capabilities, per the same 2024 report.

    established IAMAI and Kantar, "Internet in India Report 2024" (reported via Kantar director Biswapriya Bhattacharya).

  • Google's AI Mode, built on a custom version of Gemini 2.5, began rolling out in Hindi in India on September 8, 2025.

    established Google (Hema Budaraju, VP Product Management, Search), "Google Search: AI Mode now in Hindi", blog.google, September 2025.

  • In April 2025, the Government of India selected Sarvam, working with AI4Bharat at IIT Madras, to build India's sovereign LLM under the IndiaAI Mission, explicitly to be fluent in Indian languages.

    established Sarvam AI, "The government of India selects Sarvam to build India's sovereign large language model", sarvam.ai, April 26, 2025.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedThe specific, sourced findings: the Hindi-versus-Mandarin NLI gap on BlendedSkillTalk across six models; India's vernacular internet-usage scale; Google AI Mode's Hindi rollout; the Bhashini and Sarvam sovereign-AI programs.Das et al., arXiv:2507.22720; IAMAI-Kantar Internet in India 2024; Google blog.google announcement, September 2025; Sarvam AI, April 2025; Bhashini, MeitY.
emergingTreating "less Hindi training data causes more conversational hallucination" as a well-supported but not fully settled causal mechanism, and treating language-aware fine-tuning (the Sarvam/Bhashini approach) as the plausible fix.The causal claim is the study authors' own stated interpretation of their data, not an independently established mechanism tested against alternative explanations.
contestedGeneralizing "Hindi hallucinates more" into a universal claim that holds across every AI task, every model, and every language pair, including the specific claim that Hindi underperforms English (which was not directly tested in the primary study).A separate 30-language, different-task benchmark (Islam et al., arXiv:2502.12769) found no overall correlation between language resource level and hallucination rate, directly complicating a universal reading.

Reference

Glossary

Semantic hallucination
A model reply that is fluent and plausible but disconnected from what the conversation or question actually called for, as distinct from a reply that merely uses different words to say the same correct thing.
NLI (natural language inference) score
In this research, a metric that grades whether a model's generated reply is semantically aligned with the correct reference reply. The metric that actually distinguished Hindi from Mandarin in the benchmark discussed here, unlike the near-saturated FactCC score.
Low-resource language
A language for which comparatively little high-quality digital text exists to train a language model on, which can leave a model with less grounding and a greater tendency to extrapolate, and therefore hallucinate, when generating in that language.
Sovereign LLM
A large language model built, trained, and deployed within a country using that country's own infrastructure and priorities, in this case explicitly to be fluent in Indian languages rather than English-first with Indian-language support added afterward.
Machine legibility
The degree to which a business's identity, offering, and credibility are structured and corroborated in forms that AI answer engines can read, cite, and describe accurately, in whichever language a buyer actually asks in.

Straight answers

Frequently asked questions

Does ChatGPT hallucinate more in Hindi than in English?

A 2025 benchmark tested GPT-4o, the model that has powered ChatGPT, in Hindi, Farsi, and Mandarin, and found Hindi scored far lower on semantic alignment than Mandarin. The study did not include a directly comparable English test, so "worse than English" is a reasonable inference from the broader hallucination literature, not a number this specific study measured.

Why would an AI model hallucinate more in Hindi specifically?

The researchers behind the study attribute it to the comparatively smaller volume of high-quality Hindi conversational text available to train these models on, which forces more extrapolation, and extrapolation is where semantic hallucination tends to appear. Farsi showed a similar pattern in the same study.

Does this mean low-resource languages always hallucinate more?

No. A separate, larger 2025 benchmark covering 30 languages and a different task, open-domain question answering, found no overall correlation between a language's resource level and its hallucination rate. The Hindi finding is well evidenced in the specific conversational setting tested; it is not established as a universal rule.

Is this problem being fixed?

Yes, actively. The Indian government's Bhashini mission and the April 2025 selection of Sarvam, working with AI4Bharat at IIT Madras, to build India's sovereign LLM are both aimed at closing exactly this kind of gap through language-aware training, which is the same fix the original hallucination study recommends.

Why does this matter for a small business that only serves local customers?

Because a large and growing share of Indian buyers now research local businesses through AI-engine answers in Hindi rather than English; Google's AI Mode began rolling out in Hindi in India in September 2025. A business absent from that layer is not found; a business present in it is still subject to the accuracy risk this research describes.

What should a business actually do with this information?

Treat an answer from an AI engine, in any language, as something to verify rather than assume. A Machine-Readiness Score checks what search engines, the map pack, and AI answers currently say about a specific business, which is the only way to know whether the general research pattern is showing up in that business's own case.

Provenance

Sources

  1. Das, A., Hasan, M.N., Sarkar, S., Zhang, Z., Jamshidi, F., Bhattacharya, T., Raychawdhary, N., Feng, D., Jain, V., Chadha, A., "Investigating Hallucination in Conversations for Low Resource Languages", arXiv:2507.22720, 2025 (established)arxiv.org
  2. Islam, S.O., Lauscher, A., Glavaš, G., "How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination", arXiv:2502.12769, 2025 (established)arxiv.org
  3. Internet and Mobile Association of India (IAMAI) and Kantar, "Internet in India Report 2024" (established)iamai.in
  4. MXM India, "Internet users in India set to cross 900 million: IAMAI-Kantar", summary of the Internet in India 2024 report (established)mxmindia.com
  5. Google (Hema Budaraju), "Google Search: AI Mode now in Hindi", blog.google/intl/en-in, September 8, 2025 (established)blog.google
  6. Sarvam AI, "The government of India selects Sarvam to build India's sovereign large language model", April 26, 2025 (established)sarvam.ai
  7. Bhashini, National Language Translation Mission, Ministry of Electronics and Information Technology, Government of India (established)bhashini.gov.in

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

About this analysis

The findings above describe model behavior under tested conditions, not a verdict on any single business. They point to a narrower, practical question that Raveneye's machine-readiness research treats as measurable rather than assumed: what does an AI engine currently say about a given business, and in which language, when a buyer asks. Most owners have not seen that answer, because standard analytics were not built to surface it. A Machine-Readiness Score applies the same evidence-first method described here to classic search, the local map pack, AI answers, and reputation, producing a read grounded in what was actually found rather than in inference.

diagnostic Surface Intelligence Audit A measured read of where a business stands across the four surfaces buyers now use to find and evaluate it, set against the competitors appearing ahead of it, with a ranked list of the corrections most likely to matter first. See how it works

A Machine-Readiness Score is a specialist-reviewed read of where a business stands across search and AI answers, offered at no cost and with no obligation. It reports a measured position, not a guaranteed number.