MSME & Global Commerce · emerging evidence

Bhashini's Billions of Calls, Zero MSME Directory: What India's Language Stack Wasn't Built For

Last reviewed 2026-08-09. Written by Chandranshu Kumar, Founder, Raveneye Global. · 8 min read

Bhashini, India's government-built language-AI stack, now moves billions of translation calls a year across all 22 constitutionally scheduled languages and sits inside DigiLocker, UMANG, and other citizen-facing apps. It has also reached into commerce: ONDC's Saarthi app runs on Bhashini, and the Ministry of MSME's TEAM initiative is walking five lakh small enterprises onto the ONDC network with multilingual support layered on top. None of that produces what an AI answer engine actually needs to recommend a business by name, the Bhashini MSME visibility gap this analysis examines: a structured, corroborated, machine-readable record of who that business is, what it does, and where it operates. Bhashini translates sentences. It was never asked to hold identity. This piece traces what Bhashini and its adjacent schemes actually build, where the MSME layer stops short of a directory, and why closing India's vernacular translation gap and closing its MSME machine-legibility gap are two different projects that happen to share an acronym-heavy neighborhood.

What Bhashini actually is, and how big it has gotten

Bhashini launched at Gandhinagar in July 2022 as the flagship product of the Ministry of Electronics and Information Technology's National Language Translation Mission. Its job, stated plainly, is to help Indian citizens communicate and transact across languages: machine translation, automatic speech recognition, and text-to-speech, offered as free, open APIs across all 22 constitutionally scheduled Indian languages. More than 300 pre-trained AI models are available to partners under that mandate, at zero marginal cost to integrate.

By 2026 the scale is genuinely large. Trade press coverage of India's AI policy landscape reports that Bhashini processed roughly 2.5 billion API calls over the course of 2026, on pace to exceed 5 billion by year-end, and that the platform is woven into DigiLocker, the successors to Aarogya Setu, and the UMANG super-app, the everyday citizen-services layer of the Indian state. Every one of those calls is, in effect, a sentence moving between an Indian language and a system that could not otherwise read it.

That is a real achievement, and it is also a precise description of the job Bhashini was built to do: move meaning between languages inside services that already know who the citizen is. A DigiLocker session already has the citizen's identity on file. A grievance portal already has the case number on record. Bhashini's role in each is translation, not identification. Nothing in that design brief asks the platform to describe a business to a machine that has never encountered it before.

The MSME layer Bhashini touches

Bhashini's reach into commerce runs mainly through the Open Network for Digital Commerce. In September 2024, ONDC launched Saarthi, a Bhashini-powered reference application that lets businesses build their own buyer-facing apps with voice and text support in Indian languages, starting with five, Hindi, English, Marathi, Bangla, and Tamil, with a stated plan to extend to all 22. On the government-scheme side, the Ministry of MSME's Trade Enablement and Marketing initiative, a sub-scheme of the World Bank-supported RAMP program, set out in June 2024 to onboard five lakh micro and small enterprises onto ONDC, half of them women-owned, backed by an outlay of Rs 277.35 crore over three years and implemented by the National Small Industries Corporation. The support on offer is concrete: help with digital cataloging, account management, logistics, and roughly 150 e-commerce workshops planned for tier-2 and tier-3 cities.

That scheme sits inside a larger 2026 push. At the MSME Day 2026 event on 27 June, the government reported that Udyam registrations, the formal MSME identity number, had crossed 8.7 crore, an enterprise base it says anchors the livelihoods of 38 crore people and contributes 31 percent of GDP, 35 percent of manufacturing output, and close to 45 percent of exports. Alongside that figure it launched or upgraded a cluster of MSME portals, PMEGP 2.0, SAMADHAAN 2.0, MSME Global Mart 2.0 (built to integrate with the ONDC ecosystem), and a new Testing Portal, all now offering services in all 22 scheduled languages through Bhashini and the National Informatics Centre.

Set against that infrastructure, the actual behavior of MSMEs online looks thin. A 2025 SIDBI-linked survey, reported by YourStory, found that while more than 90 percent of MSMEs now accept digital payments, only 13 percent actively use digital marketing or e-commerce to reach customers, with roughly seven in ten still relying on in-person referrals, print advertising, or trade events. Payment rails arrived. Discoverability did not follow at anything like the same pace.

Citizen infrastructure and commerce-identity infrastructure are not the same build

It is worth being precise about what each of these systems actually stores. Bhashini stores language models, not business records; a call to its API carries a sentence in and a sentence out, with no requirement that the sentence describe a company, its address, its hours, or its credibility. Udyam stores a registration number and classification data, built for eligibility and scheme access, not for public discovery. ONDC's registries function, by its own technical documentation, like a DNS for commerce, a directory of network participants for the purpose of routing transactions between buyer and seller apps, not a public, crawlable business directory addressed to the open web or to the systems that read it.

ONDC's own developer documentation is candid about the difficulty this creates even inside the network: existing platforms, brands, and merchants do not share a common cataloguing schema, so the protocol has had to build standardization work just to make listings comparable to each other across sellers. That is a live, acknowledged engineering problem for commerce inside the network. It says nothing about whether a listing, once standardized for ONDC's own buyer apps, is also structured in the schema.org, entity-graph terms that an external answer engine can parse, corroborate, and cite. Those are different specifications solving different problems, and nothing in the public record indicates ONDC's catalog schema was designed with the second one in mind.

None of this is a criticism of any single system doing the job it was funded to do. It is the observation that a citizen can now file a grievance, receive a scheme benefit, or place an ONDC order in her own language, and none of those interactions produces the artifact an AI answer engine actually consumes: a structured, corroborated, standing description of a specific business that exists independent of any single transaction.

What an answer engine needs, that none of this supplies

The 2024 peer-reviewed paper that introduced "Generative Engine Optimization" tested what actually moves a source into a generated answer, and found that citations, quotations, and authoritative, corroborated content raised a source's visibility inside the engines it tested. The unit being optimized is not a page rank; it is inclusion in a synthesized answer, and the inputs that earn it are structural and evidentiary properties of a business's presence on the web, sustained over time, not a single well-translated transaction.

The consequence of missing that layer is growing for every business, in every language. A 2025 Pew Research Center browsing-panel study found that users clicked a traditional search result in about 8 percent of searches when an AI summary was present, against 15 percent without one, and clicked a link inside the summary itself only about 1 percent of the time. Ranking is worth less than it used to be; being named inside the answer is worth more.

Translation is not the same input as identity

A useful way to see the gap: Bhashini can translate a shop's WhatsApp catalog message from Marathi into English in real time, and an answer engine reading the open web will still have no idea the shop exists, because the translation happened inside a private conversation, not on a structured, indexable page that names the business, its category, its location, and something that corroborates it is real. Fluency and legibility are different properties. India's language stack has scaled the first one enormously. It has not been asked to build the second.

The open question no one has measured yet

There is no public dataset, to date, that measures whether Indic-language MSME businesses are actually surfaced or cited when Indian buyers ask AI answer engines a vernacular question, "मेरे पास एक अच्छा प्लंबर सुझाएं" instead of "suggest me a good plumber near me." Neither MeitY, ONDC, nor any independent academic body has published sampled-query results comparing citation rates for businesses with structured, corroborated identity against businesses without it, broken out by language. That comparison would be the rigorous way to test the thesis this piece argues: draw a matched sample of MSMEs across a few scheduled languages, some with structured schema and consistent citations across the web and some without, run the same buyer-intent queries in each language across two or three major answer engines, and measure how often each group is named. Until that work exists, the claim that vernacular translation scale and MSME AI-visibility are two separate gaps is a reasoned inference from what these systems actually store, not a measured finding, and it should be read that way.

What is measurable today, and already published, is the shape of the inputs: a translation layer moving billions of calls, an MSME base of 8.7 crore mostly relying on referrals and print rather than digital discovery, and a commerce network whose own documentation admits its catalog data is not yet standardized for its own buyer apps, let alone for external answer engines. Those three facts, taken together, are consistent with a business being fluently translatable and still functionally invisible to the systems now doing a growing share of the recommending. They do not, on their own, prove the size of that invisibility. That is the boundary of what public data currently shows.

How to read the evidence

Bhashini is not failing at its job. It is succeeding at a job that was never the job of making a business machine-legible. Conflating the two, treating vernacular translation capacity as if it were vernacular AI-visibility, is the mistake this piece is trying to head off, because it leads to comfortable, wrong conclusions: that a Hindi-speaking shop owner whose government portals now work in Hindi is thereby findable by an AI engine answering a Hindi query. Nothing in the public record supports that leap.

The corrective is not a bigger translation model. It is a structured, corroborated, standing record of who a business is, in whatever language its buyers actually use, built and maintained independent of any single government scheme or commerce network. That is a different build than the one India has funded so far, and it is the specific thing a business can check for itself rather than infer from national statistics.

The evidence

Key findings, with their sources

  • Bhashini launched in July 2022 under MeitY's National Language Translation Mission and offers free translation, speech-recognition, and speech-synthesis APIs across all 22 constitutionally scheduled Indian languages, with more than 300 pre-trained AI models available to ecosystem partners.

    established Wikipedia, "Bhashini" (citing bhashini.gov.in and MeitY documentation).

  • Bhashini processed roughly 2.5 billion API calls in 2026, on pace to exceed 5 billion by year-end, and is integrated into DigiLocker, the successors to Aarogya Setu, and the UMANG super-app.

    emerging The Mobile Times, "India's AI Policy and Technology Landscape 2026", themobiletimes.com.

  • In September 2024, ONDC launched Saarthi, a Bhashini-powered multilingual reference app for businesses, initially covering five languages (Hindi, English, Marathi, Bangla, Tamil) with a stated plan to scale to all 22.

    emerging Wikipedia, "Bhashini"; ONDC/Saarthi coverage via The Daily Jagran and The Startup Spectrum, 2026.

  • The Ministry of MSME's Trade Enablement and Marketing (TEAM) initiative, under the World Bank-supported RAMP programme, targets onboarding five lakh micro and small enterprises onto ONDC (50% women-owned), backed by Rs 277.35 crore over FY2024-27 and implemented by NSIC.

    established IASGyan, "MSME-TEAM Initiative" (current-affairs summary of the Ministry of MSME scheme), citing RAMP/NSIC scheme documents.

  • As of the MSME Day 2026 announcement (27 June 2026), Udyam registrations had crossed 8.7 crore, an MSME base the government says supports the livelihoods of 38 crore people and contributes 31% of GDP, 35% of manufacturing output, and nearly 45% of exports.

    established Union Minister of State for MSME Shobha Karandlaje, MSME Day 2026 remarks, reported by newkerala.com.

  • MSME Day 2026 launched or upgraded PMEGP 2.0, SAMADHAAN 2.0, MSME Global Mart 2.0 (built to integrate with the ONDC ecosystem), and a new Testing Portal, all offering services in all 22 scheduled languages via Bhashini and the National Informatics Centre.

    established Drishti IAS, "MSME Day 2026-Udyami Bharat", drishtiias.com.

  • A 2025 survey found over 90% of MSMEs now accept digital payments, yet only 13% actively use digital marketing or e-commerce to reach customers, with roughly 70% still relying on in-person referrals, print advertising, or trade events.

    emerging YourStory, "Digital transformation in MSMEs: Adoption, gaps, and what's next", 2025 (reporting SIDBI-linked survey data).

  • ONDC's own technical documentation describes network-participant cataloguing schemas as wide and varied, requiring the protocol to build standardization work for its own buyer and seller apps to interoperate.

    contested ONDC Protocol Specifications (GitHub, ONDC-Official) and independent technical analysis via thewitslab.com.

  • Adding cited statistics, quotations, and authoritative sources measurably raised a source's visibility inside generated answers in tested AI engines.

    established Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed).

  • Users clicked a traditional search result in about 8% of searches with an AI summary present, versus 15% without, and clicked links inside the summary only about 1% of the time.

    established Pew Research Center, "Do people click on links in Google AI summaries?", July 2025, pewresearch.org (browsing panel).

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedBhashini's mandate, language coverage, and launch under MeitY; the MSME-TEAM/RAMP scheme terms; the Udyam registration and MSME economic-contribution figures; the MSME Day 2026 multilingual portal rollout; the GEO literature and the Pew click-through data.Wikipedia/MeitY documentation; IASGyan RAMP/NSIC scheme summary; Minister's MSME Day 2026 remarks (newkerala.com); Drishti IAS; Aggarwal et al., KDD 2024; Pew Research Center, 2025.
emergingBhashini's specific 2026 API-call volume and named app integrations; the Saarthi language-rollout sequence; the 13%-digital-marketing MSME adoption figure.Single-outlet trade press (The Mobile Times) not yet cross-corroborated by a MeitY dashboard; ONDC/Saarthi coverage via two India tech-news outlets; YourStory's secondary reporting of a SIDBI-linked survey.
contestedThe inference that Bhashini's and ONDC's existing infrastructure could be repointed at MSME AI-answer-engine visibility without new, separate work, and the claim that ONDC's catalog-schema variance specifically blocks machine legibility for external answer engines.No public dataset yet measures Indic-language citation rates for MSMEs with structured identity versus without; the schema-variance point is drawn from developer-level technical commentary, not an audited study.

Reference

Glossary

Bhashini
India's government-built, free, open language-AI platform, offering translation, speech recognition, and speech synthesis across all 22 constitutionally scheduled Indian languages, launched by MeitY in July 2022 under the National Language Translation Mission.
ONDC (Open Network for Digital Commerce)
A government-backed protocol that lets independent buyer and seller apps interoperate over a shared network, using registries that function like a directory service for routing transactions, not a public business directory addressed to the open web.
Udyam Registration
India's official MSME registration system; a business's Udyam number establishes formal MSME status for scheme eligibility and credit access, distinct from any public-facing description of the business itself.
Machine legibility
The degree to which a business's identity, offering, and credibility are structured, consistent, and corroborated in forms that generative AI answer engines can parse and cite, independent of how well the business is translated or how it transacts.
MSME-TEAM Initiative
A Ministry of MSME sub-scheme under the RAMP programme, targeting five lakh micro and small enterprises for ONDC onboarding with cataloguing, logistics, and workshop support, funded at Rs 277.35 crore over FY2024-27.

Straight answers

Frequently asked questions

Does Bhashini make Indian MSMEs visible to AI answer engines?

Not directly. Bhashini translates text and speech across 22 languages inside services that already know the user's identity, such as DigiLocker or an ONDC transaction. It does not create or store a structured, standing public record of a business that an answer engine can independently discover, corroborate, and cite. Translation fluency and machine legibility are different properties, and Bhashini was built for the first one.

Is ONDC a business directory that AI engines can read?

ONDC's registries work more like a routing directory for the network's own buyer and seller apps than a public, crawlable business directory. ONDC's own technical documentation notes that cataloging schemas vary widely across network participants and require standardization even for the network's internal use, which is a different problem from being structured for external AI answer engines.

How many Indian MSMEs actually have a real digital presence?

Udyam registrations crossed 8.7 crore as of June 2026, but a 2025 survey found over 90% of MSMEs accept digital payments while only 13% actively use digital marketing or e-commerce to reach customers. Payment infrastructure has scaled far faster than discoverability.

Has anyone measured whether AI engines cite MSMEs in Hindi, Tamil, or other Indian languages?

No public dataset currently does this. There is no published, sampled comparison of citation rates for structurally legible MSMEs versus non-legible ones across vernacular queries on major answer engines. That is an open measurement question, not a settled finding, and this piece treats it as such.

What would actually close the gap this article describes?

Not a bigger translation model. It requires a structured, corroborated, standing description of a specific business, in the language its buyers use, maintained independent of any single government scheme or commerce network. A Machine-Readiness Score is one way to check where a given business currently stands on that specific requirement.

Provenance

Sources

  1. Wikipedia, "Bhashini" (established)en.wikipedia.org
  2. The Mobile Times, "India's AI Policy and Technology Landscape 2026" (emerging)themobiletimes.com
  3. Drishti IAS, "MSME Day 2026-Udyami Bharat" (established)drishtiias.com
  4. newkerala.com, "8.7 Crore MSMEs Driving Economic Growth, Supporting Livelihoods of 38 Crore" (Minister Shobha Karandlaje, MSME Day 2026) (established)newkerala.com
  5. IASGyan, "MSME Trade Enablement and Marketing (TEAM) Initiative" (established)iasgyan.in
  6. YourStory, "Digital transformation in MSMEs: Adoption, gaps, and what's next", 2025 (emerging)yourstory.com
  7. ONDC Protocol Specifications, GitHub (ONDC-Official) (established, primary technical spec)github.com
  8. Independent technical analysis of ONDC cataloguing challenges, thewitslab.com (contested, blog-tier)blog.thewitslab.com
  9. The Daily Jagran, "ONDC Introduces Saarthi, Multilingual Reference App Powered By Bhashini" (emerging)thedailyjagran.com
  10. Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed, established)arxiv.org
  11. Pew Research Center, "Do people click on links in Google AI summaries?", July 2025 (established)pewresearch.org

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

About this analysis

This analysis sits inside Raveneye Global's ongoing research into machine readiness, the degree to which a business's identity is structured and corroborated enough for an AI answer engine to find, verify, and recommend it by name, in whatever language its buyers use. Bhashini, Udyam, and ONDC are each doing real work at real scale for the job they were funded to do, none of them was built to answer that narrower question for a specific business. A Machine-Readiness Score is Raveneye's method for reading it directly, across classic search, the local map pack, AI answers, and reputation, rather than inferring it from a government scheme's multilingual rollout.

diagnostic Surface Intelligence Audit A measured read of where a business stands across the surfaces buyers now use to find and choose it, in the languages they actually search in, with a ranked list of the corrections most likely to matter first. See how it works

The Machine-Readiness Score is free to request and specialist-reviewed, covering where a business currently stands across search and AI answers. It offers no guaranteed number and carries no obligation.