Which Bangla Will Bangladesh’s National AI Speak?
By Mohammad Rahmatullah, Executive Editor, Tha South Asian Story
21 July 2026·6 min read
The country’s proposed language model could bring government closer to its citizens—or teach public services to misunderstand them.
Bangladesh is preparing to build a national Bangla language model. Before asking how intelligent it will be, the country must ask a more democratic question: whose Bangla will it understand?
The government’s draft National Artificial Intelligence Policy 2026–2030 makes an ambitious promise. It proposes a locally hosted, open-source national language model led by the Bangladesh Computer Council. The model would draw on archives of history, literature, public knowledge and cultural materials. The policy also plans AI-powered chatbots that can help citizens find government services, complete forms and receive answers in Bangla. Public consultation on the draft ended on 8 February 2026, and the committee is now reviewing responses before preparing the final policy.
This is an important step. Most powerful AI systems are built mainly around English and other globally dominant languages. They often understand Bangladesh through incomplete, foreign or poorly translated data. A national model could reduce that dependence. It could make public information easier to find, help civil servants process documents and allow citizens to use digital services in their own language.
But “Bangla” is not one uniform voice.
The language found in official letters, school textbooks and television news is not always the language spoken at home, in a village market, on a fishing boat or at a tea stall. A person from Sylhet may speak differently from someone in Barishal. Chittagonian speech may be difficult for a system trained mostly on standard written Bangla. Rangpuri, Noakhali and other regional forms have their own words, sounds and sentence patterns. Whether we call them dialects, regional varieties or languages, they are part of how millions of people actually communicate.
Then there are Bangladesh’s Indigenous languages, which cannot simply be treated as unusual forms of Bangla. They carry distinct histories, identities and systems of knowledge.
The draft policy recognises Bangladesh’s multilingual character. It says AI systems should support Bangla and “other nationally relevant languages,” protect cultural diversity and avoid linguistic marginalisation. It also requires public-sector training data to be inclusive and representative. These are welcome commitments. Yet the draft does not clearly explain which languages and regional forms will be included, how their speakers will participate, or how performance across different kinds of speech will be measured.
That gap matters because a language error inside a government system is not always a harmless mistake.
Imagine a woman in Sunamganj trying to ask a chatbot about a widow allowance in the Sylheti she uses every day. The system misunderstands a key word and directs her to the wrong form. A farmer describes damaged crops in a regional variety that the model reads incorrectly. An Indigenous citizen tries to report a land problem but finds that the service accepts only standard Bangla text. A patient gives health information in Chittagonian, and an automated tool changes a number, a symptom or the meaning of an ordinary expression.
In each case, the machine has made a linguistic mistake. But the citizen may lose time, money, treatment or access to a legal right.
Recent research already shows the difficulty. A 2025 study tested leading language models on medical translation involving Sylheti and Chittagonian. Even the better-performing systems continued to make errors involving medical terms, numbers and idioms. The researchers warned that standard-language health information can create barriers and safety risks for speakers of regional varieties.
The lesson is simple: a chatbot that works well in standard Bangla has not automatically passed the test of public accessibility.
This is especially serious because the draft policy places AI near high-stakes areas such as welfare, healthcare, education, justice and access to essential services. To its credit, the policy promises human review, the right to challenge significant automated decisions and safeguards against discriminatory outcomes. But these protections will become meaningful only when a citizen can report that the system misunderstood their language—and receive help without having to prove that their way of speaking is legitimate.
Bangladesh should therefore treat linguistic inclusion as part of AI safety, not as a cultural decoration added after the model has been built.
The national model should be trained and tested region by region. Accuracy reports should not provide only one national score. They should show how the system performs in Sylheti, Chittagonian, Rangpuri, Noakhali and other widely used forms of speech. Where data are limited, the government should say so openly rather than presenting the model as universally capable.
Communities must also help decide how their language data are collected. Native speakers, teachers, linguists and local organisations should take part in recording, checking and correcting material. People whose speech becomes training data should know how it will be used. They should give informed consent and, where substantial labour is involved, receive fair payment. A language should not be extracted from a community and then returned to it as a government product over which that community has no control.
There is already useful groundwork. A Bangladesh Computer Council initiative has created a digital repository containing materials from 42 Indigenous languages, including recorded speech and pronunciation data gathered from native speakers. Such resources could support more inclusive language technology, but preservation alone is not enough. Communities must remain partners in deciding whether and how these materials enter an AI system.
Most importantly, no public application should be rejected simply because a chatbot cannot understand the citizen. Every AI-assisted service must provide an easy route to a trained human official. Citizens should be able to correct a mistranscription, change an AI-filled form and appeal a decision without facing extra delay or cost.
The government’s proposed oversight body should include representatives of regional and Indigenous language communities. Its annual reports should publish language-specific error rates, complaints and corrective actions. A national model cannot honestly be called inclusive if its failures remain hidden inside a single technical score.
Bangladesh has good reason to build its own AI. Digital sovereignty matters. So does the ability to create technology from the country’s own knowledge, culture and public institutions. But sovereignty cannot mean teaching a machine only the language of the capital and asking the rest of the country to adjust.
Standard Bangla has an essential place in national life. It connects people across regions and carries a history for which Bangladeshis made extraordinary sacrifices. That history should make the country more alert—not less—to the danger of making some voices harder for the state to hear.
A national language model should not merely speak the language of government. It must learn to hear the people who speak differently from it. Otherwise, Bangladesh may build a Bangla AI that can read every government file but cannot understand the citizen standing at the door.