The rise of artificial intelligence (AI) has sparked debates around its potential risks in various domains, yet the cultural and economic threats posed by large language models (LLMs) are often overlooked. As these models, largely owned by a few multinational corporations and trained on a limited number of languages, dominate the market, developing countries—where many languages are spoken by minority communities—find themselves at a disadvantage.
Since the launch of OpenAI’s ChatGPT in November 2022, industry leaders and policymakers have touted AI’s ability to revolutionize sectors like healthcare and education while boosting productivity and income growth. However, the benefits of these advancements remain unequally distributed.
Many developing nations struggle to access AI technologies, primarily due to linguistic biases. Most LLMs are trained on just around 100 languages from the over 7,000 languages spoken globally, with English and four other “high-resource” languages—Chinese, Spanish, French, and German—accounting for nearly all training data. This linguistic imbalance risks cultural and economic isolation in the short term and threatens to stifle growth and linguistic diversity in the long run.
Furthermore, a troubling irony has emerged: as wealthier nations rapidly adopt AI, leading to the expansion of robust data centers, users in developing countries face increased challenges accessing the internet and digital tools. The demand for advanced AI technologies has driven up the costs of smartphones, pricing them out of reach for millions.
### The Need for Localized Models
To mitigate these issues, it is crucial for LLMs to better reflect local languages and cultures. Without this, communities risk losing their unique vocabularies shaped by historical events, mass migrations, and political legacies. Additionally, the lack of linguistic diversity in AI systems could hinder adoption rates. For instance, digital financial services in Peru need to communicate in colloquial language for consumers to engage effectively, while healthcare services in Vietnam require clarity achievable only with models trained in native languages.
Mobile network operators possess a unique position to develop “localized” language models, thanks to their extensive developer networks and data processing capabilities. As trusted partners to governments, they can also address a key barrier: the scarcity of training data for minority languages. Governments hold vast reserves of sensitive data that can be leveraged to create these models.
For example, Kyivstar, Ukraine’s largest mobile operator, has partnered with the Ministry of Digital Transformation to create the nation’s first large-scale language model tailored to Ukrainian and other languages used in the country, such as Russian, Bulgarian, and Crimean Tatar. Complicating this initiative are cultural sensitivities surrounding a substantial body of literary work produced during Russian rule, which has been deemed culturally inappropriate. Consequently, four advisory committees are guiding the selection of training texts, ensuring they align technically, legally, culturally, historically, and linguistically.
### Initiatives in Africa
The Global System for Mobile Communications Association (GSMA), which I lead, is pursuing a similar strategy across numerous Sub-Saharan African countries, including Nigeria, Madagascar, Togo, and the Democratic Republic of Congo, through its African Large Language Models initiative. This initiative brings together “The Big Six” mobile networks on the continent—Airtel, Axian Telecom, Ethio Telecom, MTN, Orange, and Vodacom—along with other telecommunications companies, public and private stakeholders, academia, and civil society organizations, to develop scalable language models based on African languages.
To support the growth of AI startups delivering social and economic benefits to marginalized groups, the GSMA has launched the EmergingTech program, funded by the UK Foreign, Commonwealth & Development Office. The inaugural cohort includes ToumAI, which is creating voice application interfaces in Moroccan dialects to assist rural users with low literacy levels in accessing digital financial services through simple voice commands, and Wiseyak, which is developing a multilingual AI solution to bridge the digital divide in Nepal.
These initiatives are increasingly vital as the world shifts from command-based AI to self-learning AI. Digital services—spanning banking, healthcare, and education—are progressively integrating AI into their frameworks to enhance user support and will become essential to people’s lives and livelihoods. Companies are also benefiting from open-source models.
Investing in local AI fosters a positive cycle benefiting developing economies by generating a pool of technical talent critical for achieving digital independence. With growing concerns about data sovereignty and processing, these economies must enhance operational oversight and governance measures surrounding self-learning AI deployment. This also necessitates collaboration with developers and mobile network operators to build models for minority languages that effectively meet community needs.








