Active Education & Skills History, Languages & Philosophy

Towards Globally Equitable Language Technologies (EQUATE)

In plain English

AI plain-English summary

Voice assistants, translation apps, and text-prediction tools work fluently for English and Mandarin, but fail or simply do not exist for the vast majority of the world’s 7,000-plus languages. This project tackles that imbalance head-on. The core problem is that building language technology requires vast amounts of digitised text and speech data, which most languages lack. As a result, roughly 7.9 billion people are unevenly served: those in the Global North get powerful tools, while speakers of low-resource languages get nothing. The researchers will first create an index that ranks every living language by its “readiness” for technology—how much data exists, how well models can handle it. Then they will design new machine-learning methods that work with far less data, are more transparent, and can be adapted modularly across languages. Crucially, they will work directly with local language communities to build evaluation resources that reflect real-world use, not just lab benchmarks. If successful, this could shift the field of natural language processing from a de facto English-only standard to one where a farmer in rural Senegal or a teacher in rural Nepal can use the same quality of speech recognition and translation tools as a user in London. The impact is not immediate consumer convenience; it is about making digital infrastructure genuinely global.

View original technical description
Language technologies can now offer effective support to communication, education, healthcare, and many other aspects of human life. Yet, these technologies are not distributed equally. They are only available for a small part of the world's 7.9 billion population, mainly those living in the Global North. This is because the resources needed for them are limited or lacking for the vast majority of the world's over 7,000 living languages. This situation has significant scientific and socioeconomic consequences. The goal of our project is to investigate the methodological challenges in the development of globally equitable language technologies and to design transformative approaches to overcome them, with the overall aim of creating a realistic methodological basis for multilingually equitable NLP. We will first develop an understanding of the (in)equalities in language technologies, and produce a novel index that profiles the world's languages and language populations in terms of their readiness for language technologies. We will then develop new methods for multilingual NLP that address critical aspects of equity (ranging from sample efficiency to modularity, model compactness, transparency, fairness and others), along with a novel unified approach that integrates such methods to support NLP at different levels of readiness. Working with local language populations, we will also produce new, equity-aware evaluation resources that are representative of the world's low-resource languages in terms of geographic regions, linguistic characteristics, and NLP readiness. Our novel methodology will be evaluated on downstream NLP tasks as well as in the context of useful real-life applications in languages that are currently under-served by them. This project can transform the way we approach multilingual NLP, and substantially improve our understanding of how language technologies can be made fair and inclusive at a global level.

View the original record at the funder ↗

Researchers

Anna Korhonen (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Unmute: Opening Spoken Language Interaction to the Currently Unheard
Cross-Lingual Embeddings for Less-Represented Languages in European News Media
Critical language barriers: Understanding the role of online translation tools in high-stakes professional sectors
QT21: Quality Translation 21
MTStretch: Low-resource Machine Translation

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.