Project details

Language models White
Fairness White
Bias White
Inclusive White

Towards Pluriversal Language Technology

AI-based language technology is currently limited to 3% of the world’s most widely spoken, financially and politically backed languages. Recent efforts have sought to address this “digital language divide” by extending the reach of large language models to “low-resource languages.” Despite benevolent motives, most of these efforts produce flawed solutions that adhere to a hard-wired representational preference for certain languages, which we conceptualize as “language modeling bias” (LMB).
LMB is a under-studied form of bias where language technology by design favors certain languages, dialects, or sociolects with respect to others. LMB can result in systems that, while being precise regarding languages and cultures of dominant powers, are limited in the expression of socio-culturally relevant notions of other communities, thereby contributing to discrimination of marginalized language communities.

In this project, we study LMB both in terms of its causes (e.g., that technology developer communities tend to apply surface-level understandings of diversity which do not do justice to the more profound, meaning-level differences that languages embody) and ethical and political implications (e.g., how it can lead to harms for marginalized speakers).

To address these issues, we propose that Arturo Escobar’s concept of “pluriversality” can act as a guiding compass for language technology development. As a recent example, we present the LiveLanguage initiative and discuss how far values of pluriversality are here embedded: Here attempts were made to overcome LMB in lexical resources not only through design but also through alternative methodologies (e.g., collaboration with local language communities).

Papers related to this project

Tackling Language Modelling Bias in Support of Linguistic Diversity
Diversity and language technology: how language modeling bias causes epistemic injustice