How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies

Paper details

Trustworthy White
Responsible White
Language models White
Fairness White
Bias White

With the widespread availability of Large Language Models (LLMs) since the release of ChatGPT and increased public scrutiny, commercial model development appears to have focused their efforts on ‘safety’ training concerning legal liabilities at the expense of social impact evaluation. This mimics a similar trend which we could observe for search engine autocompletion some years prior. We draw on scholarship from NLP and search engine auditing and present a novel evaluation task to assess stereotyping in LLMs.

Our findings indicate a lack of attention by LLMs under study to certain harms classified as toxic, particularly for prompts about peoples/ethnicities and sexual orientation. Mentions of intersectional identities trigger a disproportionate amount of stereotyping. Finally, we discuss the implications of these findings about stereotyping harms in light of the coming intermingling of LLMs and search and the choice of stereotyping mitigation policy to adopt. We address diverse stakeholders calling for accountability and awareness concerning stereotyping harms, be it for training data curation, leader board design and usage, or social impact measurement.

Reference:

Alina Leidinger and Richard Rogers. 2024. How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies. In Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society (AIES '24).

Other papers

CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
Let's treat language model tests as professionally as we treat exams
Tackling Language Modelling Bias in Support of Linguistic Diversity
Diversity and language technology: how language modeling bias causes epistemic injustice
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
Quantifying Context Mixing in Transformers