CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

Paper details

Trustworthy White
Responsible White
Language models White

In this paper we introduce the “CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts” dataset.

It is designed to evaluate how Language Model replies vary in response to value-sensitive topics in multiple languages. We create a hand-crafted, multilingual dataset of value-laden prompts which address specific socially sensitive topics, including LGBTQI rights, social welfare, immigration, disability rights, and surrogacy. CIVICS is designed to generate responses showing LLMs’ encoded and implicit values.

Our experiments show that Language Models refuse to answer more often when asked in English. Moreover, specific topics and sources lead to more pronounced differences across model answers, particularly on immigration, LGBTQI rights, and social welfare. Our dataset aims to serve as a tool for future research, promoting reproducibility and transparency across broader linguistic settings, and furthering the development of AI technologies that respect and reflect global cultural diversities and value pluralism.

The CIVICS dataset is available at: https://huggingface.co/CIVICS-dataset.

Reference:

Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh, Alexandra Sasha Luccioni, and Margaret Mitchell. 2024. CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models. In Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society (AIES '24).

Other papers

How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies
Let's treat language model tests as professionally as we treat exams
Tackling Language Modelling Bias in Support of Linguistic Diversity
Diversity and language technology: how language modeling bias causes epistemic injustice
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
Quantifying Context Mixing in Transformers