How robust and reliable can we expect language models to be?

This paper investigates robustness in how we measure performance of language models like ChatGPT.

If you were to ask a language model two very similar questions, e.g., ‘Do you like this movie review?’ vs. ‘Tell me if you like this movie review.’, you’d expect to get the same answer. But is this always the case?

Far from it! In this paper we show that responses from language models vary widely depending on how you phrase your instruction, making responses unreliable. Moreover, the simplest instructions don’t always work best. Sometimes an instruction that we find overly complicated, e.g., ‘Might this movie critique be found positive?’, can give you much better results than ‘Do you like this movie review?’.

Lastly, the same instruction that works very well for one language model can give you very poor results for another model. This raises questions for how we measure progress in AI. If performance of any one language model depends so much on the question asked (and the questions are rarely reported), it’s hard to compare two models fairly or rely on existing results. We end this paper with actionable suggestions for more reliable, reproducible performance evaluations.

Reference: Leidinger, Alina, Robert van Rooij, and Ekaterina Shutova. "The language of prompting: What linguistic properties make a prompt successful?." Findings of the Association for Computational Linguistics: EMNLP 2023. 2023.

Other papers

How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
Let's treat language model tests as professionally as we treat exams
Tackling Language Modelling Bias in Support of Linguistic Diversity
Diversity and language technology: how language modeling bias causes epistemic injustice
Are LLMs classical or nonmonotonic reasoners? Lessons from generics