Language models Plum
Fairness Plum
Bias Plum
Inclusive Plum

Using Language Sciences for Social Good

Blog item

Language technologies, like ChatGPT, are developing quickly. But we don’t really know how they work. And they are not always designed to help society.

adapted from Nikki Weststeijn‘s post on https://resources.illc.uva.nl/illc-blog/using-language-sciences-for-social-good/

 

Illustration by Arco Mul

 

Language Sciences for Social Good

If ChatGPT gives us an answer to a factual question, we have no idea if this answer is really true. Open AI, the makers of ChatGPT, bear no responsibility that their chatbot and text generator will reply truthfully. Similarly, if an employer uses some type of language technology to screen the CV’s of applicants, there is no promise that the technology will not be biased in some way. This is because these technologies are all self-learning algorithms. It is a type of artificial intelligence that is trained by seeing sets of exemplary inputs and correct outputs, which allow it to develop some strategy for coming up with outputs on its own. The strategy it learns is opaque, so it is unclear why exactly the algorithm will give one particular answer over another.

Floris Roelofsen, professor at the Institute for Logic, Language and Computation (ILLC), is director of a new large-scale research project called ‘Language Sciences for Social Good’ (LSG) at the University of Amsterdam. The idea behind this overarching research project is to make sure that language technology will also be used for social good, and not just guided by cost-efficiency and result-driven goals. Quote Roelofsen: “Language technology is booming, it’s really transforming society. With ChatGPT it’s very clear, but this development has been going on for the past 25 years. We have been using language technology when we search on Google, when we translate pieces of text, or when we use voice assistants on our phone.” Language technology has a lot of impact, but in areas where it could potentially have a positive impact, it is still lacking, according to Roelofsen. The LSG project has a threefold goal: to use more responsible methods in developing language technology, and make sure that language technology can contribute to a safe and inclusive society.

 

An example of LSG research (described below) is the work of Ekaterina Shutova (Associate Professor of Natural Language Processing at the ILLC). She does research on how we can prevent technologies like ChatGPT from producing hate speech or otherwise discriminatory content.

 

Shutova’s project, or “How to prevent AI from making hateful, biased and misinformed comments”

We can ask ChatGPT to translate something for us, we can ask for some factual information, we can even ask it to write an entire essay or a poem. But what if the chatbot creates a poem that is extremely sexist? Or if we ask it to predict who is most likely to commit a crime and it will give a racist answer?

ChatGPT has built-in filters that are supposed to make sure that it does not generate harmful answers or content. One of the ways in which ChatGPT is trained to create safe content is by learning with human feedback. After the initial training phase, the model is tested by a human who can tell the algorithm when the language produced is inappropriate.

However, ChatGPT can still be provoked to make hateful comments. If the systems parameters are for example set such that the chatbot should behave like a historical figure, like Donald Trump, the outputs can be much more hateful. Users can also try to find clever ways to prompt the chatbot in such a way that it creates a hateful comment.

If a language model that is used in decision making is biased against certain social groups, this could also have bigger consequences. Shutova: “If we actually want to use these types of algorithms in decision making, for example for recruitement or for medical diagnosis, then we need to systematically understand what kinds of biases these models show.”

One of the lines of research supervised by Shutova is to see if these biases exist and then to find the sources of the bias. Together with her PhD students, Shutova analyzes what stereotypical properties a language model associates with certain social groups and if there are certain emotions, such as fear or anger, associated with that group, for example with immigrants. Shutova: “And you see that there are really no surprises there: The AI mimics the biases that already exist in society.”

The next step is to find ways of decreasing the amount of bias in a model. One possible way is to look at the training data and filter out data examples that have offensive language in them. But this is harder when we want to prevent complex implicit biases. Shutova: “With explicitly offensive language, you can just filter it out. And everybody agrees that we wouldn’t want to have that. But a lot of things are really subtle. It’s really how you talk about a social group in lots of different ways across lots of different documents.”

Shutova and her student Vera Neplenbroek are therefore working on a different approach. This is to try to get rid of implicit biases after the model has had its initial training. This method tracks what specific examples in the training data caused the model to inherit a certain bias and then omits these examples at a later training stage. Shutova: “We have shown that omitting these examples during training really helps to reduce the bias. So, we hope that this is an approach that would be picked up.”

A problem with most of the existing debiasing techniques is that it is focused mostly on gender bias. Shutova’s method is designed to work for eliminating different kinds of biases, including gender bias, but also biases that are related to your level of education or economic status for example.

Read other blogs

Discover our researchers’ blogs—icons by each post indicate its themes.

Language models White
Bias White

Going beyond a mathematical investigation of bias