When considering explanations for the behavior of an AI model – for example explanations of the kind “what did the model consider important when producing this output” – confirmation bias can lead us to believe a machine is trustworthy because a few explanations comply with our beliefs.
To prevent such situations, the HUE project attempts to mitigate confirmation bias by investigating how explanations connect to human-understandable concepts. If successful, our method would allow us to ‘x-ray’ AI and verify whether it complies with our requirements, as opposed to exhibiting harmful behaviors. Building on an existing conceptual framework, this work will connect different disciplines by testing and extending the framework from Medical AI to Natural Language Processing and Computer Vision. Its application to medical use cases is already of interest to industrial partners.



