Workshop Bias in Large Language Models

On 23 June 2026 we ran a full day workshop titled “Bias in Large Language Models” at TUM Campus Heilbronn.
The day started with an overview of the TUM-IAS Dieter Schwarz Fellowship project involving Gianluca Demartini, Maribel Acosta, and Longfei Zuo. Prof. Demartini presented recent work aimed at measuring political bias in LLMs as well as the initial plans for the fellowship project that aims at using predicted LLM bias to design routing strategies that direct an input prompt to the most relevant LLMs. This will be the focus of Longfei Zuo’s PhD project.
After this first session, during his talk, Prof. Alexander Fraser from TUM discussed several research papers published by his group on the topic of multi-lingual LLMs, cross-lingual transfer, and fairness across languages in current AI models. This included examples of gender bias in image generation models, and how language may serve as a proxy for culture. The talk triggered an interesting discussion about which data we should use to build LLMs.
Our second speaker was Prof. Philippe Cudré-Mauroux from the University of Fribourg and the Swiss National Science Foundation. He talked about extreme multi-label classification (XMLC) discussing research challenges related to taxonomy-based tagging for scientific publications. He explained which challenges we need to face when we use LLMs for XMLC including the challenges with multiple-choice questions, subsumption, and popularity bias. Building smaller, fine-tuned models is a possible approach to deal with these challenges. A strong focus is put on explainability of such classification decisions as they have to be auditable and actionable by humans. Prof. Cudré-Mauroux presented work aimed at improving the quality of generated human-readable classification justifications. The talk triggered an interesting discussion on how to build explanations that mimic the reasoning process of LLMs, and on the potential to use synthetic data to train LLMs to do XMLC better.
After lunch, our third speaker Daniel Matter from TUM talked about the research currently conducted at the Chair of Computational Social Science (CSS). He discussed how LLMs are currently used in CSS research both as instruments as well as simulated participants in research studies. He then discussed a very interesting problem: bias as reachability. This is defined as the ability for LLMs to cover different areas of a multi-dimensional space. He made the point that LLM bias should not be a single position in a bias space, but rather a probability distribution where LLMs are more likely to end up in certain bias areas than others. During the discussion, the audience mentioned the potential to use reasoning traces as a source of evidence to measure LLM bias.
Finally, the last talk of the day was by Prof. Marco Steenbergen from the University of Zurich. Prof. Steenbergen is a renowned political science researcher who has conducted influential work on deliberative democracy. In this talk, he discussed the possible roles LLMs could take during a deliberation - from participants, summarizers, assistants, to moderators. During the discussion, the overarching question was whether LLMs should participate to democratic deliberations, who should decide that, and, if so, in which roles and with which level of transparency.
At the closing of the workshop, the group discussed possible next steps which included the identification of links between the different presentations and research streams.