Investigating the Text Spotting Bias of CLIP Models in Multilingual and Cross-Lingual Context
Contrastive Language-Image Pre-training (CLIP) is a multi-modal model with images and text saved in the same embedding space to be able to associate them with each other.
The standard use-case of CLIP models is visual classification, by taking a set of possible captions and an image, scoring the similarity and outputting the likelihood of each caption belonging to the picture.
There are papers showing that CLIP models tend to repeat text superimposed (or visibly depicted) on images rather than classifying what actually can be seen in these images.
For example, an image of a cat with the text "dog" written on it would rather be classified as a dog than a cat.
This text spotting bias is also called "parrot" and its consequences are either an impaired usability of the CLIP models or opens up new possibilities for downstream tasks like integrated OCR.
This thesis aims to investigate how this bias behaves with different languages, either multilingually or cross-lingually.
A series of experiments tests the bias on different models for different languages (non-english pairs of caption text and image text) as well as between languages (different languages for caption and image text).
The models differ in their image and text encoders as well as their training data sets.
Furthermore, the thesis explores the behavior of caption generating models regarding this bias.
The expected scientific contribution of this thesis is to identify possible limitations for vision language tasks as well as try to find mitigating factors or alternative use-cases for the bias.
Research Questions:
RQ1: Does the parrot bias generalize across multiple languages?
RQ2: Does the parrot bias generalize cross-lingually?
RQ3: To what extent are caption generation models affected by the parrot bias?
RQ4: Do the distribution of training data or the embedding geometry of multilingual models replicate or suppress parrot-inducing patterns?
| Attribute | Value |
|---|---|
| Title (de) | Investigating the Text Spotting Bias of CLIP Models in Multilingual and Cross-Lingual Context |
| Title (en) | Investigating the Text Spotting Bias of CLIP Models in Multilingual and Cross-Lingual Context |
| Project | |
| Type | Master's Thesis |
| Status | started |
| Student | Jacob Fehn |
| Advisor | Katharina Sommer |
| Supervisor | Prof. Dr. Florian Matthes |
| Start Date | 03.09.2026 |
| Sebis Contributor Agreement signed on | 02.09.2026 |
| Checklist filled | |
| Submission date | 03.03.2027 |