I am a master's student in computer science at TU Berlin,
supervised by
Qianli Wang.
I wrote my bachelor's thesis with Qianli Wang and
Nils Feldhus.
My recent research focuses on the interpretability of
multilingual language models: whether a concept is represented
consistently across languages, and what a model actually relies
on when it makes a prediction. Much of my work uses counterfactual
explanations, minimal edits to an input that flip a model's
prediction. My recent papers improve self-generated
counterfactuals through preference optimization in the
multilingual setting and through iterative feedback at
inference time.
Looking for internship and PhD
positions starting in 2027. Feel free to reach out by
email.