
Julian (Konstantin) Minder
About Me
I am a PhD student at DLAB at EPFL and currently an Anthropic Fellow in London. I am supervised by Prof. Robert West and co-advised by Prof. Ryan Cotterell (ETH Zurich).
I am passionate about understanding and improving artificial intelligence systems. My work focuses on understanding how models learn during pretraining and how this shapes their behavior through interpretability research. I aim to better understand how these systems work and how we can make them safer.
I completed my master's degree in computer science at ETH Zurich in 2024, following earlier studies in computer science and neuroinformatics at the University of Zurich. I wrote my master's thesis at EPFL under Bob West and Chris Wendler, investigating the mechanistic effects of fine-tuning language models (awarded the ETH medal).
I was a research scholar at MATS 7 working together with Clement Dumas under the mentorship of Neel Nanda to study the differences between base and instruct models.
Previously, I worked on Model Diffing, a research area focused on understanding the differences between two language models, which I mainly used to study the effects of finetuning. Currently, I'm moving towards understanding pretraining as a whole, as I believe it leaves strong priors that shape a model's later behavior, while its influence remains poorly understood.
Please feel free to reach out anytime!
For students:
If you're interested in doing a project with me, please reach out via email with the subject "[STUDENT PROJECT] ...", telling me a bit about yourself and your interests. I mainly look for students on master's level and/or with a strong background in machine learning. For EPFL students: Please additionally also apply via our lab application system.
News
Highlighted Publications
Synthetic Persona Pretraining: Alignment from Token Zero
Julian Minder*, Viktor Moskvoretskii*, Raghav Singhal*, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West
arXiv preprint
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Julian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt, Cameron Holmes, Robert West, Neel Nanda
Mechanistic Interpretability Workshop NeurIPS 2025 (🌟 Spotlight 🌟)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
Julian Minder*, Clement Dumas*, Caden Juang, Bilal Chugtai, Neel Nanda
NeurIPS 2025 | Mechanistic Interpretability Workshop NeurIPS 2025 (🌟 Spotlight 🌟)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel
NeurIPS 2025 (🌟 Spotlight 🌟)
Blog Posts
Synthetic Persona Pretraining: Alignment from Token Zero
Synthetic Persona Pretraining appends value-laden reflections during pretraining so alignment is installed from token zero rather than added only during post-training.
Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
Narrow finetunes leave clearly readable traces: activation differences between base and finetuned models on the first few tokens of unrelated text reliably reveal the finetuning domain.
What We Learned Trying to Diff Base and Chat Models (And Why It Matters)
This post presents some motivation on why we work on model diffing, some of our first results using sparse dictionary methods and our next steps.