Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Characterizing the linearity of LLM representations

Mechanistic interpretability

There is a lot of literature showing that LLMs/transformers/neural networks are more linear than we think:

The goal of this project would be to characterize the linearity of transformer/LLM representations and see how this can help us to build simple models for LLM/transformer behavior/output generation/training, to which we can then apply the rich mathematical literature that exists around linear mappings (dynamical systems, Markov chains, etc.), thus helping to provide a more comprehensive theory of how LLMs/transformers work.

About the project

There is a lot of literature showing that LLMs/transformers/neural networks are more linear than we think:

The goal of this project would be to characterize the linearity of transformer/LLM representations and see how this can help us to build simple models for LLM/transformer behavior/output generation/training, to which we can then apply the rich mathematical literature that exists around linear mappings (dynamical systems, Markov chains, etc.), thus helping to provide a more comprehensive theory of how LLMs/transformers work.

Theory of change

If we are able to develop a simple linear model of how LLMs/transformers work, we will be able to apply the rich mathematical literature that exists around linear mappings (dynamical systems, Markov chains, etc.) to LLMs and transformers and gain a deeper understanding of their behavior and training dynamics

Your role

  • Develop proposal
  • Conduct literature review
  • Run experiments
  • Develop theory

Prerequisites

  • Experience with mech interp

Location preference

No

Application question(s)

Do a brief literature review on linearity in transformers/LLMs. Why might we expect LLMs/transformers to behave very linearly?

About the mentor

Thomas Jiralerspong

Thomas Jiralerspong

Bengio Lab, MATS Anthropic

View profile

Hey! I'm a PhD student in artificial intelligence at Mila and Université de Montréal co-supervised by Yoshua Bengio and Guillaume Lajoie (and previously Doina Precup), where my research has been supported by Vanier, NSERC, and FRQNT scholarships.

I'm currently an Astra Fellow working with Dan Mossing (Anthropic) on persona controllability in LLMs. I'm also a Principal Investigator at Algoverse AI Research, where I oversee 20 research teams, and a research mentor for MARS V (Meridian Cambridge), where I mentor 7 researchers on AI safety projects. Previously, I was a Research Fellow at Anthropic mentored by Trenton Bricken, and an Astra Fellow in the Google DeepMind stream working on AI control/monitoring.

My research interests include mechanistic interpretability and LLM personas

Similar projects