Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Context-induced belief geometry in LLMs

Mechanistic interpretability

Toy-scale transformers trained on hidden Markov models develop representations of 'belief geometry' in their activations -- a convergent structure encoding distributions over hidden world states. This project investigates whether we can induce similar structure in production-scale LLMs purely through prompting and using linear probes to analyse the resulting representations.

About the project

Please see the live project document:

https://docs.google.com/document/d/1vqEXWRG7fS3_v2IRwr1PNuyg1KUcCtynxTIdzJ9O6Lo/edit?usp=sharing

Feel free to add comments!

Theory of change

The computational mechanics program suggests that the internal representations of intelligent systems can be understood by studying structures relevant for optimal prediction. In its most ambitious form, this framework applies to both biological and artificial intelligence -- offering a principled, substrate-independent account of cognition. Such a framework could advance AI safety on multiple fronts, from white-box scalable oversight to identifying natural latents as targets for alignment.

This project tests a key premise of that vision: whether signatures of computational mechanics can be identified in production-scale transformers. A positive result would provide evidence that these theoretical tools can ground practical interpretability research; a negative result would help clarify the framework's limits and inform when alternative approaches are needed.

Your role

Mentees will be responsible for guiding the day-to-day progress of the project, which will consist of writing code, running experiments, and analysing / communicating results. I also encourage mentees to take an active role in setting the direction of the project by either critically evaluating my proposals and / or by identifying other opportunities.

The team will have a regular 1h call each week where we discuss progress and set the direction for the following week. I suggest that mentees meet at least once prior to the team meeting to share results and set the agenda for the team discussion. Throughout the week, I will be available via Slack, where I look forward to regular updates and discussions on progress. You can expect a reply within 24h, and likely sooner for exciting progress and for addressing blockers.

Prerequisites

  • Comfortable with Python
  • Comfortable with basic linear algebra & probability
  • Comfortable with the transformer architecture
  • Familiar with Pytorch & TransformerLens
  • Familiar with common LLM APIs e.g., HuggingFace

Application question(s)

What research project are you most proud of and why? Projects with publicly available output are preferred (max 400 words).

Define a new HMM and describe why this process is interesting. Provide the transition matrices and a visualisation of the corresponding belief geometry (max 200 words).

About the mentor

Xavier Poncini

Xavier Poncini

Simplex

View profile

Xavier is a researcher at Simplex, an AI safety non-profit working toward a principled understanding of the internal structure of machine learning systems. He is interested in drawing on ideas from physics -- particularly a field called computational mechanics -- to study convergent structures that emerge in systems performing optimal prediction, and to identify these structures in production-scale transformers. Before moving into AI safety, Xavier worked as a post-doc in mathematical physics, studying abstractions of two-dimensional models of statistical mechanics.

Similar projects