Explore sparse representations in LLMs using SAEs, LoRA, latent geometry analysis, and formal verification tools. We'll build toy models, benchmark structured priors, and probe "deceptive" features in compressed networks.
About the project
This project investigates the latent geometry of LLM representations and their implications to interpretability and formal verification. Building on recent research into structured priors, variational SAEs, subliminal learning, and compression artifacts, we aim to explore how concepts are organized in hidden space — and how we can validate these representations using formal guarantees.
Key directions:
- Analyze latent structures using spectral and geometric tools (e.g., PCA, LDA, Manifold learning, Wishart theory)
- Prototype structured autoencoder variants for efficient probing (e.g., Probabilistic SAE, crosscoders)
- Investigate subliminal learning and deceptive features via latent superposition and composed phenomenon
- Apply symbolic verification tools to detect representation stability
- (Information Theory Track) Protocol development for agentic AI systems (e.g., cooperation, security concerns)
I'm also open to any research topics you're interested in!
Theory of change
Interpretability tools like SAEs and LoRA offer a middle ground between symbolic verification and end-to-end alignment. This project focuses on making representations more structured, verifiable, and understandable, enabling scalable monitoring and steering of LLMs. Our theoretical framing (e.g., compression bounds, latent modularity) aims to clarify when and how models store dangerous or deceptive features — a key bottleneck in understanding model generalization and failure modes.
Your role
Mentees will:
- Design and run experiments on toy models and transformers
- Contribute to code, analysis, and interpretation
- Co-author research writeups and/or blog posts
- Participate in weekly meetings and async discussion
They will be treated as junior collaborators, with substantial autonomy in choosing sub-questions and implementation strategies.
Prerequisites
- Comfortable with Python and PyTorch or JAX
- Familiarity with AI interpretability and transformers
- Basic linear algebra and probability
- Ability to read ML papers and experiment independently
- Bonus: Interest in information theory, spectral methods, physical models, or symbolic verification
Time commitment
10–20 hrs/week
Location preference
Most work is async and flexible.
Application question(s)
- Briefly describe a project where you implemented or modified an ML model/theoretical research. What went well, and what was challenging? (250 words)
- How would you detect “subliminal” features (features that store knowledge but don’t affect output) in a sparse autoencoder trained on transformer activations? (200 words)
- (Optional) Link to any relevant code or writing sample.
About the mentor

Yuxiao is an independent researcher in mechanistic interpretability. Before she was a postdoc at the Basque Center for Applied Mathematics (BCAM) and an AI Safety researcher at the Beneficial AI Foundation (BAIF). She was also a SERI MATS scholar in 2022 and a MATS scholar in 2023 both Summer and Winter tracks. Her research interests include information theory, probabilistic frameworks, and their applications for building more theoretically sound and trustworthy AI systems. She has a background in statistical inference, machine learning, and deep generative models. She completed her PhD in Electronic Engineering at Tsinghua University and has mentored research teams with SPAR and Algoverse.