Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Understanding LLM generalization through fine-tuning

Alignment Behavioral evaluation of LLMs

Alignment is, at its core, a generalization problem. Thus, we would really like to build a better understanding of how LLMs generalize. By fine-tuning models on small datasets, we can study surgical interventions on models to see how far new information generalizes.

About the project

Detailed project proposal: https://docs.google.com/document/d/1-Pxn2y8ZZKpM-eZlyHS0Ayri6t1aq7SMr-AJggGAwUs/edit?tab=t.0

Here’s a bunch of questions about LLM generalization that I think could well be studied by simple fine-tuning experiments. I expect a good project will aim to answer one of these properly, but expect it to be fairly likely that we will explore multiple of these questions initially, or pivot from one to another if we move quickly or decide that one of them is unpromising.

Knowledge to action: Here is a list of settings where you fine-tune a model with some new information/knowledge/opinion, and see how far it generalizes to downstream model behavior. For instance, if you fine-tune a model on documents about factory farming, does this impact what recipes the model will recommend you?

How far do different models personas generalize? If you train the model to give a somewhat more socially conservative answer in some case, how does this generalize?

Can models tell what they’ve recently been fine-tuned on? There is some evidence that models can remember what they have been fine-tuned on most recently. This could be a concern in cases where we want to fine-tune models to believe in a certain evaluation scenario, or different kinds of honeypotting setups. We would like to test this.

This project is a synthesized version of Ryan Greenblatt's previous project proposal: https://docs.google.com/document/d/17qXPbqbWsuKazXjtFAzvwlSKCsXq69NlWASzTtAk2Ec/edit?tab=t.0#heading=h.gimq5t8clw6i

Theory of change

This project falls under the category of research which is "build better understanding of how LLMs work in ways that will be critical to ensuring robust alignment". More specifically, alignment is at its core a generalization problem. And while there are capabilities externalities to a project studying generalization, we think a better understanding of generalization in LLMs (in particular in regards to the generalization of different "personas" of the model) is a safety differential project.

Your role

See proposal for specific experiments that we expect mentees to run.

At a high level, we expect mentees to be quite independent in setting up and running experiments. We will be responsive on Slack, but won't be very available to help set up experiments/write code/other technical unblocking.

The experiments will be doing simple SFT runs on LLMs using either to OpenAI API or open-source libraries.

Prerequisites

Able to run small experiments on LLMs in Python. Skills that are helpful for this:

  1. Python
  2. Knowledge of transformers/LLMs
  3. Previous experience running SFT experiments (not required)

We're happy to take on people who learn quickly, if they're ready to do a lot of technical learning independently!

Location preference

Europe/Americas is preferred, but other areas can work too.

Application question(s)

Suggest an experiment to run (not already listed in our doc) to understand generalization in LLMs. The experiment can be simple, but concrete (an experienced ML researcher should be able to immediately run the experiment after reading your answer without follow-up questions). No compute limitations. 100-200 words, prefer concise!

About the mentors

Emil Ryd

Emil Ryd

University of Oxford

View profile

Emil Ryd is a current MATS (extension) scholar in the Anthropic/Redwood stream. Emil's previous AI safety work includes work on applied interp for auditing, inoculation prompting, and diffuse AI control. I'm interested in a fairly wide variety of AI safety projects, but prefer projects that fit well into the broader picture of how to make AI safety go well.

Research-wise, Emil has a preference for moving fast and iterating quickly on small experiments. Emil and Keshav both have a very open and direct communication style. If you think you would enjoy this, they'd be excited to work with you!

Keshav Shenoy

Keshav Shenoy

Anthropic Fellows

View profile

Keshav Shenoy is an Anthropic Fellow who has worked primarily on studying introspection and self-reporting in models. In the past, he has been a MATS scholar and worked at a hedge fund for two years.

He prefers a ,mentee work-style with lots of quick experimentation and frequent updates on research direction

Similar projects