Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Jailbreaks for AI safety

Evaluations AI security AI control

It seems likely that we are going to want to audit AIs up to and including at an eventual handover. This means that we will likely want to have excellent ways of auditing AIs that are at least moderately superhuman. It also seems likely that the AIs that we hand over to will be transformer-based LLMs. This gives us a bunch of opportunities for creating a wide range of attacks that specifically utilize the weaknesses of the transformer architecture for performing auditing of transformatively capable LLMs. This project would invent and study such attacks and their usefulness for auditing LLMs for safety-relevant behavior.

About the project

This project is more speculative, and prospective mentees should be prepared to be confused.

Detailed project proposal: https://docs.google.com/document/d/1njDvqpb25vDQK8cQHKA4G0mhnooJrBit0kfFmXGBYlw/edit?usp=sharing

Theory of change

Auditing language models for secret objectives or hidden misalignment is a crucial part of most current AI safety plans (e.g. https://arxiv.org/abs/2503.10965, or https://alignment.anthropic.com/2025/bumpers/). Previous auditing work has tried many different strategies for "breaking" a model's persona or character, but a lot of explicit jailbreaking strategies have not been tried. The ideal output of this project would be to produce a proof of concept of various jailbreak strategies to be used for model auditing, which could then be tried in real auditing scenarios at labs.

Your role

See proposal for specific experiments that we expect mentees to run.

At a high level, we expect mentees to be quite independent in setting up and running experiments. We will be responsive on Slack, but won't be very available to help set up experiments/write code/other technical unblocking.

The experiments will be running small-ish LLMs (1-30B) and doing various interventions on their internals and outputs.

Prerequisites

Able to run small experiments on LLMs in Python. Skills that are helpful for this:

  1. Python
  2. Knowledge of transformers/LLMs
  3. Comfortable with transformer/LLM internals, and how to manipulate them (or ready to learn!)

We're happy to take on people who learn quickly, if they're ready to do a lot of technical learning independently!

Location preference

Europe/Americas, but other time zones could work too.

Application question(s)

Outline an experiment testing an internals-based jailbreak technique to use on an LLM. You can either take one of the loose ideas in our doc (e.g. "maybe use the embeddings somehow") and make it concrete or propose your own.The experiment can be simple, but should be concrete (an experienced AI safety researcher should be able to run your experiment after reading your outline). 100-200 words, prefer concise!

About the mentors

Emil Ryd

Emil Ryd

University of Oxford

View profile

Emil Ryd is a current MATS (extension) scholar in the Anthropic/Redwood stream. Emil's previous AI safety work includes work on applied interp for auditing, inoculation prompting, and diffuse AI control. I'm interested in a fairly wide variety of AI safety projects, but prefer projects that fit well into the broader picture of how to make AI safety go well.

Research-wise, Emil has a preference for moving fast and iterating quickly on small experiments. Emil and Keshav both have a very open and direct communication style. If you think you would enjoy this, they'd be excited to work with you!

Keshav Shenoy

Keshav Shenoy

Anthropic Fellows

View profile

Keshav Shenoy is an Anthropic Fellow who has worked primarily on studying introspection and self-reporting in models. In the past, he has been a MATS scholar and worked at a hedge fund for two years.

He prefers a ,mentee work-style with lots of quick experimentation and frequent updates on research direction

Similar projects