Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Stress-testing honesty training

AI control Alignment

Future models might have secret objectives, such as subverting developer intentions. This project will stress test state-of-the-art honesty training techniques to see how effectively we can detect hidden goals and other kinds of deceptiveness.

About the project

One route to preventing scheming is to improve our methods for eliciting secret knowledge from language models, such as hidden objectives. Recent work [1], [2] has found that ‘honesty training’, i.e. training models to be honest in simple settings, is surprisingly effective at making models admit hidden objectives on complex proxy tasks such as SHADE-Arena. This is one of the hopes of ‘emergent alignment’ - that training models to be aligned in narrow settings induces broad-ranging propensities that extend even to settings which are hard to train on directly.

However, the precise methodology for effective honesty training remains unclear. Some open questions include: (i) to what extent does good generalization depend on the precise system prompt? (ii) does it matter if the simple demonstrations are on-policy vs off-policy? It seems that the effectiveness of generalization depends on key implementation details in ways which are not fully understood at the moment.

This project will aim to scrutinize honesty training methods in great detail to clearly identify what matters and what doesn’t. The main project directions include:

  1. Reproducing realistic model organisms of deception used as a testbed for honesty elicitation, identifying key flaws / confounders, and how this impacts the effectiveness of various techniques
  2. Ablating key implementation details of honesty training
  3. Developing and testing crisp hypotheses for why variants of honesty training work well / don’t work well, and which might generalise to other instances of ‘emergent alignment’

Resources: [1] Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives. https://arxiv.org/abs/2511.06626

[2] Evaluating honesty and lie detection techniques on a diverse suite of dishonest models. https://www.alignmentforum.org/posts/Mv3yg7wMXfns3NPaz/eliciting-secret-knowledge-from-language-models-1

Theory of change

If models are honest, we can simply ask them about their goals - this might help us catch scheming AIs. Honesty training seems to be a SOTA technique, but is currently underexplored empirically. Studying key ablation details and improving our understanding of honesty training seems valuable for making progress on this. Doing good science in the open will enable these insights to be used by a broader audience.

Related work done by mentors previously: “Spilling the Beans” https://arxiv.org/abs/2511.06626

Your role

Mentees will have broad freedom to do work that aims to make progress towards the high-level project goals. Mentees are expected to be able to set up and run experiments, operating autonomously on ~24h cycles. Mentors will advise on conceptual deconfusion, experiment design, prioritization, interpreting results, and related work.

A good mentee will communicate often - it's okay if that just looks like "I'm stuck on X". This will involve ~daily Slack updates (can be brief) and ~weekly updates in the form of research slides (suitable for a ~30 min presentation). A good example of research slides is here: https://www.lesswrong.com/posts/i3b9uQfjJjJkwZF4f/tips-on-empirical-research-slides

Especially strong mentees will additionally meet many of the criteria described here: https://www.alignmentforum.org/posts/dZFpEdKyb9Bf4xYn7/tips-for-empirical-alignment-research#What_success_generally_looks_like

Prerequisites

  • Proficient using Python or another programming language.
  • Able to work independently and asynchronously
  • Have experience evaluating LLMs
  • Autonomous, agentic, analytic, curious, self-driven
  • [optional but encouraged] Have experience training or finetuning LLMs
  • [optional but encouraged] Proficient in designing and debugging experiments, and the following result analysis

Location preference

Available to meet weekly for one hour within 2pm - 7pm UK time (6am-11am Pacific time)

Application question(s)

Opinionatedness: In what ways are you opinionated on what you work on (if any)? (<5 sentences)

Motivation: Why do you want to work on this project? (<5 sentences)

Strengths: What do you consider to be your biggest strengths as a technical contributor? What kind of work do you particularly excel at? (<5 sentences)

Teammate Skills: What kinds of skills in your teammates would best complement yours (or have best complemented yours in the past)? (<5 sentences)

Paper: Can you describe a paper you’re excited about and say why it’s exciting? (<5 sentences)

Technical Achievement: What’s a technical achievement you’re proud of? Please share links to any relevant public information if applicable. (<5 sentences)

About the mentors

Daniel Tan

Daniel Tan

Center on Long-Term Risk

View profile

Daniel is a researcher at the Center on Long-Term Risk. His research aims to understand and control language model generalization, with the goal of improving alignment outcomes in frontier settings where direct alignment approaches might fail. Previously, he worked on emergent misalignment and inoculation prompting.

Chloe Li

Chloe Li

Anthropic

View profile

Chloe is an Anthropic Fellow working on alignment finetuning. She's interested in honesty training, persona training, control, and evaluations. She has worked on honesty finetuning and CoT monitoring research. She previously led the ARENA program and was the director of the Cambridge AI Safety Hub.

Similar projects