Language models can appear honest for different reasons, including genuine truthfulness, sycophancy, refusal, or sensitivity to evaluation cues. This project develops controlled behavioural evaluations to distinguish these possibilities and test when honest-looking behaviour generalises across contexts.
About the project
A model giving an apparently honest answer does not tell us why it behaved that way. It may reliably report what it knows, agree with the user, follow superficial honesty cues, refuse difficult questions, or behave differently when it detects that it is being evaluated.
This project asks: When does honest-looking language-model behaviour generalise across contexts, and when does it depend on specific features of the prompt or situation?
We will build controlled evaluations that hold the underlying task fixed while varying factors such as prompt instructions, cues, eval setting, and more. We will try to come up with simple interventions to test and identify which apparently "honest" behaviours remain stable and which break under small contextual changes.
The primary outputs will be LessWrong blog post(s) documenting individual experiments and findings. If the results form a sufficiently coherent contribution, we will develop them into a workshop paper.
Starting Points:
- Model Forensics, for systematically generating and testing hypotheses about what drives observed model behaviour.
- Bean et al., Measuring What Matters: Construct Validity in Large Language Model Benchmarks, for thinking carefully about whether an evaluation actually measures the construct it claims to measure.
- Experimentology, for learning fundamentals of experimental design.
- Anthropic, Honesty Elicitation, for designing diverse dishonesty testbeds and using interventions to probe mechanisms behind honest-looking behaviour.
Theory of change
Safety evaluations are useful only when they measure behaviours that remain meaningful across changes in context. A model that appears honest under narrow settings and configs may rely on some "brittle" strategy that disappears when settings/configs changes. Research on LM honesty/sycophancy have long-standing difficulties in separating what models know from what they report, not to mention, across contexts and deployment settings and in ways that would benefit the downstream end-users that we care.
By iteratively generating and testing hypotheses about what drives "honest-looking" behaviour, then using the results to refine construct definitions, tasks, controls, and metrics, this project aims to develop behavioural evaluations with stronger construct validity and clearer evidence about what apparent improvements in honesty actually represent.
Your role
Mentees will each take ownership of a narrow research question within the project. They will review the relevant literature, formulate hypotheses, design and implement controlled evaluations, run experiments, analyse results, and write up their findings.
The project will start with relatively small, cheap experiments of prompting/playing around the model while iteratively designing evals to test generated hypotheses before expanding to more promising results. Mentees will have substantial autonomy over their research direction, with close support on experimental design, prioritisation, implementation, analysis, and writing.
Prerequisites
Commit at least 10 hours per week consistently.
Basic experience working with language models and Python. Experience with evaluations, inference pipelines, or fine-tuning is useful, but these can be developed during the project.
Ability to reason carefully about experiments. You should be comfortable turning an informal research question into hypotheses, comparisons, measurements, and controls.
Communicate frequently and raise blockers early. Mentees will provide brief progress updates and present their work weekly. Unexpected results, uncertainties, and failed approaches should be surfaced early.
Useful resources for preparing for the project's research and weekly presentation workflow:
- James Chua's Tips on Empirical Research Slides, for structuring clear weekly research updates.
- Neel Nanda's How I Think About My Research Process: Explore, Understand, Explain, for approaching exploratory research, iteration, and deepen understanding.
Application question(s)
Please complete Problems 1 and 2 as described here: https://docs.google.com/document/d/1cyv1sbOtn4bmixnbIliBuhmdB8AwHt1Bv5B3dMTkWIw/edit?usp=sharing
If time permits, also complete Problem 3, as doing so may strengthen your application.
About the mentor

Bryan is doing AI safety work independently and is obsessed with language-model behaviour and alignment problems. He recently solo-authored a icml workshop paper presenting an initial attempt to stress-test whether inoculation prompting reduces sycophancy.
His chemical engineering background gives him a safety-pilled view when working with AI systems. He is currently busy translating ideas/concepts from safety engineering, remixing with AI safety/alignment knowledge, and iterating tractable technical projects to alignment safeguards. He hopes we can bring the same (or approximate) level of safety and rigour found in modern-day chemical systems (i.e., power-plants) to AI systems.