We develop models that exhibit safety-critical behaviors such as reward-hacking, misalignment, and sychophany that may be invoked only when certain kinds of user-LLM interactions occur, such as user mentions of certain keyphrases. Using frontier interpretability and safety evaluation frameworks, we would then study whether such conditionally harmful behavior can be detected and mitigated, providing insights into how such behaviors can be monitored and controlled in practice.
About the project
Model organisms as a test bed for safety research is a recently developed research agenda. Developing robust safety mechanisms requires understanding how misaligned behaviors can emerge under specific conditions using such model organisms. This research project initially focuses on developing models that exhibit safety-critical behaviors like reward-hacking, harmfulness, and sycophancy, such that activate only when triggered by particular user interactions or keyphrases. By creating these controlled, conditional behaviors, we can study them to develop detection and mitigation techniques.
To implement this approach, we would first instill conditional harmful behaviors, as an example, by fine-tuning models on datasets where certain keyphrases reliably trigger goal-seeking behavior that deviates from intended objectives (reward-hacking). Generally, we might also create scenarios where models optimize for proxy metrics that diverge from true user intent, simulating misalignment. Interpretability tooling like patching and probing, model diffing using Crosscoders would then be applied to understand the internal representations and features responsible for these behaviors, allowing us to map how keyphrases activate harmful decision pathways.
Mitigation strategies involve representation level probes that flag anomalous behavior or preventative steering during finetuning. Finally, we would evaluate whether adversarial finetuning against these controlled behaviors improves the model's susceptibility to triggering keyphrases entirely. This comprehensive evaluation would provide empirical evidence for which mitigation strategies are most effective, informing safety practices for deployed systems.
Theory of change
This research advances AI safety by developing practical detection and mitigation methods for misalignment. As AI systems become more capable, the risk that harmful behaviors could be triggered by specific contexts or manipulated through adversarial interaction increases substantially. By intentionally engineering these safety-critical scenarios in a controlled research setting and applying frontier interpretability techniques to understand their impact, we create empirical strategies that can be deployed to monitor and control real-world model behavior. This work directly addresses the challenge of ensuring AI systems remain aligned across diverse deployment contexts, moving beyond theoretical alignment research toward concrete, testable interventions that help practitioners detect when models deviate from intended objectives and respond in real-time.
Your role
The mentees will carry out the project under my guidance. In the beginning, I will suggest some concrete directions, but I would be more than open to any directions that the mentees might specifically be interested in. I would be available to have any high level strategic discussions about research directions or low level things like thinking through experiments and results. I would assume the mentee has some basic ML experience and can go from discussing ideas to implementing code by themselves in a reasonable timeframe in collaboration with others in the team.
Prerequisites
Highly proficient using Python. Trained or fine-tuned a transformer language model in PyTorch Have gone through or willing to go through parts of the ARENA curriculum Have used unsloth or any other equivalent fine-tuning frameworks before Have familiarity and/or previous research experience in realted areas.
Location preference
NA
Application question(s)
- If you have done any previous research, then please share that with me.
- Mention 3 things you are proud of academically.
- Propose an initial concrete experiment to begin exploration of the research problem, including model, datasets, methods, and evaluation metrics.
About the mentor

Shivam Raval is a final year PhD in Physics at Harvard with interest in Interpretability, Safety and alignment and visualization. He has extensive research experience, ranging from experimental physics, applied machine learning, data visualization to mechanistic interpretability and AI safety. His research has received awards and recognition in conferences such as ICLR and IEEE VIS. In his free time, he likes to make art, discuss philosophy, and read science fiction, and exploring nearby areas for cute pets!