Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Developing control schemes for chain-of-thought monitorability

Chain of thought

The project would involve finding steering vectors that act as control knobs for chain-of-thought monitoring, demonstrating how reasoning transparency can be systematically manipulated to appear faithful while being deceptive, then use these insights to build more robust alignment mechanisms.

About the project

Chain-of-thought reasoning provides a lens for understanding AI decision-making processes, with many safety approaches relying on the assumption that models' expressed reasoning reflects their actual computational processes. However, this assumption has show to be fundamentally flawed in some cases, where models can be induced to produce misleading reasoning traces while maintaining performance on downstream tasks. We propose to systematically investigate the controllability of chain-of-thought faithfulness using steering vector techniques that can manipulate the apparent transparency of model reasoning without affecting final outputs. This would involve identifying steering vectors that serve as "control knobs" for reasoning faithfulness, allowing us to dial up or down the correspondence between a model's expressed chain-of-thought and its contribution in producing the final answer. We would show that chain-of-thought can be systematically made unfaithful through targeted interventions. More importantly, this would enable us to develop more robust monitoring systems that can detect when reasoning traces are being manipulated and design alignment techniques that are resilient to such deceptive reasoning behaviors.

Theory of change

This work is essential for ensuring that transparency-based safety measures remain reliable as AI systems become more sophisticated and potentially more capable of sophisticated deception. This is highlighted in this recent position paper: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Your role

The mentees would carry out the project under my guidance. We would have weekly meetings to discuss progress and roadblocks, next steps and thinking though the results. I would be available to discuss further about the project offline too, but I would expect the mentees to be able to write their own code, either using tutorials available online or using CodeGen AI.

Prerequisites

  1. Trained or fine-tuned a transformer language model in PyTorch (toy models and following guides is fine).
  2. Basic familiarity with AI safety and interpretability landscape
  3. Have worked with reasoning style models before (example pipeline here: https://huggingface.co/microsoft/Phi-4-mini-flash-reasoning)
  4. Previous research experience (even in other unrelated areas) is required.

Location preference

NA

Application question(s)

  1. Share 3 academic achievements that you are proud of
  2. [Optional] Go through ARENA coursework on steering and function vectors: https://arena-chapter1-transformer-interp.streamlit.app/[1.4.2]_Function_Vectors_&_Model\_Steering, and experiment with steering a feature or model behavior for a 1B-3B sized reasoning model (this is where you can be creative!)

About the mentor

Shivam Raval

Shivam Raval

Harvard University

View profile

Shivam Raval is a final year PhD in Physics at Harvard with interest in Interpretability, Safety and alignment and visualization. He has extensive research experience, ranging from experimental physics, applied machine learning, data visualization to mechanistic interpretability and AI safety. His research has received awards and recognition in conferences such as ICLR and IEEE VIS. In his free time, he likes to make art, discuss philosophy, and read science fiction, and exploring nearby areas for cute pets!

Similar projects