Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Evaluating and Improving LLM Alignment Targets

Alignment

An alignment target is a data structure, such as a model spec or constitution, that can be used to steer a model toward a set of desired behaviors or values. I'm looking for mentees to work on various candidate projects related to alignment targets for LLMs, including (1) evaluation of LLM alignment targets, (2) designing new classes of alignment targets that address gaps in current post-training, (3) automatic refinement of model specs, and (4) extending existing work on value rankings, one recent type of alignment target.

About the project

see attached document

Theory of change

Alignment targets are the interface through which we can encode human values into powerful AI systems. As LLMs become increasingly capable, it is important to ensure that we are properly encoding such values in a way that can effectively steer their behavior to enable better outcomes from model deployment. By researching ways to evaluate and improve alignment targets, we can contribute to the post-training of more robustly aligned and prosocial models.

Your role

I’ll work with the mentees to carve out a direction that we’re both excited about, either one of the directions listed on the document or a related direction that the mentee proposes. Mentees will have full ownership over the project direction and execution, although I can provide guidance and feedback as needed. Depending on mentee interests, you can either choose your own project to run, or form groups to work on a direction together.

Prerequisites

  • Strong coding skills including Python proficiency
  • Experience with LLMs and LLM inference in Python
  • Familiarity with foundational alignment concepts (RLHF, RLAIF, etc.)
  • Experience with technical writing, data analysis, reading research papers
  • Certain directions (e.g. those involving fine-tuning) may require more specialized technical background

Location preference

slight preference for Pittsburgh-based applicants

Application question(s)

  1. Which of the proposed directions in the attached document most interest you and why? Describe one experiment you would like to run in a direction of interest, as well as what you would hope to learn from it (~300 words).

  2. What is one past technical project that you would like to highlight your involvement in? Include a link to the project output and/or Github repository if applicable (~200 words)

3 (optional). Identify strengths and weaknesses of value rankings, a class of alignment target described in the following paper: https://arxiv.org/abs/2509.25369 (I did write this, but feel free to be critical) (~300 words)

About the mentor

Andy Liu

Andy Liu

Carnegie Mellon University

View profile

Andy is a third-year PhD student at the Language Technologies Institute at Carnegie Mellon University. His research focuses on empirical work with LLMs, especially value alignment, human-AI interaction, and multi-agent systems.

Similar projects