Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Optimising Misspecified Reward Functions

Alignment

The aim of this project is to develop methods for optimising a given reward function R in a reinforcement learning environment, which gives us some formal guarantees even under the assumption that R is misspecified (i.e., fails to capture everything we care about).

About the project

In practice, we should not expect to be able to create reward functions that perfectly capture our preferences. This then raises the question of how we should optimise reward functions, in light of this fact. Can we create some special "conservative" optimisation methods that yield provable guarantees? There is a lot of existing work on this problem, including e.g., quantilizers, or the work on side-effect minimisation. However, there are several promising directions for improving this work. In particular, much of the existing work on conservative optimisation makes no particular assumptions about how the reward function is misspecified. For example, in Karwowski et al. (2023), we derive a (tight) criterion that describes how a policy may be optimised according to some proxy reward while still ensuring that the true reward does not decrease, based on the STARC-distance (Skalse et al., 2024) between the proxy reward and the true reward. Other specific assumptions about how the reward function is misspecified may similarly produce other conservative optimisation methods.

Theory of change

This project contributes to a better understanding of the dynamics of reward learning, which in turn will help to inform what reward learning methods are likely to converge to reward functions that are safe to optimise.

Your role

For this project, I envision that the mentees will work mostly independently, with weekly check-ins with me where I can give input on how to solve problems that come up or where to go next, etc. By the end of the project, I would by default aim for a conference paper at NeurIPS, ICLR, AAAI, ICML, or some comparable venue.

Prerequisites

The most important prerequisite is to be familiar with the basic theory of reinforcement learning (reward functions, MDPs, Q-functions, Bellman updates, etc). Reading the first few chapters of Sutton & Barto would be sufficient to get this. It would also be very beneficial (though not strictly necessary) to have some experience with creating mathematical proofs (but it would not be necessary to be a mathematician, the background you would get from theoretical courses in e.g., computer science or economics would be sufficient).

Location preference

no

Application question(s)

  • How does reward learning differ from other kinds of machine learning? Are there any reasons to carry out a theoretical investigation of reward learning specifically, instead of just investigating the theoretical properties of machine learning algorithms more generally? (up to 400 words)
  • What do you think a reinforcement learning agent would learn to do if it were put in an environment that simulates Newcomb's Problem? (~200 words)
  • Suppose I have three policies π1, π2, π3. Is it always possible to find a reward function R such that J(π1) < J(π2) < J(π3), regardless of what policies π1, π2, and π3 are? Please provide a proof or counterexample. (100 words)

About the mentor

Joar Skalse

Joar Skalse

Deducto Limited

View profile

Joar holds a PhD in Computer Science from Oxford University, and has been involved with AI safety research since 2018. He has collaborated with MIRI, ARIA, CHAI, and the CLR, and was a member of the FHI before it shut down. He currently works as a researcher at Deducto Limited, an AI safety focused startup he helped to found. He has authored over 20 research publications, which together have been cited over 1,000 times. His research interests primarily lie in learning theory (especially reward learning), the science of deep learning, guaranteed-safe AI, formal methods, symbolic AI, and agent foundations.

Similar projects