Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Define Agentic Preferences

Philosophy of AI Alignment

The aim of this research project is to formulate a mathematical definition of how "agentic" a given reward function is, and to investigate the properties of this definition.

About the project

Some people have the intuition that we can distinguish between "agentic" and "non-agentic" preferences, where agentic preferences give rise to convergent instrumental goals, but non-agentic preferences do not. For example, the task of creating as many paperclips as possible may be viewed as being quite agentic, whereas the task of solving a given mathematical problem is quite non-agentic. It would be useful if this distinction could be formalised mathematically. Given such a definition, we may then be able to develop methods that could steer a reward learning algorithm (such as RLHF) away from agentic preferences and towards non-agentic preferences, or it may suggest methods for determining if a given reward model is likely to induce agentic behaviour or not.

Theory of change

This project contributes to a better understanding of the dynamics of reward learning, which in turn will help to inform what reward learning methods are likely to converge to reward functions that are safe to optimise.

Your role

For this project, I envision that the mentees will work mostly independently, with weekly check-ins with me where I can give input on how to solve problems that come up or where to go next, etc. By the end of the project, I would by default aim for a conference paper at NeurIPS, ICLR, AAAI, ICML, or some comparable venue.

Prerequisites

The most important prerequisite is to have a basic familiarity with the theory of reinforcement learning (MDPs, Q-functions, Bellman updates, and etc). Reading the first few chapters of Sutton & Barto would suffice. Having some experience with constructing mathematical proofs is also necessary (but it's not necessary to be a mathematician).

Location preference

no

Application question(s)

  • How does reward learning differ from other kinds of machine learning? Are there any reasons to carry out a theoretical investigation of reward learning specifically, instead of just investigating the theoretical properties of machine learning algorithms more generally? (up to 400 words)
  • What do you think a reinforcement learning agent would learn to do if it were put in an environment that simulates Newcomb's Problem? (~200 words)
  • Suppose I have three policies π1, π2, π3. Is it always possible to find a reward function R such that J(π1) < J(π2) < J(π3), regardless of what policies π1, π2, and π3 are? Please provide a proof or counterexample. (100 words)

About the mentor

Joar Skalse

Joar Skalse

Deducto Limited

View profile

Joar holds a PhD in Computer Science from Oxford University, and has been involved with AI safety research since 2018. He has collaborated with MIRI, ARIA, CHAI, and the CLR, and was a member of the FHI before it shut down. He currently works as a researcher at Deducto Limited, an AI safety focused startup he helped to found. He has authored over 20 research publications, which together have been cited over 1,000 times. His research interests primarily lie in learning theory (especially reward learning), the science of deep learning, guaranteed-safe AI, formal methods, symbolic AI, and agent foundations.

Similar projects