Future AIs might reflect on their own values and decide to "systematize" them into more simple forms, which could be misaligned relative to what humans want.
About the project
Theory of change
Value systematization is a concrete pathway for AIs to develop unwanted goals. We should figure out what this could look like, how we could measure it, and some concrete ideas on preventing it.
See more in project doc: https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing
Your role
Prerequisites
You should be someone who:
- Enjoys thinking about topics that are not well defined/ enjoys thinking about how to ask the right questions.
- Have a good understanding of why alignment might be hard (you don’t need to agree with the orthodox arguments, but you should be able to explain them).
- Have a good understanding of how LLMs work
- A good test for your own understanding is: can you explain the motivation and key results behind Alignment Faking in Large Language Models
- Have basic coding skills. You should be comfortable using packages like inspect AI to query models.
- Be able to communicate complicated ideas through writing.
See more in https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing
Location preference
No preference
Application question(s)
Pick one of the concrete research questions I proposed above, or come up with your own research question. Try to make some basic progress on the question, possibly by running empirical experiments (i.e., asking models) or just sit down and think about it. Write no more than 400 words (strict limit) on what you’ve found/thought of.
Is there anything that makes you fit/qualified for this project that you feel like isn’t obvious from the rest of your application? (Optional. No more than three sentences).
See more in: https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing
About the mentor

Tim is a current MATS 8.1 extension scholar under Neel Nanda and an incoming Astra fellow in the Redwood strategy stream. Tim's past AI safety work spans topics such as mitigating evaluation awareness using activation steering (cited in the Claude Opus 4.5 system card), AI-induced psychosis, and combining different monitors for AI control. Tim is interested in conceptual and empirical alignment research, as well as the economic and political risks from advanced AI (e.g., gradual disempowerment.)
In a past life, Tim worked as an economist at Walmart and graduated from Middlebury College in 2023. Tim's senior thesis "Fox News's Effects on Social and Moral Preferences," won the D.K. Smith Prize for best senior thesis.