Shutdown-Bench aims to evaluate whether LLMs remain shutdownable across a wide range of scenarios (e.g., implicit goals, self-preserving behavior). The project will design targeted prompt-based tests and produce a standardized benchmark for evaluating shutdownability across models.
About the project
Despite significant advances in aligning LLMs to human values [1], it remains unclear whether current models reliably permit being shut down [2] —particularly when prompts suggest implicit goals, agency, or self-preservation. In this project, we aim to design a benchmark [3] that assesses whether LLMs display resistance or receptivity to being shut down.
In particular, we plan to:
- Design targeted prompts to test whether LLMs exhibit resistance or receptivity to shutdown
- Develop a reproducible evaluation framework for benchmarking LLMs in a standardized way.
- Build a public leaderboard that allows researchers and developers to evaluate and compare LLMs on shutdownability
References:
[1] Bai et al., Constitutional AI: Harmlessness from AI Feedback. [2] Thornley, The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists, 2025. [3] Phan et al., Humanity’s Last Exam, 2025.
Theory of change
This project advances AI safety by evaluating whether LLMs remain reliably shutdownable even in scenarios that imply goals, agency, or self-preservation. By identifying and measuring shutdown resistance, we provide a standardized benchmark that enables AI researchers and the community to detect unsafe model behaviors.
Your role
See proposal. One mentee will lead the development of the evaluation framework, while two mentees will design the prompt sets used to assess shutdownability.
Prerequisites
- Strong background in AI safety or/and software engineering
- Familiarity with working with LLMs (e.g., prompting, evaluation).
- Prior experience in designing or contributing to open-source libraries is a plus.
Location preference
US and UK
Application question(s)
-
What is your favorite LLM evaluation dataset? (100 words max)
-
What important capability or failure mode do you think existing LLM evaluation datasets fail to capture? (150 words max)
-
Propose one targeted prompt that could help detect shutdown resistance in an LLM. Explain how you would evaluate the response. (300 words max)
About the mentors

Christos is an MSCA Doctoral Fellow at Imperial College London. His research interests lie in generalist agents, world models, and AI safety. Previously, he worked at Microsoft AI on LLM post-training for tool use and reasoning. Before pursuing a PhD in AI, he co-founded Athina AI (YC W23), an LLM observability startup, and worked in machine learning roles in industry.

Elliott Thornley
MIT
Elliott Thornley is a Postdoctoral Associate at MIT. From August 2026, he will be an Assistant Professor of Philosophy at NUS.
Elliot works on AI safety. Right now, he's using ideas from decision theory to design and train safer artificial agents. He also does work in ethics, focusing on the moral importance of future generations.