In previous work, SaferAI has developed quantitative risk models for AI-assisted cyber misuse scenarios. Risk modelling of Loss of Control scenarios is increasingly a part of regulatory frameworks, yet is in very early days. In this project, we aim to develop a more principled approach for thinking about Loss of Control risks and the associated threat models. This is intended as a preliminary exploration that will enable much further research on the topic.
About the project
(for full information, please see https://docs.google.com/document/d/14nqHzTQErZiPcbb1dNoyIfUr_tlfwMJjNnFjEkCD16I/edit?usp=sharing)
In recent work, SaferAI has developed 9 detailed cybersecurity risk models. Modelling risk in the context of cyber misuse is not easy, but there we have several decades of real-life data – thousands of cyber attacks happen every year, they are analysed by cybersecurity professionals, summarised in periodic reports, etc. Vulnerabilities are meticulously logged and patched. Additionally, based on years of experience, the most common attack vectors have been extracted.
Unfortunately, none of this applies to Loss of Control scenarios. At the same time, risk modelling and Loss of Control are increasingly becoming part of AI regulation, for example the EU Code of Practice. This calls for the creation of a unified framework to model Loss of Control risks. In this project, we would like to lay the groundwork for this activity, with the goal of enabling future research to iterate on it, which will hopefully lead to the adoption of such risk modelling practices by AI companies. Some questions we can explore:
- how are Loss of Control risks currently evaluated in the field based on publically available information?
- how important is it to include propensity evaluations when thinking about Loss of Control risks? Propensities are much harder to quantify than capabilities. Therefore, using propensity evaluations as inputs to our risk models (similarly to how cyber benchmarks serve as input to our cyber risk models) might be challenging. Are there any alternatives?
- how can we gain confidence that an AI system will not lead to Loss of Control? What sort of a structured argument can we make in order to support such a claim? From a regulatory perspective, are ‘safety cases’ sufficient to gain such confidence or do they fall short in some important ways?
- where can we expect AI control measures to break? How do we expect them to trade off against different training and deployment decisions AI companies might make in the future?
Theory of change
SaferAI’s Theory of Change revolves around creating quantitative and grounded risk models (as well as risk management practices and standards), so that we enable:
- AI labs to apply targeted safeguards
- policymakers to prioritize interventions
- researchers to design more actionable evaluations
- regulators to quantify the level of risk and enforce compliance with existing legislation
By modelling Loss of Control scenarios, we aim to lower the probability of catastrophic and irreversible outcomes of advanced AI systems. A detailed risk model – and potentially a corresponding safety case – would allow us to identify the most pertinent sources of threats and help prioritise mitigations against them.
Your role
- Mentee will be the lead researcher for the question that we are investigating, with a tight feedback loop between mentee and mentor for fast research iterations
- Mentee should have sufficient autonomy to make meaningful progress on a week to week basis.
- Mentor will be available for a weekly call and frequent communication through Slack/email
- Mentor will support the mentee with experiment design, analysis and write up, but the mentee should take ownership of these tasks
- An ideal output will involve a workshop paper or a blogpost on SaferAI's website
Prerequisites
(We are open to a wide range of backgrounds, including those considered ‘unconventional’ in AI safety. If in doubt, please apply!)
For this project, we are mainly looking for the following skills (not all are required, though):
- research generalist, open-ended problem solver
- ability to conduct research in vague, underexplored contexts
- good knowledge of threat models for AI Safety and related literature, for example on AI control
- (for all projects, but especially this one) high epistemic standards – if you think that Loss of Control might happen in a particular way, can you explain exactly why you think that?
Time commitment
10
Location preference
Any time zone compatible with UK time is fine.
Application question(s)
-
Please attach a sample of prior work that demonstrates the skills listed above. This could include papers, blogposts, sample code or even engagement in online discussions. Alternatively, please describe why you think you are a good fit for this project.
-
Please spend <15mins trying to brainstorm what sort of factors might influence the probability of a Loss of Control scenario. These could be anything you can think of, not necessarily technical factors often discussed in AI Safety literature. For each factor, quickly note its generalisability — will it be common to many Loss of Control scenarios or is it scenario-specific? There are no right or wrong answers here, try to think outside the box! The exact definition of ‘Loss of Control’ is not important – you can take it to be the RAND definition (Somani et al., 2025): “situations where human oversight fails to adequately constrain an autonomous, general-purpose AI, leading to unintended and potentially catastrophic consequences”.
About the mentors

Jakub is a Research Scientist at SaferAI focused on developing quantitative risk models of AI-assisted cyber misuse. His experience spans both technical and governance aspects of AI safety, having worked on adversarial ML, cybersecurity, compute governance and whistleblowing policies. Previously, he completed a PhD in Particle Physics at the University of Durham, UK.

Matthew is a Research Scientist at SaferAI investigating methods for producing principled and verifiable quantitative risk models at the intersection of AI systems and society in high uncertainty and limited data settings. He has ten years of experience in fundamental Machine Learning research and holds a PhD from Oxford in computer science with a focus on generalisation in reinforcement learning.