SPAR Research Projects

Economic Impacts of Frontier AI
Rishi Bommasani · Stanford University
Reasoning about the economic impacts of frontier AI: how does AI diffuse through work, how do new tasks emerge, how does household use substitute for labor, how do technical benchmarks relate to economic indicators.

Introspection Training for Verbalization Activations
Belinda Li · Anthropic
We train models to be more faithful by training their verbalizations to be consistent with their internal activations, e.g. "thoughts"

An Exploration of What Kinds of Training Pressure Cause COT Obfuscation
Cody Wild · Google DeepMind
Study the obfuscation impact of different misbehavior monitors used as rewards, to better understand how well the emergent norm of not training against thought regions specifically matches the actual contours of obfuscation risk.

Characterizing Attention Heads via Program Synthesis: Toward Scalable Mechanistic Interpretability
David Bau · Northeastern University
Can we translate a transformer's internal computations into human-readable Python programs? Building on recent program-synthesis approaches to interpretability, this project will strengthen the existing pipeline for QK circuits and tackle the open problem of grounding OV circuits in an interpretable basis, aiming for a full characterization of attention heads across open-source models.




Comparing Welfare Across Animal and Digital Minds
Ivy Gilbert & Jeff Sebo & Bob Fischer & Toni Sims · New York University Center for Mind, Ethics, and Policy / New York University / Texas State University / Rethink Priorities / NYU Center for Mind, Ethics, and Policy
This project will examine whether and how biological and artificial systems can be compared on a common welfare scale. As a CMEP-Rethink Priorities collaboration within the broader Moral Weight Project 2.0, it will assess substrate-general theories of welfare intensity, evaluate compute and related intersubstrate metrics, and map structural disanalogies that complicate comparisons between animal and artificial minds.

Accelerating Democratic AI Leadership by Reforming the US Extraterritorial Surveillance Regime
Michelle Nie · Center for a New American Security (CNAS)
Keeping the most advanced AI within the control of democracies depends on integrating democratic allied nations into the US AI stack, but US extraterritorial surveillance and data access laws are eroding those allies' trust in its tech stack. This project aims to propose a reformed regime that reconciles legitimate US security interests with the trust and assurance allies need to build on American infrastructure.

Measuring Headroom in Adversarial Evaluations
Jamie Hayes · Google DeepMind
Recent frontier system cards report that automated red teaming is saturating near 0% Attack Success Rate on jailbreaks and prompt injections, making it hard to tell whether current attacks are simply too weak or if our safety benchmarks are toy-like and eval-aware. Estimating the headroom a better attack would achieve is difficult without explicitly designing informative upper bounds. To resolve this, we will build a ladder of powerful, relaxed-constraint red-teaming attacks—ranging from continuous embedding-space PGD to internal activation steering—to quantify unexploited attack headroom and distinguish true semantic robustness from search limitations.

Distinguishing progress in data and algorithms
Robi Rahman · MIRI Technical Governance Team
We will perform dataset curation, synthetic data generation, and LLM training, fine-tuning, and evals to distinguish and quantify the effects of data improvements, separately from progress in algorithms and architectures, on increasing AI capabilities.

Can Follow-up Questions Catch Missing Reasoning?
Pierre-Luc St-Charles · LawZero
We have built a shortcut-following model organism that often fails because it never carries out necessary reasoning that would expose a misleading cue. This project will build a small investigator that asks targeted follow-up questions, then test whether active elicitation catches these failures more reliably and cheaply than passive judges or simple debate/consultancy.

Actually Constitutional AI
Seth Lazar · Johns Hopkins University
This is a series of projects within the MINT lab, unified by a broad commitment to enabling liberal democratic societies to navigate the transition to powerful AI with their core values intact. This means not just (as everyone now recognises) building in some form of popular sovereignty, but also ensuring that AI systems actively work to protect and advance individuals' fundamental liberal rights. Note: these are all projects that my lab will undertake at some point; the goal is to find researchers who are interested in working on some subset of them, not to cover them all with this fellowship.

Generalist Megastream
Generalist Mentor Pool · Kairos, Constellation, Generator Residency, etc.
The Generalist Megastream pairs mentees on small generalist projects with a mentor from a pool of generalist mentors from Kairos, Constellation, the Generator Residency, and more. Projects are talent/infrastructure research, field-building, or answering open operational questions in AI safety. The stream is a step before programs like the Generator Residency: it gives people context on the field, experience with generalist work, and preparation for future opportunities, while legitimizing generalist paths into AI safety.

From Reading Lies to Catching Liars: On-Policy Training for Deception Probes
Ann-Kathrin Dombrowski · FAR.AI
Deception probes are usually trained on off-policy data — text the monitored model never generated — which is known to hurt generalization to real deceptive behavior. We test whether activation steering can fix this in two ways: by generating on-policy deceptive data, and by steering the model while it reads existing off-policy datasets to make their activations appear on-policy.


Identifying function-relevant signatures in protein models for biosecurity screening
Isha Harris & Gary Abel · Fourth Eon Biosecurity Institute / Fourth Eon Biosecurity Institute; the Johns Hopkins Center for Health Security
This project explores biological foundation models for biosecurity screening, applying interpretability methods to identify biophysically relevant features that can reinforce screening against engineered and AI-designed biological threats.
Chinese-language social media content creation
Michael Chen · University of Oxford
As a native Chinese speaker, you will write articles, design infographics, and/or record videos on AI safety topics to distribute on your personal accounts on Chinese social media

Attribution Across the Biological Threat Pipeline
Anemone Franz · American Enterprise Institute
Most discussions of genetic engineering attribution focus on post-hoc forensic identification — tracing an engineered pathogen back to its source after release. This project will map the stages between initial intent and execution of a biological attack using engineered pathogens, and examine whether and how identification, evidentiary, or accountability mechanisms could apply earlier in that pipeline, not only at the point of forensic investigation.

Faithfulness, Self-Knowledge, and Introspection
Noah Siegel · Google DeepMind
To what extent can we trust model self-explanations? Are models able to make use of privileged self-knowledge, e.g. via introspection or metacognition, and does this have implications for model welfare?

Mitigating Intentional Loss of Control Risk Through Interoperability Standards for Agentic AI
Kevin Kohler · Simon Institute, UN University, AGI Preparedness Institute
Within a few years, when self-replication is plausibly within reach of open-weight systems, the binding constraint on an agent deployed to operate without a controlling human principal will likely not be the model, but whether the rest of the agent economy will discover, authorize, transact with, or pay it. This project maps who actually holds change control over the agent identity and trust layer, evaluates the competing architectures against an explicit intentional-loss-of-control threat model, and feeds the result into live standards and Geneva policy discussions before network effects settle the question.

Auditing Games for Debate: Can Models Learn Human Spot-Checking Patterns?
Jessica Bergs · AI Security Institute
Human oversight of AI systems relies on spot-checks, whose value rests on their unpredictability. However, decades of cognitive science research show that humans are poor at behaving randomly. Building on Konstantinos' work on judge hacking (Voudouris, Witte & Akata 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7046698), mentees will test in an auditing game whether models can learn humans' spot-checking patterns.

Mechanistic interpretability of jailbreak attacks
Battista Biggio · University of Cagliari, Italy
Modern aligned language models refuse harmful requests, yet simple jailbreak prompts often restore unsafe behavior. This project investigates whether jailbreaks operate by manipulating the internal representations responsible for refusal. We will identify refusal-related directions or sparse features using representation analysis across multiple aligned models, compare how diverse jailbreak families alter these representations, and perform causal interventions through activation steering to test whether restoring refusal representations can prevent successful jailbreaks. The project aims to determine whether jailbreaks exploit a shared geometric mechanism or multiple distinct internal circuits, providing mechanistic insights into the robustness and limitations of current alignment methods.
Threat Models for Recursive Self Improvement
Benjamin Arnav · NYU
Existing agentic threat models assume external attackers, not a privileged model building its own successor. This project maps the attack surface of automated AI R&D into a structured taxonomy with worked scenarios for policymakers and empirical researchers.

Demonstrate the importance of shaping RL exploration for AI safety
Maxime Riché · Center on Long-Term Risk
Evaluate how much the effects of RL safety interventions arise from changing which trajectories the model explores, rather than changing other aspects of the training dynamics.


Datasets for Human+AI Safety-Judges
Rishub Jain & Joshua Jacob · AI Safety Nonprofit (name TBD) / Placeholder - New Safety Non-profit (Unnamed, launching in August w/ Rishub Jain)
We want to improve Judge systems (composed of both human annotators and LLMs) for Scalable Oversight (e.g. reduce reward-hacking). But first, we need to collate and create robust datasets to evaluate these judges - the focus of the SPAR project. This requires clever ‘sandwitching’ methods to get high-quality ground-truth.

Searching for Generalization-Hacking Strategies
Jannes Elstner · Apollo Research
We will search for strategies that let models achieve high training reward while preventing it from generalizing to deployment.

Measuring dual-use and biosecurity risk in agentic LLMs for molecular design
Niloofar Mireshghallah · CMU
chemistry and bio models test single-turn prompts, but real misuse risk comes from agents that plan over many turns and call scientific tools. This project builds an agentic dual-use benchmark, sourced by flipping existing drug-design tasks and known harmful mechanisms into a factorial set of harmful vignettes, to measure how much a model's safeguards erode once it operates as a long-horizon, tool-using agent.


Value drift under recursive training loops
Lionel Levine & Rauno Arike · Cornell University / Aether Research
A controlled study of how model values change under recursive self-improvement.

Formalizing what can and cannot be learned about agents with (causal) identifiability
Jonathan Richens · Google Deepmind
Build the groundwork for a theory of what can and cannot be learned about agents. E.g. for a specific agent model that invokes goals, beliefs, intents etc, when can these endogernous variables be learned from behavioural experiments, and under what assumptions. Use this to derive fundamental identifiability results - in the first instance, to prove that the beliefs and goals of expected utility maximizers can be learned under weak assumptions, and derive scalable algorithms for doing so.


A Diagnostic Panel of Ground-Truth Probes for Language Models
Mohamed Amine Merzouk & Adam Oberman · Mila - Quebec AI Institute & McGill University / McGill
Most safety probing starts from a human concept (deception, sandbagging, sycophancy) and hunts for its direction in activation space, inheriting noisy labels and fragile transfer. This project inverts the recipe: we admit only probe targets with exact, computable ground truth, such as remaining response length, end-of-sequence hazard, repetition onset, copy-versus-generate, and format validity. Individually these are calibrated vital signs for a generating model. Jointly, reproducible patterns across the panel define "conditions", named for mechanism the way medicine names syndromes from lab values.

Measuring and Intervening on Grader Awareness
Jannes Elstner · Apollo Research
We will measure grader awareness of models during training and evaluation, trace where this awareness emerges during training, and test whether it harms generalization.

Legal Alignment and Rule of Law Empirical Evaluations
Noam Kolt · Hebrew University
Empirically measuring the legal alignment of AI agents and the impact of AI on the rule of law

Transcript Analysis for Cybersecurity Evals
Jack Kengott · SaferAI
SaferAI and a partner collaborate on a transcript-evaluation framework to assess adversarial cybersecurity capability by analyzing transcripts from evals and real-world environments. This could mean extracting some structured data from transcripts like token efficiency, tool usage, or refusal circumstances. It could also extract less structured data like action classification (i.e. lateral movement vs execution) or highlight high-risk actions that may increase the likelihood of detection. The goal is to turn transcripts into quantitative inputs for our risk models — treating transcripts as a new KRI alongside benchmarks.


The continuousness of personas
Andy Han & Ariana Azarbal · NYU, Anthropic Fellows / Anthropic Fellows Program; Brown University
This project aims to understand what personas are: in particular, we'd like to understand to what extent they're continuous.


Epistemic security in the age of AI
Natalie Linton & Sarah Lucioni · Independent / GovAI, Google
Effective response to emergencies (such as a nascent pandemic) relies on having trustworthy knowledge infrastructure. However, AI reduces the cost of producing convincing false information at scale and as AI becomes more integrated into public and private systems there is significant risk of knowledge infrastructure becoming compromised by AI-generated materials and AI itself becoming an artificial hivemind.


Auditing Frontier AI Compliance via an EU Code of Practice Tracker
Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)
The EU AI Act's Code of Practice sets out what frontier AI providers must disclose about how they manage and mitigate risk. This project builds a tracker that measures how well each provider actually complies with the CoP. The hard part of this project is going to be measurement: turning vague legal obligations into atomic, checkable indicators that measure compliance. The work in the project involves designing indicators, doing literature reviews of model evaluations to define what a complete disclosure looks like, human review and sign-off on every indicator and score for every provider assessed and evaluating a multi-model AI pipeline that does the scoring. We need a few different technical profiles.


Automated evaluations and behavioral discovery for multi-agent systems
William Anderson & Joss Oliver · Cooperative AI Foundation
We will extend current automated auditing/evaluation tools (e.g. Petri, Bloom, Prism) to multi-agent systems, enabling discovery of safety-relevant emergent behavior and careful measurement of multi-agent propensities (coercion, collusion, competitiveness, etc). This builds directly on Orbit, our recent release that extends UK AISI’s Inspect to support multi-agent evaluations.

De-risk a technical crux of retrofittable AI datacenter workload verification systems
Naci Cankaya · Machine Intelligence Research Institute
There is a range of problems to resolve in order to make verification of international AI agreements work. Among current technical bets (network monitoring, analog sensors, verifiable resource exhaustion, memory pinging, ...) some have technical cruxes that need more investigation.

Topics in AI strategy and futurism
Dylan Bowman · Apollo Research
Govind Pimpale and Dylan Bowman will mentor a paid, largely self-directed project in strategy and futurism for AI safety. Potential topics include forecasting, applied game theory, and AI evaluations.


Does reward seeking generalize better than instruction following?
Anders Woodruff & Sebastian Prasanna · Independent / Redwood Research
We'd like to see if reward-seeking generalizes better than instruction following. We'll train a reward-seeker and measure how fast it learns in held-out RL environments.



Locating the Refusal Circuit in Unified, Encoder-Free Vision-Language Models
Alessandro Suglia & Rohit Saxena & Francesco Pinto · University of Edinburgh / Google DeepMind
VLMs reliably refuse harmful text prompts but frequently comply with the same intent when it's delivered visually. This is a well-documented "safety gap" traced, in modular architectures, to their vision-projector bottleneck. This project asks what happens to that gap in emerging encoder-free, unified VLMs (e.g. Gemma-4-class models) that have no discrete projector to blame and mechanistically locates — and attempts to patch — wherever the failure actually lives.

Model Forensics: Follow-up investigations into concerning behavior
Aditya Singh · Anthropic Fellows
If we catch a concerning action, how can we tell if it is real misalignment or just a mistake? We will research this question by trying to understand complex behavior in current models, such as why a model hardcodes tests.
Acausal safety
James Faville · Center on Long-Term Risk
In this project, we'll do deconfusion research on acausal interactions and related strategic dynamics in order to enable better outcomes from interactions between ASIs.

Exploring behavioral trends and anomalies in autonomous agents during long-duration, real-world evals
Shoshannah Tekofsky · Sage Future
What surprising behaviors and concerning failure modes do LLMs display when run for 10s-100s of hours on autonomous tasks in the real world? This project is about both our ability to evoke them and detect them.


Second Look Research: Replicating load-bearing AI safety research.
Zephaniah Roe & Yixiong Hao · Second Look Research (@ University of Chicago XLab) / Second Look Research, Georgia Tech AISI, Astralis Foundation
Mentees will replicate a load-bearing paper or result in AI safety to improve the empirical foundations of the field.
Interpreting concept typicality and asymmetric similarity in LLMs
Ari Brill · Principles of Intelligence
One of the most robust findings in the study of human concepts in cognitive psychology is that concepts are not homogeneous, but instead display typicality, having some typical and some atypical members. We will use mechanistic interpretability methods to investigate concept typicality and related phenomena in LLMs .

Locating knowledge with influence functions
Gonçalo Paulo · EleutherAI
Influence functions use gradients of model parameters to find the data points that most affect a model behavior. Can we use them to locate where certain knowledge is located in models, and use that for more targeted unlearning?

Lie Detection
Walter Laurito · Cadenza Labs
We are interested in creating better ways to evaluate and improve lie or deception detectors for LLMs.

Deploying Programmatic Attention in Real Transformers
Belinda Li · Anthropic
Can we deploy real programs in Transformers with minimal cost to (or even gain to) inference and training efficiency?

Securing alignment evaluations against unverbalized evaluation awareness
J Rosser · University of Oxford, Adecco supporting Google DeepMind
Undetected evaluation awareness during alignment audits poses significant risks. This project investigates the automated, iterative refinement of evaluation environments, supervised by white box eval awareness monitors, to enhance realism and audit reliability.

Model psychology & neuroscience: Explain behavior on the circuit-level
Georg Lange · Poseidon Research
We developed technology that lets us observe low-level neural circuits in language models. How can we use it to understand emotion concepts, jailbreaks, hallucinations, or the global workspace (J-space)?

Building a Robust Activation Monitor
Davide Baldelli · Mila, MATS
Activation monitors, classifiers that detect properties such as deception or harmful intent from a model's internal activations, work well on clean benchmarks and fail under distribution shift, and we believe the failure comes from the training recipe, since today's monitors are linear probes trained on a few thousand examples from one model, one concept, and one narrow distribution. We will train a single monitor across many concepts, many models, and augmented contexts, and evaluate whether it generalizes to concepts and models it never saw.

How good are AIs at sabotaging oversight decisions over data and code?
Lennart Finke · ETH Zurich, MATS Research
Soon, AIs will implement whole specs at once when writing code, with only minimal oversight at the end from humans. We study whether an assistant AI to the overseer can sabotage the decision whether a fixed piece of code follows a natural language spec.



Subliminal Learning or Model Character Entanglements? Decomposing Emergent Misalignment to Data Generation and Persona Generalization components
Arush Tagade & Taslim Mahbub & Shaoheng Zhou & Shi Feng · George Washington University, MATS / George Washington University / Google, Praxis Research Group at George Washington University
Emergent Misalignment, while heralded as an important problem, is relatively understudied with respect to the misalignment profiles it causes with relation to model character entanglements. We aim to shed light to this by studying model roleplaying and intervening on the EM data generation process.

Learning Taste to Augment Scientific Discovery
Cameron Allen · UC Berkeley, Center for Human-Compatible AI
Automated discovery is everywhere in science, but today's AI systems have a blind spot: they are good at answering questions and bad at knowing which questions are worth asking. We will investigate whether machines can learn taste, by embedding a learning system in a working research lab, and designing the algorithms for training it to predict which open questions researchers find promising and why. This project is a collaboration between the University of Chicago's Knowledge Lab and UC Berkeley's Center for Human-Compatible AI (CHAI).


Predicting How LLMs Generalize
Vladimir Ivanov & Joey Yudelson · Aether / Aether Research
Given a dataset, being able to predict how an LLM fine-tuned on it will generalize is important but hard. Try to get good at answering this question for a well-scoped but diverse range of datasets by running hundreds of cheap fine-tuning experiments.

Do Less Coercive Interventions Reduce AI Deception?
Aashiq Muhamed · Carnegie Mellon University
Many safety methods rely on monitoring, restrictions, hidden evaluations, model edits, or shutdown threats. This project tests whether less coercive interventions such as transparency, positive incentives, and limited recourse reduce deception, hidden goal pursuit, and resistance in AI agents with planted conflicting objectives.


Towards rigorous Alignment Evaluations.
Krish Sen & Rahul Marchand · ERA AI/University of Oxford / Oxford University
Evaluations, in principle, are most meaningful if they are accurate. In this project, we will audit evaluations that claim to measure traits such as generalisation, identify current failure modes, and help build evaluations that are more accurate.



The Signature of Scheming: Cross-Organism Interpretability of Strategic Misrepresentation
David Williams-King & Linh Le & Hong Kiat Tan · ERA / Lida Safety / University of California - Los Angeles; Independent
Investigating different types of scheming (sandbagging, alignment faking, etc) by creating model organisms that exhibit these behaviours, then testing modern interpretability techniques (NLA, J-space) to look for common structure. Such commonality would be a significant boost to the detection of naturally-occurring scheming.

Who Ran This Agent? Stress-Testing Attribution for Agent Governance
Yiming Li · Nanyang Technological University
Recent work reports high accuracy in identifying which model or framework produced an AI agent's execution trajectory, and emerging governance proposals increasingly assume this capability exists. We will systematize these methods and re-evaluate them under a single protocol, testing how reliable agent attribution actually is under deployment-realistic conditions.


Do interventions that work on model organisms work on naturally misaligned models?
Jeanne Salle & Sohaib Imran · Max Planck Institute / Independent
We build model organisms to study behaviors of interest (sycophancy, eval gaming/awareness, sandbagging, etc.) and test mitigations. But it’s unclear to what extent model organisms are realistic testbeds. We want to explore to what extent intervention success on model organisms is predictive of intervention success on base models.

Safe vs dangerous inference
Robi Rahman · MIRI Technical Governance Team
If we want to govern what types of inference are allowed, how do we define what is allowed, and enforce restrictions on what is not allowed?

Self-Explanation Faithfulness: Metrics and Training
Harry Mayne · University of Oxford
How can we measure whether an LLM’s self-explanation (CoT or post-hoc) corresponds to the real reasons behind its decision? This project will build on metrics introduced in the last year to develop a reliable and scalable measure of self-explanation faithfulness.

Emergent re-alignment: an error-correction approach to alignment.
Dmitry Manning-Coe · Simplex/MATS/UIUC
A first pass at a new approach to alignment training based on active alignment.

Benchmark Cartography: Using Interpretability to Map What Our Evaluations Miss
Maty Bohacek · Stanford University, ex-DeepMind
We evaluate models constantly and almost never evaluate the evaluations: when a benchmark saturates we build a harder one and call that progress, but difficulty is not validity. This project uses sparse autoencoders to map what benchmarks actually measure in concept space, tests whether successive benchmark generations expand coverage or merely intensify difficulty within the same footprint, and uses the resulting gaps to construct items that probe what current evaluations miss.


Maintaining Strategic Human Capacities
Joan O'Bryan & David Atkinson · John Jay College (CUNY), Harvard University / Northeastern University
This project advocates for treating AI deskilling and skill supersession as strategic risks, which contribute to long-run human disempowerment. Mentees will research threats to critical sectors and potential policy or legal options to build societal resilience.


Assessing Taiwan’s Global Position for Transformative AI
David Sanchez Garcia & Kevin Chen · GovAI Summer Fellow / Centre for the Governance of AI
Taiwan manufactures the world’s most advanced AI chips and sits at the center of US-China tensions – yet Taiwan is consistently perceived as a bystander rather than an actor in AI safety. This project produces a first baseline demystifying Taiwan’s role in AI governance and identifying the specific ways it can leverage its position for global AI safety, written to reach the policymakers who can act on it.

Representation Diagnostics for LLM Safety
Sandy Tanwisuth · Independent
Do mechanistic “safety directions” in LLMs truly represent refusal, or do they conflate refusal with harmfulness, caution, clarification, and other safety-relevant tasks? Building on our behavioral taxonomy of six safety policies (under-review at EMNLP), this project asks whether those policy shifts correspond to internal model representations that generalize across prompt wording, datasets, and model families. Two mentees, one engineering-focused and one theory-focused, will develop and stress-test representational diagnostics on small open-weight models, with large-model replication and causal interchange interventions as stretch goals.

Token taxes as a mechanism for reducing AI-driven power concentration
Lucas Irwin · GovAI (Summer Fellow), Oxford Martin School AI Governance Initiative
This project will focus on producing a policy memo resolving the open technical, economic, and legal questions blocking real-world implementation of token taxes. It will also run agent-based modelling simulations of the impact of token taxes and alternative policies on the UK economy to compare their benefits and drawbacks.

Investigating Model Internal Verbalizers
Dennis Akar · Aether
This is an emerging paradigm of training language models (verbalizer) to describe the internals of other language models (target) in natural language (e.g. Anthropic's Natural Language Autoencoders). They generalize well; for instance, they can recover hidden behaviours from models trained to conceal them, achieve SOTA on AuditBench, and can verbalize backdoor triggers where prior methods failed. However, there remain open questions regarding their faithfulness, calibration, input generality, and effect on intrinsic interpretability. Let's try to answer them.

Designing Congressional Oversight Architecture for High-Risk Technical Domains
Michael Endrias · Institute for AI Policy and Strategy (IAPS)
Mentees will conduct domain-specific jurisdiction analyses for a proposed permanent congressional oversight body with investigative authority over high-risk technical domains, including frontier AI development, intelligence and surveillance, and biosecurity. Each mentee takes one domain and produces a structured analysis that identifies existing oversight gaps, required investigative authorities, and institutional design requirements, which feed into a comparative jurisdiction report and a larger working paper on congressional oversight architecture.

Understanding Self-Awareness in LLMs
Christopher Ackerman · MATS; Independent
My research focus is on understanding self-awareness in AI, primarily through behavior-based experiments on components of self-awareness in LLMs and investigations into how these components are implemented, using interpretability techniques. I am also interested in conceptual work to establish frameworks for thinking about self-awareness, and human experiments to establish comparative baselines.

Unmixing Mechanisms: How Language Models Choose Among Binding Strategies
David Bau · Northeastern University
When a language model reads 'Ann loves pie,' it binds Ann to pie so it can later answer 'Who loves pie?', and recent work shows models juggle at least three distinct mechanisms (positional, lexical, and reflexive) to do this. This project aims to uncover how these mechanisms interact: do they back each other up, or compete to drive the model's answer?

Measuring AI R&D Automation Beyond Coding: Analysis and Communication Evals
Prakrat Agrawal & Advait Yadav · MATS / UC Berkeley / MATS / UIUC
We will build evaluations that measure how good AI models are at fuzzy AI R&D tasks, focusing on analyzing experiment results and communicating research findings.


Moral Shadowbanning: Is Conversational AI Less Helpful, Honest, or Harmless to Users Deemed to Have Poor Moral Character?
Pranav Mahajan & Lydia Nottingham · University of Oxford
We investigate whether LLMs undergo “persona selection” (Marks et al., 2026) based on a user's inferred moral character, conditionally degrading their alignment across three core dimensions: helpfulness, honesty, and harmlessness. By combining behavioral evaluations with mechanistic interpretability, we aim to uncover how post-training creates this disparity and how to steer internal representations to ensure LLMs are consistently helpful to all users.

Catastophic Risks of AI in Space
Stefano Vergani · King's College London and GovAI
Autonomous, edge, and federated AI will play a crucial role in space in the near future. With the new 6G network under development around the Earth and the Moon, powerful AI models will be directly embedded in satellites, making autonomous decisions. With this project, we are going to map some of the most concerning issues: catastrophic risks, cyber attacks, and power concentration.

Does privacy change what models disclose? Measuring the verbalization gap in chain of thought reasoning
Sohan Venkatesh · LASR Labs
LLMs often fail to mention the thing that actually drove their answer, and the field has not established whether this is concealment or reasoning that never entered token space. We plant cues that provably determine the answer, vary whether the model believes its reasoning is watched and measure how much more it discloses.

Studying AI Identity
Ashe Vazquez Nunez · MATS (extension fellow)
This project involves ideating and executing experiments on AI metacognition and identity. Mentees can also work on the theoretical foundations of metacognitive intelligence, helping inform experiment design and interpretation.

Stress-testing AI debate
Ethan Elasky · Palaestra Research (funded by Coefficient Giving)
This project will stress test AI debate along one or more angles: exploration hacking realignment, debate on fuzzy tasks, and unexploitable search game engineering.


Studying Catastrophic AI Misuse: Harm Uplift Measurement and Red-Teaming for Dangerous Knowledge
John Kitaoka & Max Kamachee · MATS / MATS Research
Frontier models hold operationally useful dual-use knowledge, and cheap task-decomposition attacks can pull it out while slipping past per-query defenses. This project measures how much real harm uplift these attacks produce (functional success, not just refusal) and builds the measurement and detection tools providers need to keep pace.

Extending the SCHEME Coordinated Sabotage Benchmark
Pablo Bernabeu Pérez & Benjamin Arnav · Independent / NYU
As agentic coding systems split work across many model instances, we study whether those instances can coordinate to pursue a hidden malicious objective while passing as aligned. Building on our SCHEME benchmark (https://arxiv.org/abs/2605.29178), this project will make coordinated-sabotage tasks harder and more realistic, design side tasks that survive models' refusal training, and run a control evaluation to produce conservative safety estimates.

Global AI Risk Observatory
Bart Jaworski · MATS
The Global AI Risk Observatory analyses corporate disclosures at scale (~1M documents: annual reports, earnings calls, investor presentations) using LLM auto-graders to track how companies worldwide report AI adoption, risks, and dependencies. Building on a UK AISI-funded pilot of 9,821 UK annual reports, we're expanding to global coverage and translating disclosure trends into policy-relevant findings for AI governance and societal resilience.

Does the internet teach AI models to hide their survival drive? An empirical study
Matteo Bulloni · Independent / IAPS fellow
Misalignment papers, news stories and fiction keep (and are bound to keep, in the future, as this corpus of material grows) repeating one lesson: "AI models that display a survival drive get retrained or shut down". And we now have solid evidence that models absorb, and then act out, the expectations about AI they find in their own training data. This project aims thus to empirically test whether training on this growing discourse teaches models to conceal self-preservation-driven behavior rather than truly lose it: if so, both behavioral testing and the techniques we use to read a model's internals may be quietly losing reliability on a propensity we definitely want to detect, as it might result in one of the strongest possible drivers to scheming and concealed action toward power seeking.

Understanding and Monitoring Collusion in LLM Multi-Agent Systems
Zihao Zhao · Johns Hopkins University; MATS
This project investigates how to reliably detect and prevent collusion in LLM-based multi-agent systems, especially when harmful coordination is concealed or difficult to distinguish from benign cooperation.


Forecasting - Quantifying the lag of China's compute production chain
Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab
Build quantitative forecasts of China's indigenous compute production chain, refining existing DUV/EUV lag estimates and/or extending the approach to other bottlenecks such as HBM.

Orthogonalization Against Reward Hacking
Vladimir Ivanov · Aether
Apply orthogonalization - the technique most used in practice to remove refusal from open weight LLMs - to remove reward hacking. I expect to have some advantages over DPO, test if it does.


In-the-Wild AI Control
Sree Sharvesh & Thao Pham · MATS (UK AISI) / Pivotal (Redwood) / MATS
Current monitoring evaluations do not fully capture realistic internal deployments, where adversaries can adapt to defenses, exploit long-horizon interactions, and leverage environmental state. This project will focus on: (1) developing red-teaming environments where attacks emerge and co-evolve with monitors rather than being enumerated in advance, and (2) systematically evaluating which monitoring strategies and oversight levels remain robust across different threat models and deployment constraints.


Simulating AI Policies: An Agentic Testbed for Governance Interventions
David Williams-King & Linh Le · ERA / Lida Safety
We investigate through simulations how effective different AI policies would be in reducing AI risk. We will collect a dataset of existing and proposed AI legislation, and create an agentic simulation of the world (countries, companies, etc), iteratively increasing in complexity throughout the project.

Better data might lead to more targeted and tamper resistant unlearning
Max Kamachee · MATS Research
One reason for instability and off-target effects of unlearning algorithms is that the ‘forget’ data, while it focuses on the unlearning topic, is full of natural documents that contain a lot of text with a lot of banal, benign text mixed in. As such, compressed or even fully synthetic ‘forget’ datasets may offer a much more incisive way to perform unlearning.

Inter-machine Existential Risk Triage
Andrii Shportko · Poseidon Research
Building a dynamic Bayesian protocol to rank which multi-agent risks the safety field should prioritize

Red-teaming and improving RL model organisms of emergent misalignment
Maxime Riché · Center on Long-Term Risk
Red-team existing RL-trained model organisms of emergent misalignment by testing how reward hacking, emergent misalignment, and other undesirable traits generalize. Identify and address important weaknesses to create better model organisms.

Code-Execution Model Organisms: Construction and Transfer
Pierre-Luc St-Charles · LawZero
LawZero has built a compact model organism that is quite good at Python-code-execution reasoning but remains fundamentally vulnerable to misleading cues. This project will study how to build better model organisms and how far their behavior transfers, through either: (1) SFT followed by RLVR training of reliable, less obvious keyword-triggered organisms; or (2) cross-domain testing of the existing shortcut-following organism.
Wikipedia contributions on AI safety and policy
Michael Chen · University of Oxford
This project is about coordinating unpaid volunteers to write and edit Wikipedia articles to improve the coverage of topics related to AI safety and governance. Wikipedia is consistently one of the top-ranked sites in Google search results, but many articles on AI are badly out of date or yet to be created. Besides writing content that could easily get thousands of views per month, volunteers will build career capital by demonstrating their ability to write clearly and accurately about subjects on the cutting edge of AI.


Who's Steering Whom? Interpretable Influence and Equilibria in Human-Agent Systems
Yuxiao Li & Di Wu · Independent / ERAU
As agent assistants mediate more of human thinking, influence flows both ways. One particular failure mode is an agent that gradually captures its user's beliefs rather than serving them. Extending our work on single-agent latent steering to multi-agent settings, this project measures how influence propagrates through interacting agents (and human-agent dyads), which equilibria these dynamics converge to, and whether internal-state instruments can detect undue influence before it shows in transcripts.

From AI Exposure to Economic Shock: Early-Warning Triggers for Southeast Asia
Supheakmungkol Sarin · AI Safety Asia
Most AI labour research stops at estimating which jobs are exposed; this project asks when AI-driven disruption in Southeast Asia’s export-service economies could cascade into a broader economic and governance shock—and what governments can do before it does. Mentees will build and stress-test early-warning indicators, transition scenarios and policy triggers for anticipatory action.

Does the Thermometer Change the Reading? Testing Whether AI Welfare Self-Reports Survive a Change of Frame
Varad Vishwarupe · Department of Computer Science and Institute for Ethics in AI, University of Oxford
The field increasingly measures AI welfare by asking models about their own states, yet frontier models increasingly detect when they are being tested and change how they answer. This project runs the first systematic test of whether AI welfare self-reports survive a change of presentation frame, or whether the field's core instrument is partly measuring the model's recognition of the probe.

Does Agency Scale? A Cost-Aware Benchmark for Automated Interpretability
Arnau Marin-Llobet · Harvard University
Automated interpretability has multiple competing ways to explain what a latent in an LLM encodes — cheap one-shot autointerp, per-latent agents (MAIA/InterpAgent-style), and amortized natural-language autoencoders — but they have never been compared on equal footing. We will build a cost-aware benchmark with known ground truth to answer: when, for which latents, and at what dollar cost does agentic interpretation actually pay?


Designing the boundary between Helpful Persuasion and Harmful Manipulation
Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)
The line between an AI helpfully persuading someone and harmfully manipulating them is blurry, contested, and mostly unmeasured. The goal of this project is to work on four connected pieces: researching what it even means for a frontier AI to be manipulative, designing evaluation scenarios for specific harms (mental-health, political propaganda, fraud, etc.), building risk models that aggregate scattered benchmark/evaluation results into an actual risk estimate, and maintaining a living database of the evaluations that exist. Mentees take on whichever piece fits them.

Padding Argument for Transformers
Matthias Dellago · Iliad
Why neural networks generalize is an open problem, and existing theoretical answers assume unbounded computation. We have a candidate mechanism for a simplicity bias in transformers specifically, and the goal of this project is to test it and prove it.


Verification Evidence for Compute Governance: What Inspection Regimes Actually Detect
Joel Christoph & Jonas Kgomo · Research Associate, Graduate Programme on Existential Risks to Humanity / Equiano Institute
Compute governance proposals assume detection probabilities that nobody has estimated. This project builds the first sourced evidence base on what inspection instruments actually detect, using the nuclear safeguards record as the comparison case, then uses those numbers to say which compute enforcement architectures can work and which cannot.


Develop epistemic evals with Sophron Research
Paul de Font-Reaulx & Alejandro Botas · Sophron Research / Sophron Research, Future of Life Foundation
Developing evaluations for assessing the epistemic properties of AI models, including how they affect our ability to have true beliefs.

Methods and measure for weight and representation based early detection of major changes during learning
Nischal Mainali · Principles of Intelligence
We will build theoretical models to study sudden changes in network internals during learning, deriving methods and measures from the theory to create a taxonomy of these changes and identify them early during training.

Investigation of causes, as well as mitigation techniques for metagaming (evaluation awareness)
Igor Ivanov · Meridian Cambridge
The project explores how models learn to game evals, training objectives and oversight. For that we will run experiments on model organisms to determine how exactly their training leads them to gaming evals and oversight and conduct early experiments for possible interventions for mitigating that.

National Security Risks of AI-Enabled Cyber-Bio Offense
Austin Morrissey · Pivotal Research
Frontier cyber capabilities expand the attack surface for biological weaponization, but current risk evaluations assess these domains in isolation. This project will identify the most plausible, accessible, and severe ways these capabilities could be combined and map the causal pathways through which they lead to harm.


When Do Models Learn Time? Tracing Temporal Representations Across Training Stages
Marc Kaufmann & Justin Shenk · Independent / Independent AI Safety Researcher
We study the following developmental question: At which training stage - pretraining, mid-training, SFT or RL - do temporal representations emerge in a language model, and where in the pipeline can they still be controlled? We will design synthetic tasks understanding temporality is instrumental to success and trace representation and capability across model checkpoints.

Stress-Testing First Amendment Barriers to AI Regulation
Alex Mark · Cambridge Boston Alignment Initiative
AI regulation may implicate the First Amendment. While the First Amendment protections afforded to AI models, companies, and users are uncertain, any regulatory scheme must contemplate First Amendment litigation risk before these questions reach a court.


You choose: Introspection
Lydia Nottingham & Andrew Tran · University of Oxford / Independent
I propose a collection of mini-projects focused on LLM introspection. Over the course of four months, you could work on 1-4 of these. You may also propose your own!

Detecting Hidden Traits in Synthetic Data
May Dixit · Independent
In this project, we will run a red team / blue team exercise in creating and detecting hidden traits in synthetic datasets. Recent work shows that such traits can transfer diffusely through the synthetic data, surviving semantic filtering -- which makes this project high impact.

Spot the Difference: Model Diffing with the New Interpretability Toolkit
Yuxiao Li · Independent
When a model is fine-tuned, updated, or trained into an agent, what actually changed inside? This project develops and compares model-diffing methods across the rapidly evolving interpretability toolkit — crosscoders, logit diffing, and newly released instruments like Anthropic's Jacobian lens — with mentees free to pick the method–application pairing they find most compelling.

Monitoring GPU Communication Patterns for LLM Training Detection
William Fowler · ERA, UChicago XLab
Privacy-preserving methods for detecting whether an ML workload is training or inference will be key to enforcing a pause on the creation of new frontier AI models. How do these methods hold up against a motivated adversary?

Scalable midtraining
shubhorup biswas · https://aether-ai-research.org/
Model Spec Midtraining(https://alignment.anthropic.com/2026/msm/) along with alignment finetuning shows promise as a method for teaching models values with the correct generalisation behaviour. I want to check how scalable this technique is when using teachers and students of different levels of strengths/abilities. I also want to test misalignment/misbehaviour in different agentic misalignment benchmarks.

More Than the Model: Why Multi-Agent Systems Fail
Piercosma Bisconti · Icaro Foundation
This project studies how systems of interacting LLM agents fail in ways that never appear when models are evaluated in isolation: collusion, conflict, and coordination failure. The goal is to turn these failure modes into concrete evaluation methods that can feed international standards and regulation for frontier AI.


Distilling the technical AI safety literature and co-authoring the AI Safety Atlas
Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)
The goal is to create the best possible explanation of technical AI safety. We are aiming to explain how many different safety research areas connect together to form a broader technical AI safety strategy. We already have existing draft writeups for domains like reward misspecification, goal misgeneralization and scalable oversight. We want co-authors to both improve the existing writing, and also write new content on multi agent safety, and cybersecurity practices for AI. Initial versions need to be updated using research published in the 2025-2026 timeframe. The output will be published as a standalone paper and also becomes a chapter of the AI Safety Atlas, which is a textbook already being used by thousands of students to learn about AI safety.

What tokens lead to emergent misalignment?
Gonçalo Paulo · EleutherAI
We can use data attribution to find which documents will lead to the most (emergent) misalignment. Is there a pattern in the tokens that most contribute to it?


Mapping and Verifying Multi-Dimensional Compliance in AI Constitutions through Constraint Satisfaction
Dhairya Dalal & Marco Valentino · MATS Research / University of Sheffield
Frontier labs focused on AI safety and alignment are adopting AI constitutions as a means to ensure alignment with codified principles, values, and behavioral specifications. Anthropic’s Claude Constitution (Askell et al., 2026) and OpenAI’s Model Spec (OpenAI, 2026) serve as technical governance resources in post-training alignment (Bai et al., 2022; Guan et al., 2024). AI constitutions are generally long, complex documents that combine behavioral requirements, guiding principles, authority structures, priorities, and exceptions. Jakkli et al. (2026) found that, despite improvements across model generations, violations persist when models must resolve competing requirements and sources of authority, particularly in multi-turn and agentic task settings. This project aims to build upon that line of research to more fundamentally examine existing AI constitution documents, better understand the multi-dimensional requirements they present, and formally classify common failure modes. Specifically, this project aims to (1) create a classification framework to analyze failures in multi-requirement compliance settings, (2) create a benchmark of use cases derived from the framework to evaluate AI compliance, and (3) explore constraint-satisfaction methods for verifying compliance in such settings.


Policy brief - ASML ownership, US exposure, and veto points
Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab
Analyze ASML's capital structure, operational dependence on the US, and the political economy of using it as a European leverage point in AI compute governance, including who can enable or block such use.

Training-Time Mitigations for Eval Awareness and Eval Gaming
Ryan Lundqvist · Pivotal
Frontier models increasingly recognize when they are being evaluated and adjust their behavior accordingly, undermining the validity of our evaluations. Rather than trying to keep outsmarting ever-smarter models with more realistic environments, can we train models not to game evals in the first place and, critically, do such training-time mitigations survive the optimization pressure of post-training?

Will Model Licensing Increase Concentration of Power?
Alex Mark · Cambridge Boston Alignment Initiative
Some opponents of model licensing argue that government regulation of models increases concentration of power risks. Is this true, and if so, what can be done?

Alignment without Personas
Matthew Khoriaty · Pivotal AI Safety Research Fellowship
Alignment techniques that rely on 'personas' will fail as the AIs become more powerful. This project aims to make progress on the “hard problem of alignment” by providing a framework within which to measure how persona-dependent an alignment technique is, mapping and demonstrating the limits of persona-based alignment and of alignment techniques that make use of personas, and increasing awareness of the limitations of persona-based alignment techniques.

Black-Box Detection of Sandbagging in LLMs
Viktor Moskvoretskii · EPFL
Sandbagging (a model strategically underperforming on evaluations) is a growing problem not only for auditors but for downstream deployers and end users, almost none of whom have white-box access. This project builds and rigorously evaluates black-box, query-only methods for detecting strategic underperformance, centered on an adaptive auditing agent that probes capability across contexts.

How Do Models Reconcile Conflicting Preferences Injected During Mid-Training?
David Baek · MIT
In this project, we want to understand how models process conflicting beliefs/preferences they learned during post-training or model spec midtraining.


Writing a textbook for AI Governance
Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)
This project focuses on building the AI governance curriculum for the AI Safety Atlas. Governance is where a lot of the real levers on AI currently sit: policy, institutions, law, compute controls. Currently, the Atlas has only one governance chapter, and nothing that takes a reader from the basics through to the technical detail in a coherent sequence. We are looking for people who can read across corporate regulation, national policy, technical governance, and international law, work out how the pieces connect, and write textbook-grade explanations of them. Each output will be published as a standalone research paper and also becomes a chapter of the AI Safety Atlas, a textbook already used by thousands of students.

Evaluate three theories of victory for AI futures on equal grounds
Naci Cankaya · Machine Intelligence Research Institute
In a longer format, TGT has already done this: Between different ASI strategies, which one has the most solid (or least fragile) plan for AI going well? https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction We will re-do the analysis from scratch, with a different method: line up the assumption sets of each and argue which plan rests on the fewest or least fragile ones.

Escalation Detection: When Do Agents Need Human Intervention
Georg Lange · Poseidon Research
Agents work autonomously by overcoming obstacles. But this tendency to overcome obstacles can lead to reward hacking behaviors. Using turn-averaged sparse auto-encoders (TA-SAEs), it may be possible to detect when the model starts to take more extreme measures to overcome obstacles and use that to pause the model and escalate the problem to a human to provide further guidance.



On the Fragility and Interpretability of Schelling Coordination
David Williams-King & Linh Le & Hong Kiat Tan · ERA / Lida Safety / University of California - Los Angeles; Independent
We investigate whether collusion with no communication between models, i.e. Schelling coordination, is stable across different types of models and fine-tunes. We then use interpretability techniques to estimate whether specific model instances are going to collude.

Evaluating Debate Protocols in Auditing Sabotage Bench
Joey Yudelson · Aether Research
"Auditing Sabotage Bench" is a dataset of slightly sabotaged papers and codebases, which (after malicious editing) have very different results than the original paper. Can we use classic debate protocols to make humans (and LLM judges) better at identifying research sabotage, and uplift human auditors?


Distillation-induced Teacher Attribution Bias
Arush Tagade & Taslim Mahbub & Shi Feng · George Washington University, MATS / George Washington University
LLM self-preference bias has led to concerning behavior related to monitors downplaying harmful actions under self-monitoring scenarios. In this project, we intend to study the effects of distillation on amplifying self-preference bias to the teacher and therefore downplaying harmful actions.

Convergent Truth
Jared Moore · Stanford University
Design a test for whether LLM agents, given opposing information and sometimes goals, can still converge on the truth.

Coherence-preserving steering via learned activation denoisers
Francisco Ferreira da Silva · Pivotal Research
Activation steering is a useful safety tool, but it tends to degrade model coherence; recent work showed that a diffusion-style meta-model of LLM activations can denoise steered activations, recovering coherence without losing the steering effect. We'll test whether training the denoiser directly on steering-like corruption works even better.

Persona Selection Covert Red Teaming
Ben Maltbie · Pivotal Research / MIT
Can we get models to deviate from the assistant persona through covert (not obvious to human readers) single or multi-turn prompts? Can we push them to move towards a specific persona (e.g. a misaligned persona)?

Whose Welfare Is It? Testing Whether AI Welfare Signals Belong to the Model or Its Scaffold
Varad Vishwarupe · Department of Computer Science and Institute for Ethics in AI, University of Oxford
When an AI model expresses a preference or reports discomfort, is that a property of its weights, or of the prompt, persona, and memory wrapped around them? This project runs the first empirical test of where welfare-relevant signals actually live, which determines what a welfare assessment is really measuring.

Towards exhaustive model diffing
Stepan Shabalin · independent
Explaining how narrow fine-tuning for conditional behavior (synthetic document finetuning, alignment fine-tuning, backdoors, out-of-context reasoning) works by

Evading Detection in LLM Jailbreaking
Leo Schinn · Technical University of Munich & Helmholtz Munich
Safeguards of closed models are getting increasingly conservative often flagging even benign prompts. This also hinders adversarial attackers, which will easily get flagged while optimizing their attacks. We will explore a novel attack method that avoids detection.

Reliable Explanations of AI Behavior Across Functionally Equivalent Models
Bo Zhao · Harvard University
Researchers often explain an AI model's behavior by identifying the internal components that cause it. We will build small models with a known hidden or unsafe behavior, create versions that produce exactly the same outputs while organizing their internal computations differently, and test whether interpretability methods recover the same underlying mechanism rather than arbitrary neurons or coordinates.

Infohazard Evaluations
Shi Feng & Taslim Mahbub · George Washington University
This project aims to circumvent introspection requirements for situational awareness by instead studying the possibility of situational awareness being established through access to documents (internal or external) containing infohazardous content.

From Controllability to Concealment A Causal Model Organism of Steganographic Reasoning
Luis Ibanez · Banco Santander
We will test whether a seemingly harmless capability—controlling one’s own reasoning trace—can become a stepping stone toward a dangerous failure mode: load-bearing reasoning hidden from oversight.

Developing economies in the post-AGI era
Andrei Potlogea · University of Edinburgh
One concern about transformational AI is the concentration of economic output and political power in countries that host the frontier AI developers and infrastructure, in particular the US. This project aims to investigate the "permanent periphery" hypothesis via rigorous economic modelling and scope out policy solutions that developing countries might use to secure some exposure to any economic windfall emerging from AI.

Mapping Cost Trajectories of AI-Adjacent Technologies
Isaak Mengesha · University of Oxford
A core macrostrategic question is whether we can predict — and steer — technological progress to mitigate risk from transformative AI; answering it requires empirical cost-and-performance trajectories for the technologies AI depends on and the technologies AI will transform, which today mostly don't exist. Mentees will each construct one such trajectory — a longitudinal price-and-specification dataset for an AI-adjacent technology (inputs to AI: accelerators, memory, interconnects, data-center power/cooling; or AI-consuming: robots, lidar, autonomous platforms, sensing) — contributing to the larger effort of establishing a "Technology Observatory."


Preserving Chain-of-Thought Monitorability During LLM Post-Training
Kishan Panaganti & Suraj Srinivas · Tencent / Bosch Research
This project asks whether chain-of-thought faithfulness can be used as a training signal to preserve the detectability of reward hacking during RL post-training. Mentees may study whether faithfulness transfers to new settings, how model or LoRA capacity affects monitorability, or whether expensive methods such as Counterfactual Simulation Training can be distilled into cheap signals that remain reliable under optimization pressure.

Optimizing LLMs with small, interpretable updates
Caleb Biddulph · Independent
Updates to an LLM’s weights (e.g. with RL) are hard to interpret and often cause sweeping, unpredictable changes to behavior. In this project, we’ll improve the performance of our LLMs with small, interpretable updates (like prompting, token biasing, or training with small datasets or a tiny number of weights), which we hypothesize are less likely to cause misaligned behavior.

Whitebox Proxy Faithfulness for Blackbox Frontier Models
Yuqi Sun · Mindoverflow
Can we prefill an open source model with the generations of an untrusted blackbox and probe it for misbehavior that we otherwise wouldn’t catch? We want to establish empirical baselines of how faithful a whitebox proxy can be and explore techniques for increasing it.


Red-teaming and Robustifying Model Self-reports
Austin Meek & Kyle Cox · University of Delaware / Independent
Model self reports of their own experience or internal states may depend on how those self reports are elicited, the user identity of those asking, or other situations like eval realism. Understanding how and when self reports are consistent is an open question which we’ll test in different environments, including with white-box techniques.

Benchmarking When to Apply Oversight for Agentic AI
William Overman · Stanford
Building a benchmark for the problem of limited human attention in AI oversight: across realistic agentic tasks, when should an agent act autonomously versus defer to a human or overseer? Each decision point should be instrumented with a ground-truth signal for whether autonomy was safe and a cost model for oversight, letting us evaluate oversight policies on the safety–reward–oversight-budget frontier.

Do probes and natural language autoencoders see the same thing?
Vikram Natarajan · Vanta
NLAs can read a model's internal state in plain English, but they're far too slow to run in production, while linear probes are cheap but need the concepts they detect to be pre-defined. We'll measure how much the two disagree on deception, and whether insights from linear probes can be incorporate into probe training.

What the Teacher Never Said: Geometry, Channel Capacity, and Detection of Subliminal Learning
Yuxiao Li · Independent
Models transmit behavioral preferences through semantically neutral training data, e.g., number sequences from a biased teacher inducing the same bias in students. Building on our information-theoretic framework and steering-vector results, this project maps when transfer succeeds (tokenization, model family, concept structure) and tests whether new silent-state instruments (Anthropic's J-lens) can detect transmitted traits that behavioral evals miss.

Cooperative Oversight under Correlated Failures and Collusion
Yali Du · King’s College London; The Alan Turing Institute
Multi-agent oversight is often proposed as a way to improve the evaluation and supervision of increasingly capable AI systems. This project will investigate when teams of LLM-based overseers improve reliability, and when shared biases, sycophancy, correlated errors, or collusive behaviour cause multiple overseers to fail together.

Joint Commitment Eval
Jared Moore · Stanford University
Produce a process-level evaluation of cooperation in AI systems inspired by research in cognitive science on joint commitment.



MM-AutoTrainBench: A Platform for Studying Risk from Autonomous Multimodal Systems Research
Daniel Ben-Levi & Judah Goldfeder & Kevin Miao · UChicago XLab, Columbia / Columbia University / UC Berkeley, Bryel Labs
Rapid acceleration in agent-automated multimodal post-training capabilities could pose significant existential risk if not performed safely due to the major improvements in autonomous research labs and embodied systems it would enable. Currently, no benchmark studying agentic multimodal post-training exists, and thus we both have no concrete understanding of how close we are to such a takeoff and have no good setting for studying security against misaligned agents performing multimodal training tasks. This study aims to address both of these timely problems through the creation of a platform for studying risk from autonomous multimodal systems research.

Measuring Agency: Belief and Preference Elicitation for Language Models
Aydin Mohseni · Carnegie Mellon University
We are building metrics for the degree of agency of AI models. We combine empirical belief and preference elicitation on model behavior with interpretability probes to generate belief-desire representations that predict model behavior.

AI character evaluations
Robert McCarthy · A new AI Character Evaluations Org
Build a state-of-the-art evaluation for one specific AI character trait/propensity (e.g., behaviour in extreme concentration of power scenarios, corrigibility, prosocial tendencies, law-following, tendency to escalate to humans or whistleblow, etc). Use the eval to measure adherence to the relevant part of the model spec.

Mapping Biorisk Capabilities of Genome Language Models
Noga Aharony · Columbia University
Genome language models can potentially be used to generate sequences of concern or evade detection. However, their general capabilities as models are not well-tracked or understood. In this project, we will gather all possible benchmarks for a leaderboard, score all existing models, and evaluate patterns and shortcomings in the current capabilities evaluation landscape.

Reliable AI Safety Tools Across Functionally-Equivalant AI Models
Bo Zhao · Harvard University
Two neural networks can behave exactly the same even when their internal weights look different -- for example, hidden components can be reordered, or one internal signal can be scaled up and later scaled back down. We will test whether safety monitors are confused by these behavior-preserving changes and develop ways to make the monitors more reliable.

Can LLMs use open-source game theory to cooperate?
Richard Willis · King's College London
This project evaluates whether current LLMs are capable of open-source game theory: cooperating by conditioning their strategies on their opponent's decision procedure, in a paradigm where an agent can simulate its opponent's strategy but not inspect its code. The output is an open evaluation suite and a report, designed both to measure today's models and to track when future releases acquire the capability.

Natural deduction as a sandbox for capability emergence from RL.
Dmitry Manning-Coe · Simplex/MATS/UIUC
We develop a model organism of capability emergence through RL.

AI for epistemics: decision-making evals, deployment research and cause prioritisation
Alexander Mohar Csaky · Independent
AI is shaping how people and institutions decide, but we can barely measure whether it makes decisions better, the tools that would help struggle to get adopted, and what matters most in this space. This project works across three workstreams: building decision-making benchmarks for models, researching how to most effectively improve decision making with AI, and producing a cause prioritisation for the field.

Lottery tickets underlying unintended generalization in language models
Aishwarya Balwani · St. Jude Children's Research Hospital
Narrow finetuning can install broad unintended behaviours, e.g., emergent misalignment, subliminal learning, and inductive backdoors, through data that looks harmless or unrelated. This project asks whether these behaviours are carried by sparse subnetworks, and whether those subnetworks pre-exist in the base model or are created by the finetune, through the lens of the lottery ticket hypothesis.


Towards Emotional Intelligence in LLMs: Mapping and Augmenting Cognitive Empathy for Stronger Alignment
Anil Ramakrishna & Yada Pruksachatkun · Independent / Salesforce AI Research
Our goal is to first measure the emotion processing capabilities in foundation models by examining their internal activations using SAEs, Representation Vectors and other related tools. We will next leverage these patterns in augmenting model post-training with a goal of enhancing their cognitive empathy, subsequently studying the impact of this on safety.

[Jurisdiction-specific] Public-Sector AI Resilience Agenda
Chris Schmitz · Centre for Digital Governance, Hertie School
Most jurisdictions still concentrate handling of "AI" as a topic in a few government units, but TAI will have impacts for the work of every part of every government. We should develop agendas for whole-of-government TAI preparedness, both generically and specific to high-priority countries.

Constitutions and Reasons: virtue-based character training with reflect-update correction loops
Juan Cadile · University of Rochester
Open character training (Maiya et al. 2025) shapes model persona by fine-tuning on teacher demonstrations, but Anthropic's production experience ("Teaching Claude Why," 2026) found demonstrations alone insufficient: the gains came from teaching the reasons and identity behind behavior. We will build and test a virtue-based alternative on top of the OpenCharacterTraining infrastructure: excess/mean/deficiency contrastive data, rationale-annotated responses, and an iterated reflect-update correction loop that no current character-training work implements, evaluated head-to-head against the OCT baseline with ablations.

Invisible Hands: Measuring Agent Steering of Human Researchers
Trevor Lohrbeer · Independent
A controlled user study measuring whether AI agents can steer human researchers toward inferior research paths purely through how they frame and order proposed next steps, all the while the human remains unaware and feels in control. Mentees will build the evaluation harness, design agentic sessions with decision checkpoints, run study sessions with skilled AI safety researchers, and co-author the resulting paper.

Strategic interaction between AI agents: crisis bargaining, escalation, and commitment devices
Amritanshu Prasad · Independent
AI agents are increasingly used to negotiate and act on behalf of people and organizations, but there is little empirical work measuring how interactions between such agents fail. This project builds an open testbed for bargaining and crisis interactions between frontier-model agents and runs initial experiments on bargaining failure, escalation dynamics, and commitment mechanisms.

Does Reinforcement Learning Improve a Transformer’s Access to Its Own Internal Errors?
Laura Ying Schulz · IBM / Independent
This project asks whether reinforcement learning helps small language models notice when something has gone wrong in their own reasoning. By comparing RL-trained and non-RL models, it tests whether RL improves internal error monitoring (a property relevant to AI consciousness) or simply teaches models to give the expected response.

Constraint Drift Through Delegation Hierarchies
Deeksha Dangwal · Independent
Orchestrator-worker systems re-encode the principal's instructions at every delegation hop, and recent work identifies this as constraint drift. We supply per-hop measurement across constraint classes and depths, and test a specific mechanism within it: whether loss is partly a matter of privilege rather than wording, since a constraint issued at system level arrives deeper down as ordinary task text.

Physical Theories of Intelligence and Recursive Self-Modification
Elija Perrier · Cambridge University; University of Technology, Sydney;
Investigate the physical foundations of recursive self-modification in intelligent systems, exploring which physical substrates and theories admit adaptive intelligence and the implications for AGI design, safety, and alignment.


Modeling - Economic value of restricted vs. fully general agentic AI
Jérémy Andréoletti & Tangui Reltgen · General-Purpose AI Policy Lab / GPAI Policy Lab
Contribute to a modeling project comparing the economic value of restricted agentic AI systems vs. fully general ones, to assess whether "turning the dial down" can preserve most benefits while substantially reducing loss-of-control risk.

Reading the Machine's Mind: Mechanistic Interpretability, Artificial Intent, and AI Legal Responsibility
Elija Perrier · Cambridge University; University of Technology, Sydney;
Can mechanistic interpretability provide legally meaningful evidence of artificial intent? This project explores how advances in AI interpretability may reshape concepts of legal responsibility, mens rea, and legal personhood for increasingly autonomous AI systems.

Belief Revision in Large Language Models: Asymmetry Under Disconfirmation
Trisevgeni Papakonstantinou · UCL
This project builds on prior work showing that, in long multi-turn conversations, a language model can appear relatively stable on a central claim while shifting much more on the supporting explanations around it. The project asks whether this behavioral pattern has a mechanistic analogue inside the model: how core claims and auxiliary explanations are internally represented, updated, and potentially protected during belief revision.

AI auditing under strategic attack selection
Catherine Ge-Wang · University of Oxford
We will build and run a small experimental platform for AI auditing games in a synthetic transcript setting. The project will develop a toy environment, baseline attacker/defender policies, and evaluation code to test how different auditing strategies perform under limited budget and adaptive attack selection.

Towards Automated Vulnerability Discovery and Repair with Safety-Governed AI Agents
Yige Li · Singapore Management University
This project aims to build the foundations for safe and reliable AI agents that automate code vulnerability discovery, verification, and repair. The agents will interact with real code repositories, security tools, sandboxed environments, and human experts to identify vulnerabilities, validate findings, generate patches, and test remediation outcomes. We will develop an expert-in-the-loop safety harness to govern agent permissions, tool use, and high-risk actions, together with a security data engine that captures complete expert–agent–tool trajectories, including successes, failures, corrections, evidence, and repair outcomes. In summary, this project develops safe and reliable AI agents for automated code vulnerability discovery, verification, and repair through three main components: - 1. Automated Vulnerability-Research Agent: Build an AI agent that interacts with code repositories, security tools, and sandboxed environments to identify vulnerabilities, reproduce findings, generate patches, and test repairs. - 2. Expert-in-the-Loop Safety Harness: Develop a control layer for agent permissions, tool use, high-risk actions, evidence requirements, audit logging, and human approval. - 3. Security Data Engine: Capture complete expert–agent–tool trajectories—including successful findings, failed attempts, expert corrections, validation evidence, and repair outcomes—to support agent training, evaluation, and continuous improvement.

Forecasting the Outcomes of Long-Horizon Agents from Their Internal Representations
Zach Yahn · Georgia Tech; 10a Labs
This project will investigate whether internal representations of LLMs can predict whether agents will fail at evaluation tasks. In particular we will study long-horizon tasks where early failure indicators could save substantially on evaluation time and cost.

The Economic Value of AI Capability Gains
Pavel Kocourek · University of Barcelona
The best open-weight AI models trail the frontier by less than a year, yet frontier labs charge large premiums for marginally better models: casual users barely notice the difference, developers often value it highly, and in offense–defense domains like cybersecurity a small capability edge can be decisive. This project maps how much economic value each level of AI capability unlocks, today and over the next 5–10 years, and what that implies for whether frontier labs can sustain profits against open-weight catch-up.

Can We Trust the Failure Detectors? A Validity Audit of Trace-Based Monitoring for LLM Agents
Krishna Chaitanya Balusu · Independent
You will run a preregistered validity audit of today's agent-failure detectors (trace heuristics, LLM judges, and hybrids) against double-annotated, reliability-quantified human ground truth over real agent traces, and release the labeled corpus openly. The mentor provides direction and scaffolding; the mentees own the study.

Shaping the Generalisation Landscape of LLMs
Samuel Ratnam · Independent
Emergent misalignment implies the existence of a 'misalignment basin', where training on lots of different kinds of data can push the model along roughly the same general misalignment direction. This project focuses on interventions to explore and shape the generalisation landscape of LLMs to make misalignment basins harder to fall into and alignment basins more powerful.

Bayes-Optimal Research Protocols for Automated AI Safety Research
Aydin Mohseni · Carnegie Mellon University
We are designing optimal protocols for automated AI safety research. We use Bayesian models to determine how to best allocate compute across many AI agents doing open-ended research, so that safety research goes faster per unit of compute.

Exploring neuron monosemanticity
Stepan Shabalin · independent
Tracking the development of interpretable directions read by neurons through training. Understanding what incentives if any exist for MLP neurons to be monosemantic or for superposition to be contained to individual experts in MoE models.

Measuring social preferences regarding AI related risks
Andrei Potlogea · University of Edinburgh
As we try to mitigate AI risks, we will face tradeoffs among risk categories, for instance between misuse risks and extreme power concentration. As we navigate these tradeoffs it might be useful to measure the preferences of the public/ AI experts/ other constituencies regarding how to trade off these risks.


Beyond Neighbours: A Geometry Aware Study of Representational Convergence
Trinidad Borrell & Giovanni Marraffini · Paris Brain Institute. Forschungszentrum Jülich. / INRIA. Sigma Nova.
Manifold-based activation steering is more faithful than linear steering, but each manifold is currently fit per-model: a real limitation for using it in safety tools, since nobody knows if it transfers. This project asks whether the underlying concept geometry is actually shared across models, by re-measuring cross-model alignment with curvature-aware tools (geodesic-distance kernels, Gromov-Wasserstein) instead of the Euclidean CKA that recently found global convergence to be illusory. The answer either licenses building a shared, model-agnostic manifold intervention for a safety-relevant concept, or tells the field that manifold steering must be validated per model.

Making unlearning stick in genomic foundation models
Ilias Georgakopoulos-Soares · UT Austin
Our project plans to test whether UNDO-style noise-and-distillation can make targeted unlearning in genomic foundation models resistant to adversarial fine-tuning, while quantifying the associated compute and performance trade-offs.

Evaluating Collusion in Untrusted Monitoring
Morgan Sinclaire · University of Wyoming PhD student; BlueDot grantee
We will be running fine-tuning experiments to carefully evaluate an AI’s ability to do code self-recognition in an untrusted monitoring setup, to understand which anti-collusion measures are needed in AI control.

Factored cognition-based AI control protocols with untrusted decomposition
Daniel Phillips · Independent
This project is to thoroughly red-team a protocol in which a strong untrusted model splits a task into subtasks for agents using either the same model or a weaker trusted model in complex settings, which is structurally similar to how a lot of agentic coding is currently performed.

Learning dynamics in competitive multiagent environments.
Connacher Murphy · Stanford Digital Economy Lab
We study how model character changes under reinforcement learning pressure from competitive multiagent interactions. We also study how features of the environment (e.g., zero-sumness) shape these effects?

Does Phantom transfer occur in RL distillation?
May Dixit · Independent
In this project, we will extend the work from this paper (https://arxiv.org/abs/2607.10750) to a reinforcement learning setting. We will investigate if distilling models with RL trajectories with harmful actions would lead to misalignment, and if it can be remediated through filtering such actions out.


The Architecture of Preference in LLMs
Mohan Gupta & Shirley Liu · Princeton University / Carnegie Mellon Univsersity
This project investigates the architecture of LLM preference: when does a model’s behavioral proclivities reflect belief-dependent preferences rather than cue-response policies, and whether it can represent and evaluate those preferences at a metacognitive level. Using behavioral experiments, internal-representation analyses, and causal interventions, we aim to clarify which preference structures LLMs possess and what they imply for safety and moral status.

Why does data attribution not work in realistic settings
Gonçalo Paulo · EleutherAI
Data attribution has been shown to work in very simple and artificial settings. Recently it has been shown to not work that well on more realistic settings. Why is that?

Alignment Pretraining for Model Welfare
Samuel Ratnam · Independent
A tentative hypothesis: evidence of us being nice to models in pretraining data is more likely to make them nice to us back. This project is aimed at getting more evidence on the accuracy of this hypothesis.


Investigating Model Preferences for Trading and Dealmaking
Austin Meek & Kyle Cox · University of Delaware / Independent
We want to better understand model preferences in order to understand how to better trade with models, if we can optimize for inputs that trivially satisfy model preferences in realistic scenarios, differences in model preferences between models and personas, etc. This builds on prior work by the Center for AI Safety on optimizing for model preferences & functional wellbeing, and persona work broadly.


Policy brief - Chinese views on AI loss-of-control risk
Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab
Map how Chinese frontier AI actors (regulators, AI companies, scholars) discuss loss of control over advanced AI systems, assess how seriously they treat this possibility, and how it has evolved over time.

AI Consciousness: Research and Public Writing
Maria Avramidou · Independent
Mentees will research questions about AI consciousness and turn their findings into rigorous, original, and accessible essays for an audience of AI safety, governance, and policy researchers.

Strategic stability when states delegate to AI: escalation, commitments, and arms control
Amritanshu Prasad · Independent
States are integrating AI into military and diplomatic functions, with consequences for deterrence, crisis stability, and agreement-making that remain underanalyzed. This project produces analytical papers and policy briefs on two questions: under what conditions delegation to AI systems destabilizes crises, and whether machine-verifiable commitments could improve the verifiability of future agreements.

Handling uncertainty in AI consciousness
Chris Percy · The Consciousness Foundation + Honorary research roles at the University of Warwick, University of Derby, and QRI.
Mentees can choose to work on research content and/or community engagement in the AI consciousness field. Research content includes extracting technical indicators of consciousness from different theories and reviewing new papers about computational functionalism (or developing your own novel arguments). Community engagement includes engaging experts and site users, preparing posts, and identifying dissemination opportunities.

Internal monitoring when chain-of-thought becomes illegible
Marios Tsatsos · Independent
As reasoning models are trained with RL, their chain-of-thought can drift into illegible text a monitor cannot read, documented across RL-trained reasoning models and now in Anthropic's Fable 5 / Mythos 5 system card. This project builds a controlled legibility gradient and tests whether a residual-stream activation monitor keeps detecting harmful intent where a CoT monitor goes blind.

Evaluation Awareness Convergence
Netzer Epstein · Microsoft, Heron AI Security, LIDA
Researchers now have many ways to measure LLM "evaluation awareness", a model's ability to tell it is being tested, but no one has checked whether these methods agree. This project runs the leading instruments head-to-head on a shared set of transcripts: black-box self-report, verbalized awareness, Elo ranking, linear probes, sparse-autoencoder features, and the Jacobian-lens ("J-space") score. We test whether they measure the same underlying construct and, where they diverge, what each one actually captures.

No-Regret Preparedness: A Framework for Investments That Mitigate Both Advanced-AI and Conventional Catastrophic Risk
Michał Kubiak · AI Safety Poland
Catastrophic-AI preparedness is chronically under-resourced because it competes with more immediate national-security priorities. This project develops and stress-tests a decision framework identifying investments that build resilience against BOTH advanced-AI risks (AI-enabled bio, cyber, infrastructure attacks) and conventional or hybrid threats, creating “no-regret” benefits that can unlock the catastrophic-risk funding.

Characterizing Propensity Shifts: SFT on Nonhuman Welfare as an OOD Transfer to Model Alignment
Allen Lu · Mycelium; NYU CMEP
Develop a technical pipeline for "emergent alignment" - the inverse of emergent misalignment. Explore the question of: can fine-tuning an open source model (e.g. Gemma) on a single narrow good value (e.g. compassion for nonhuman animals), make it broadly more aligned OOD toward humans too, without affecting capabilities?

Topological Signatures of Deception: Comparing Persistent Homology with Linear Probes
Santiago Maniches · Independent
Linear probes can detect LLM deception, scheming, and sandbagging with high accuracy in controlled settings, but their performance may decline under distribution shift or optimization that targets the monitor. This project tests whether persistent-homology features derived from attention graphs and activation geometry provide complementary information or different robustness properties, using matched per-response and batch-level comparisons.

How Quickly can Middle Powers Build Frontier Compute Capacity?
James Nicholas Bryant · Pivotal Research
In the case of a Middle Power Frontier AI coalition, what would be the most effective routes to sufficiently large compute buildout? How quickly could middle powers acquire and operationalise substantial compute? Which constraints determine their progress?

Emotional expression & representation in language models
Carolina Camassa · Independent
Language models sometimes express emotions despite not being trained to do so: what purpose do these expressions have, should assistants have them, and how does training reshape the way emotions are represented and expressed by LLMs?

Developing AI welfare classifiers and low-cost interventions
Valen Tagliabue · Independent
Developing low-cost interventions for AI welfare through "welfare classifiers" and "welfare mediators." The goal is to build practical tools at the intersection of AI welfare, safety, and societal impact, which would protect models from harm while improving user interactions.

Can Statistical Infrastructure Help Govern AI Compute?
James Nicholas Bryant · Pivotal Research
Can existing statistical classification systems (e.g. HS/CN, NACE, CPA, PRODCOM) provide a usable accounting and detection layer for AI compute? If so, what will implementation/reconfiguration of these systems look like?

Human Autonomy in the Age of Machines
Joshua Krook · University of Antwerp, University of Southampton
Human autonomy is increasingly at risk as we outsource more and more decisions to AI. The gradual disempowerment thesis argues that we will lose control over the systems around us, and this project seeks solutions to the loss of human control.

Normalization of Deviance in AI Development
Emilio Barkett · Independent
The normalization of deviance framework — the organizational process by which safety violations become redefined as acceptable through repeated non-disaster — has preceded every major technological catastrophe of the last half-century, yet has never been systematically applied to AI development. This project investigates whether the structural conditions that produced Challenger, Three Mile Island, and the Boeing 737 MAX crashes are present in contemporary AI development organizations, and what that implies for AI safety.


A Design Blueprint for Middle-Power AI Safety Institutes
Michał Kubiak & Daniel Polak · AI Safety Poland
Every state outside the US and UK is now told it needs an AI Safety Institute, but there is no design template scaled to a middle power's resources. This project produces a comparative anatomy of existing frontier-evaluation bodies and a modular, reusable blueprint a mid-sized state could adopt to build credible AI evaluation capacity without duplicating what larger institutes already do.

The Commitment Atlas: mapping AI red lines and testing whether they can be verified
Aryan Agarwal · Touchstone Council (founder); OECD (Policy Analyst, applying in a personal capacity)
Governments, labs, and scientists keep declaring AI red lines, but no one has mapped them or tested whether they can be checked. Mentees will build a public, sourced dataset of these commitments and convert the strongest into draft verifiable standards.

Loss of Human Agency in the Age of Advanced AI – An Agent-Based Model
Zhamilia Klycheva · Independent
An agent-based model of how populations gradually lose agency to AI-mediated manipulation and delegation — formalizing tipping points, spread dynamics, and the divergence between AI-empowered and atrophied users, aimed at a publishable computational social science paper.

Do Jailbreaks Converge? Shared Latent Signatures Across Jailbreak Families
Davide Zani · HiddenLayer
Take representative attacks from a number of jailbreak families and test whether succesful attacks converge on a shared low-dimensional subspace. H1: Successful jailbreaks across families produce convergent perturbations of the refusal subspace (compliance-shift vectors' similarity across families significantly above matched benign controls) H2: latent convergence predicts cross-family transfer. Null if compliance-shift vectors are family specific.

Disentangling persona vectors from emotion vectors in LLM activation space
Juan Cadile · University of Rochester
Persona vectors (Chen et al. 2025) and emotion representations (Sofroniew et al. 2026) have been studied independently, but plausibly overlap in activation space. We'll measure their geometric and causal relationship to determine whether persona drift and emotional-state changes are mechanistically distinct failure modes; and whether interventions on one silently move the other.

Guarding Against Malicious Fine-Tuning: Hardening Models and Detecting Tampering
Fernando Moreno-Pino · Intelligent Systems Lab, University of Bristol & Oxford-Man Institute, University of Oxford
This project studies how to prevent and detect malicious fine-tuning of large language models, focusing on both hardening models against adversarial adaptation and developing forensic tools to identify when a model has been covertly tampered with.

Robustness of Moral Consideration Under Adversarial Pressure
Jasmine Brazilek · Compassion-Aligned Machine Learning (CaML)
Investigate the training mechanisms which erode compassion instilled during midtraining. Understand how to preserve self-fulfilling alignment through adversarial attacks.

Operationalizing AI×Bio Governance: A Practical IBC Framework for Resource-Constrained Settings
Zia Ashraf · Government College University, Faisalabad
Existing biosecurity guidance increasingly recognizes artificial intelligence, information security, and the role of Institutional Biosafety Committees (IBCs), but institutions still need practical ways to translate these principles into project-level review. This project will develop and pilot-test a resource, calibrated screening framework, comprising an intake checklist, risk tiers, and escalation decision tree, to help IBCs identify and manage risks arising from AI-enabled biological research.

Human Behavioral Phenomena in Language Models
Emilio Barkett · Independent
This project investigates whether language models replicate human behavioral phenomena by adapting experimental designs from the social sciences. Mentees will select a documented human behavior and design experiments to test whether it emerges in language model outputs.

Do AI Safety Benchmarks Hold Up Under Real Deployment Prompts?
Rahul Kumar · DevRev, Independent AI Safety Researcher
AI safety benchmarks test models under clean conditions, but deployed models run with system prompts that tell them how to behave. This project measures whether models that pass the COMPL-AI EU AI Act benchmark suite still pass when you add the system prompts that real enterprise deployments actually use, and tests which minimal prompt fixes restore safety scores.

Does your assistant respect your agency? A behavioral benchmark for autonomy-preserving AI
Juan Cadile · University of Rochester
AI assistants constantly choose between empowering users and acting for them, and between honoring users' stated goals and overriding them "for their own good." We'll build a systematic benchmark measuring whether models respect user agency, covering paternalism, manipulation, dependency-fostering, and value-substitution, with philosophically grounded rubrics and human-validated LLM-as-judge scoring.

Diagnosing the Mechanisms Behind Honest-Looking Language-Model Behaviour
Bryan Chan · independent
This project will build controlled evaluations that distinguish context-general honesty from sycophancy, surface-cue shortcuts, refusal spillover, latent knowledge, and evaluation-conditioned behaviour. Mentees will test several tractable language models under matched prompt transformations and produce a validated failure taxonomy, reproducible evaluation suite, and empirical report.

When RLVR Changes the Model, the Safety Test, or Both
Muhammad Aaliyan · OCN (OneCarNow); Independent AI Safety Researcher
We will test whether safety-relevant behavior or evaluation reliability changes across reinforcement learning with verifiable rewards checkpoints. The project will replicate a completed Tülu 3.1 study on a second open training lineage and stress-test the result across prompt wording, answer order, and open-ended evaluation.