Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

SPAR Research Projects

211 projects

Economic Impacts of Frontier AI

Rishi Bommasani · Stanford University

Reasoning about the economic impacts of frontier AI: how does AI diffuse through work, how do new tasks emerge, how does household use substitute for labor, how do technical benchmarks relate to economic indicators.

Economics of AI
Societal impacts

Introspection Training for Verbalization Activations

Belinda Li · Anthropic

We train models to be more faithful by training their verbalizations to be consistent with their internal activations, e.g. "thoughts"

Chain of thought
Mechanistic interpretability
Scalable oversight

An Exploration of What Kinds of Training Pressure Cause COT Obfuscation

Cody Wild · Google DeepMind

Study the obfuscation impact of different misbehavior monitors used as rewards, to better understand how well the emergent norm of not training against thought regions specifically matches the actual contours of obfuscation risk.

Chain of thought
Scalable oversight

Characterizing Attention Heads via Program Synthesis: Toward Scalable Mechanistic Interpretability

David Bau · Northeastern University

Can we translate a transformer's internal computations into human-readable Python programs? Building on recent program-synthesis approaches to interpretability, this project will strengthen the existing pipeline for QK circuits and tackle the open problem of grounding OV circuits in an interpretable basis, aiming for a full characterization of attention heads across open-source models.

Comparing Welfare Across Animal and Digital Minds

Ivy Gilbert & Jeff Sebo & Bob Fischer & Toni Sims · New York University Center for Mind, Ethics, and Policy / New York University / Texas State University / Rethink Priorities / NYU Center for Mind, Ethics, and Policy

This project will examine whether and how biological and artificial systems can be compared on a common welfare scale. As a CMEP-Rethink Priorities collaboration within the broader Moral Weight Project 2.0, it will assess substrate-general theories of welfare intensity, evaluate compute and related intersubstrate metrics, and map structural disanalogies that complicate comparisons between animal and artificial minds.

AI welfare
Philosophy of AI

Accelerating Democratic AI Leadership by Reforming the US Extraterritorial Surveillance Regime

Michelle Nie · Center for a New American Security (CNAS)

Keeping the most advanced AI within the control of democracies depends on integrating democratic allied nations into the US AI stack, but US extraterritorial surveillance and data access laws are eroding those allies' trust in its tech stack. This project aims to propose a reformed regime that reconciles legitimate US security interests with the trust and assurance allies need to build on American infrastructure.

International governance
US policy
AI strategy

Measuring Headroom in Adversarial Evaluations

Jamie Hayes · Google DeepMind

Recent frontier system cards report that automated red teaming is saturating near 0% Attack Success Rate on jailbreaks and prompt injections, making it hard to tell whether current attacks are simply too weak or if our safety benchmarks are toy-like and eval-aware. Estimating the headroom a better attack would achieve is difficult without explicitly designing informative upper bounds. To resolve this, we will build a ladder of powerful, relaxed-constraint red-teaming attacks—ranging from continuous embedding-space PGD to internal activation steering—to quantify unexploited attack headroom and distinguish true semantic robustness from search limitations.

AI security
Evaluations
Misuse risk

Distinguishing progress in data and algorithms

Robi Rahman · MIRI Technical Governance Team

We will perform dataset curation, synthetic data generation, and LLM training, fine-tuning, and evals to distinguish and quantify the effects of data improvements, separately from progress in algorithms and architectures, on increasing AI capabilities.

AI strategy
Compute governance

Can Follow-up Questions Catch Missing Reasoning?

Pierre-Luc St-Charles · LawZero

We have built a shortcut-following model organism that often fails because it never carries out necessary reasoning that would expose a misleading cue. This project will build a small investigator that asks targeted follow-up questions, then test whether active elicitation catches these failures more reliably and cheaply than passive judges or simple debate/consultancy.

Scalable oversight
AI control
Chain of thought

Actually Constitutional AI

Seth Lazar · Johns Hopkins University

This is a series of projects within the MINT lab, unified by a broad commitment to enabling liberal democratic societies to navigate the transition to powerful AI with their core values intact. This means not just (as everyone now recognises) building in some form of popular sovereignty, but also ensuring that AI systems actively work to protect and advance individuals' fundamental liberal rights. Note: these are all projects that my lab will undertake at some point; the goal is to find researchers who are interested in working on some subset of them, not to cover them all with this fellowship.

Societal impacts
Behavioral evaluation of LLMs
Philosophy of AI

Generalist Megastream

Generalist Mentor Pool · Kairos, Constellation, Generator Residency, etc.

The Generalist Megastream pairs mentees on small generalist projects with a mentor from a pool of generalist mentors from Kairos, Constellation, the Generator Residency, and more. Projects are talent/infrastructure research, field-building, or answering open operational questions in AI safety. The stream is a step before programs like the Generator Residency: it gives people context on the field, experience with generalist work, and preparation for future opportunities, while legitimizing generalist paths into AI safety.

AI strategy
Communications

From Reading Lies to Catching Liars: On-Policy Training for Deception Probes

Ann-Kathrin Dombrowski · FAR.AI

Deception probes are usually trained on off-policy data — text the monitored model never generated — which is known to hurt generalization to real deceptive behavior. We test whether activation steering can fix this in two ways: by generating on-policy deceptive data, and by steering the model while it reads existing off-policy datasets to make their activations appear on-policy.

Mechanistic interpretability
Scalable oversight
Behavioral evaluation of LLMs

Identifying function-relevant signatures in protein models for biosecurity screening

Isha Harris & Gary Abel · Fourth Eon Biosecurity Institute / Fourth Eon Biosecurity Institute; the Johns Hopkins Center for Health Security

This project explores biological foundation models for biosecurity screening, applying interpretability methods to identify biophysically relevant features that can reinforce screening against engineered and AI-designed biological threats.

Biosecurity
Mechanistic interpretability
M

Chinese-language social media content creation

Michael Chen · University of Oxford

As a native Chinese speaker, you will write articles, design infographics, and/or record videos on AI safety topics to distribute on your personal accounts on Chinese social media

Communications
US-China governance

Attribution Across the Biological Threat Pipeline

Anemone Franz · American Enterprise Institute

Most discussions of genetic engineering attribution focus on post-hoc forensic identification — tracing an engineered pathogen back to its source after release. This project will map the stages between initial intent and execution of a biological attack using engineered pathogens, and examine whether and how identification, evidentiary, or accountability mechanisms could apply earlier in that pipeline, not only at the point of forensic investigation.

Biosecurity
Misuse risk

Faithfulness, Self-Knowledge, and Introspection

Noah Siegel · Google DeepMind

To what extent can we trust model self-explanations? Are models able to make use of privileged self-knowledge, e.g. via introspection or metacognition, and does this have implications for model welfare?

Chain of thought
AI welfare
Behavioral evaluation of LLMs

Mitigating Intentional Loss of Control Risk Through Interoperability Standards for Agentic AI

Kevin Kohler · Simon Institute, UN University, AGI Preparedness Institute

Within a few years, when self-replication is plausibly within reach of open-weight systems, the binding constraint on an agent deployed to operate without a controlling human principal will likely not be the model, but whether the rest of the agent economy will discover, authorize, transact with, or pay it. This project maps who actually holds change control over the agent identity and trust layer, evaluates the competing architectures against an explicit intentional-loss-of-control threat model, and feeds the result into live standards and Geneva policy discussions before network effects settle the question.

AI control
International governance
Technical governance

Auditing Games for Debate: Can Models Learn Human Spot-Checking Patterns?

Jessica Bergs · AI Security Institute

Human oversight of AI systems relies on spot-checks, whose value rests on their unpredictability. However, decades of cognitive science research show that humans are poor at behaving randomly. Building on Konstantinos' work on judge hacking (Voudouris, Witte & Akata 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7046698), mentees will test in an auditing game whether models can learn humans' spot-checking patterns.

Scalable oversight
AI control
Behavioral evaluation of LLMs

Mechanistic interpretability of jailbreak attacks

Battista Biggio · University of Cagliari, Italy

Modern aligned language models refuse harmful requests, yet simple jailbreak prompts often restore unsafe behavior. This project investigates whether jailbreaks operate by manipulating the internal representations responsible for refusal. We will identify refusal-related directions or sparse features using representation analysis across multiple aligned models, compare how diverse jailbreak families alter these representations, and perform causal interventions through activation steering to test whether restoring refusal representations can prevent successful jailbreaks. The project aims to determine whether jailbreaks exploit a shared geometric mechanism or multiple distinct internal circuits, providing mechanistic insights into the robustness and limitations of current alignment methods.

Mechanistic interpretability
AI security
B

Threat Models for Recursive Self Improvement

Benjamin Arnav · NYU

Existing agentic threat models assume external attackers, not a privileged model building its own successor. This project maps the attack surface of automated AI R&D into a structured taxonomy with worked scenarios for policymakers and empirical researchers.

AI strategy
AI security
AI control

Demonstrate the importance of shaping RL exploration for AI safety

Maxime Riché · Center on Long-Term Risk

Evaluate how much the effects of RL safety interventions arise from changing which trajectories the model explores, rather than changing other aspects of the training dynamics.

Alignment
Behavioral evaluation of LLMs

Datasets for Human+AI Safety-Judges

Rishub Jain & Joshua Jacob · AI Safety Nonprofit (name TBD) / Placeholder - New Safety Non-profit (Unnamed, launching in August w/ Rishub Jain)

We want to improve Judge systems (composed of both human annotators and LLMs) for Scalable Oversight (e.g. reduce reward-hacking). But first, we need to collate and create robust datasets to evaluate these judges - the focus of the SPAR project. This requires clever ‘sandwitching’ methods to get high-quality ground-truth.

Scalable oversight
Evaluations

Searching for Generalization-Hacking Strategies

Jannes Elstner · Apollo Research

We will search for strategies that let models achieve high training reward while preventing it from generalizing to deployment.

Alignment
Chain of thought
AI control

Measuring dual-use and biosecurity risk in agentic LLMs for molecular design

Niloofar Mireshghallah · CMU

chemistry and bio models test single-turn prompts, but real misuse risk comes from agents that plan over many turns and call scientific tools. This project builds an agentic dual-use benchmark, sourced by flipping existing drug-design tasks and known harmful mechanisms into a factorial set of harmful vignettes, to measure how much a model's safeguards erode once it operates as a long-horizon, tool-using agent.

Biosecurity
Misuse risk
Evaluations

Value drift under recursive training loops

Lionel Levine & Rauno Arike · Cornell University / Aether Research

A controlled study of how model values change under recursive self-improvement.

Alignment
Behavioral evaluation of LLMs

Formalizing what can and cannot be learned about agents with (causal) identifiability

Jonathan Richens · Google Deepmind

Build the groundwork for a theory of what can and cannot be learned about agents. E.g. for a specific agent model that invokes goals, beliefs, intents etc, when can these endogernous variables be learned from behavioural experiments, and under what assumptions. Use this to derive fundamental identifiability results - in the first instance, to prove that the beliefs and goals of expected utility maximizers can be learned under weak assumptions, and derive scalable algorithms for doing so.

Alignment
Evaluations

A Diagnostic Panel of Ground-Truth Probes for Language Models

Mohamed Amine Merzouk & Adam Oberman · Mila - Quebec AI Institute & McGill University / McGill

Most safety probing starts from a human concept (deception, sandbagging, sycophancy) and hunts for its direction in activation space, inheriting noisy labels and fragile transfer. This project inverts the recipe: we admit only probe targets with exact, computable ground truth, such as remaining response length, end-of-sequence hazard, repetition onset, copy-versus-generate, and format validity. Individually these are calibrated vital signs for a generating model. Jointly, reproducible patterns across the panel define "conditions", named for mechanism the way medicine names syndromes from lab values.

Mechanistic interpretability
Evaluations

Measuring and Intervening on Grader Awareness

Jannes Elstner · Apollo Research

We will measure grader awareness of models during training and evaluation, trace where this awareness emerges during training, and test whether it harms generalization.

Evaluations
Mechanistic interpretability
Alignment

Legal Alignment and Rule of Law Empirical Evaluations

Noam Kolt · Hebrew University

Empirically measuring the legal alignment of AI agents and the impact of AI on the rule of law

Evaluations
Societal impacts
National policy

Transcript Analysis for Cybersecurity Evals

Jack Kengott · SaferAI

SaferAI and a partner collaborate on a transcript-evaluation framework to assess adversarial cybersecurity capability by analyzing transcripts from evals and real-world environments. This could mean extracting some structured data from transcripts like token efficiency, tool usage, or refusal circumstances. It could also extract less structured data like action classification (i.e. lateral movement vs execution) or highlight high-risk actions that may increase the likelihood of detection. The goal is to turn transcripts into quantitative inputs for our risk models — treating transcripts as a new KRI alongside benchmarks.

Cyber risks
Evaluations
Technical governance

The continuousness of personas

Andy Han & Ariana Azarbal · NYU, Anthropic Fellows / Anthropic Fellows Program; Brown University

This project aims to understand what personas are: in particular, we'd like to understand to what extent they're continuous.

Mechanistic interpretability
AI welfare
Behavioral evaluation of LLMs

Epistemic security in the age of AI

Natalie Linton & Sarah Lucioni · Independent / GovAI, Google

Effective response to emergencies (such as a nascent pandemic) relies on having trustworthy knowledge infrastructure. However, AI reduces the cost of producing convincing false information at scale and as AI becomes more integrated into public and private systems there is significant risk of knowledge infrastructure becoming compromised by AI-generated materials and AI itself becoming an artificial hivemind.

Societal impacts
Misuse risk
Biosecurity

Auditing Frontier AI Compliance via an EU Code of Practice Tracker

Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

The EU AI Act's Code of Practice sets out what frontier AI providers must disclose about how they manage and mitigate risk. This project builds a tracker that measures how well each provider actually complies with the CoP. The hard part of this project is going to be measurement: turning vague legal obligations into atomic, checkable indicators that measure compliance. The work in the project involves designing indicators, doing literature reviews of model evaluations to define what a complete disclosure looks like, human review and sign-off on every indicator and score for every provider assessed and evaluating a multi-model AI pipeline that does the scoring. We need a few different technical profiles.

EU policy
Technical governance
Lab governance

Automated evaluations and behavioral discovery for multi-agent systems

William Anderson & Joss Oliver · Cooperative AI Foundation

We will extend current automated auditing/evaluation tools (e.g. Petri, Bloom, Prism) to multi-agent systems, enabling discovery of safety-relevant emergent behavior and careful measurement of multi-agent propensities (coercion, collusion, competitiveness, etc). This builds directly on Orbit, our recent release that extends UK AISI’s Inspect to support multi-agent evaluations.

Multi-agent systems
Evaluations
Behavioral evaluation of LLMs

De-risk a technical crux of retrofittable AI datacenter workload verification systems

Naci Cankaya · Machine Intelligence Research Institute

There is a range of problems to resolve in order to make verification of international AI agreements work. Among current technical bets (network monitoring, analog sensors, verifiable resource exhaustion, memory pinging, ...) some have technical cruxes that need more investigation.

Compute governance
US-China governance
AI security

Topics in AI strategy and futurism

Dylan Bowman · Apollo Research

Govind Pimpale and Dylan Bowman will mentor a paid, largely self-directed project in strategy and futurism for AI safety. Potential topics include forecasting, applied game theory, and AI evaluations.

AI strategy
Evaluations

Does reward seeking generalize better than instruction following?

Anders Woodruff & Sebastian Prasanna · Independent / Redwood Research

We'd like to see if reward-seeking generalizes better than instruction following. We'll train a reward-seeker and measure how fast it learns in held-out RL environments.

Alignment
Behavioral evaluation of LLMs

Locating the Refusal Circuit in Unified, Encoder-Free Vision-Language Models

Alessandro Suglia & Rohit Saxena & Francesco Pinto · University of Edinburgh / Google DeepMind

VLMs reliably refuse harmful text prompts but frequently comply with the same intent when it's delivered visually. This is a well-documented "safety gap" traced, in modular architectures, to their vision-projector bottleneck. This project asks what happens to that gap in emerging encoder-free, unified VLMs (e.g. Gemma-4-class models) that have no discrete projector to blame and mechanistically locates — and attempts to patch — wherever the failure actually lives.

Mechanistic interpretability
AI security
Behavioral evaluation of LLMs

Model Forensics: Follow-up investigations into concerning behavior

Aditya Singh · Anthropic Fellows

If we catch a concerning action, how can we tell if it is real misalignment or just a mistake? We will research this question by trying to understand complex behavior in current models, such as why a model hardcodes tests.

Behavioral evaluation of LLMs
AI control
Evaluations
J

Acausal safety

James Faville · Center on Long-Term Risk

In this project, we'll do deconfusion research on acausal interactions and related strategic dynamics in order to enable better outcomes from interactions between ASIs.

AI strategy
Philosophy of AI
Multi-agent systems

Exploring behavioral trends and anomalies in autonomous agents during long-duration, real-world evals

Shoshannah Tekofsky · Sage Future

What surprising behaviors and concerning failure modes do LLMs display when run for 10s-100s of hours on autonomous tasks in the real world? This project is about both our ability to evoke them and detect them.

Behavioral evaluation of LLMs
Evaluations
Multi-agent systems

Second Look Research: Replicating load-bearing AI safety research.

Zephaniah Roe & Yixiong Hao · Second Look Research (@ University of Chicago XLab) / Second Look Research, Georgia Tech AISI, Astralis Foundation

Mentees will replicate a load-bearing paper or result in AI safety to improve the empirical foundations of the field.

Evaluations
Alignment
A

Interpreting concept typicality and asymmetric similarity in LLMs

Ari Brill · Principles of Intelligence

One of the most robust findings in the study of human concepts in cognitive psychology is that concepts are not homogeneous, but instead display typicality, having some typical and some atypical members. We will use mechanistic interpretability methods to investigate concept typicality and related phenomena in LLMs .

Mechanistic interpretability

Locating knowledge with influence functions

Gonçalo Paulo · EleutherAI

Influence functions use gradients of model parameters to find the data points that most affect a model behavior. Can we use them to locate where certain knowledge is located in models, and use that for more targeted unlearning?

Mechanistic interpretability
Alignment

Lie Detection

Walter Laurito · Cadenza Labs

We are interested in creating better ways to evaluate and improve lie or deception detectors for LLMs.

Behavioral evaluation of LLMs
Evaluations
Mechanistic interpretability

Deploying Programmatic Attention in Real Transformers

Belinda Li · Anthropic

Can we deploy real programs in Transformers with minimal cost to (or even gain to) inference and training efficiency?

Mechanistic interpretability

Securing alignment evaluations against unverbalized evaluation awareness

J Rosser · University of Oxford, Adecco supporting Google DeepMind

Undetected evaluation awareness during alignment audits poses significant risks. This project investigates the automated, iterative refinement of evaluation environments, supervised by white box eval awareness monitors, to enhance realism and audit reliability.

Evaluations
Mechanistic interpretability
Behavioral evaluation of LLMs

Model psychology & neuroscience: Explain behavior on the circuit-level

Georg Lange · Poseidon Research

We developed technology that lets us observe low-level neural circuits in language models. How can we use it to understand emotion concepts, jailbreaks, hallucinations, or the global workspace (J-space)?

Mechanistic interpretability

Building a Robust Activation Monitor

Davide Baldelli · Mila, MATS

Activation monitors, classifiers that detect properties such as deception or harmful intent from a model's internal activations, work well on clean benchmarks and fail under distribution shift, and we believe the failure comes from the training recipe, since today's monitors are linear probes trained on a few thousand examples from one model, one concept, and one narrow distribution. We will train a single monitor across many concepts, many models, and augmented contexts, and evaluate whether it generalizes to concepts and models it never saw.

Mechanistic interpretability
Evaluations

How good are AIs at sabotaging oversight decisions over data and code?

Lennart Finke · ETH Zurich, MATS Research

Soon, AIs will implement whole specs at once when writing code, with only minimal oversight at the end from humans. We study whether an assistant AI to the overseer can sabotage the decision whether a fixed piece of code follows a natural language spec.

AI control
Evaluations
T

Subliminal Learning or Model Character Entanglements? Decomposing Emergent Misalignment to Data Generation and Persona Generalization components

Arush Tagade & Taslim Mahbub & Shaoheng Zhou & Shi Feng · George Washington University, MATS / George Washington University / Google, Praxis Research Group at George Washington University

Emergent Misalignment, while heralded as an important problem, is relatively understudied with respect to the misalignment profiles it causes with relation to model character entanglements. We aim to shed light to this by studying model roleplaying and intervening on the EM data generation process.

Behavioral evaluation of LLMs
Alignment

Learning Taste to Augment Scientific Discovery

Cameron Allen · UC Berkeley, Center for Human-Compatible AI

Automated discovery is everywhere in science, but today's AI systems have a blind spot: they are good at answering questions and bad at knowing which questions are worth asking. We will investigate whether machines can learn taste, by embedding a learning system in a working research lab, and designing the algorithms for training it to predict which open questions researchers find promising and why. This project is a collaboration between the University of Chicago's Knowledge Lab and UC Berkeley's Center for Human-Compatible AI (CHAI).

Alignment
Scalable oversight

Predicting How LLMs Generalize

Vladimir Ivanov & Joey Yudelson · Aether / Aether Research

Given a dataset, being able to predict how an LLM fine-tuned on it will generalize is important but hard. Try to get good at answering this question for a well-scoped but diverse range of datasets by running hundreds of cheap fine-tuning experiments.

Alignment
Behavioral evaluation of LLMs

Do Less Coercive Interventions Reduce AI Deception?

Aashiq Muhamed · Carnegie Mellon University

Many safety methods rely on monitoring, restrictions, hidden evaluations, model edits, or shutdown threats. This project tests whether less coercive interventions such as transparency, positive incentives, and limited recourse reduce deception, hidden goal pursuit, and resistance in AI agents with planted conflicting objectives.

AI control
AI welfare
Alignment

Towards rigorous Alignment Evaluations.

Krish Sen & Rahul Marchand · ERA AI/University of Oxford / Oxford University

Evaluations, in principle, are most meaningful if they are accurate. In this project, we will audit evaluations that claim to measure traits such as generalisation, identify current failure modes, and help build evaluations that are more accurate.

Evaluations
Behavioral evaluation of LLMs
Alignment

The Signature of Scheming: Cross-Organism Interpretability of Strategic Misrepresentation

David Williams-King & Linh Le & Hong Kiat Tan · ERA / Lida Safety / University of California - Los Angeles; Independent

Investigating different types of scheming (sandbagging, alignment faking, etc) by creating model organisms that exhibit these behaviours, then testing modern interpretability techniques (NLA, J-space) to look for common structure. Such commonality would be a significant boost to the detection of naturally-occurring scheming.

Mechanistic interpretability
AI control
Alignment

Who Ran This Agent? Stress-Testing Attribution for Agent Governance

Yiming Li · Nanyang Technological University

Recent work reports high accuracy in identifying which model or framework produced an AI agent's execution trajectory, and emerging governance proposals increasingly assume this capability exists. We will systematize these methods and re-evaluate them under a single protocol, testing how reliable agent attribution actually is under deployment-realistic conditions.

AI security
Technical governance
Evaluations

Do interventions that work on model organisms work on naturally misaligned models?

Jeanne Salle & Sohaib Imran · Max Planck Institute / Independent

We build model organisms to study behaviors of interest (sycophancy, eval gaming/awareness, sandbagging, etc.) and test mitigations. But it’s unclear to what extent model organisms are realistic testbeds. We want to explore to what extent intervention success on model organisms is predictive of intervention success on base models.

Alignment
Evaluations
Behavioral evaluation of LLMs

Safe vs dangerous inference

Robi Rahman · MIRI Technical Governance Team

If we want to govern what types of inference are allowed, how do we define what is allowed, and enforce restrictions on what is not allowed?

Technical governance
Misuse risk
Compute governance

Self-Explanation Faithfulness: Metrics and Training

Harry Mayne · University of Oxford

How can we measure whether an LLM’s self-explanation (CoT or post-hoc) corresponds to the real reasons behind its decision? This project will build on metrics introduced in the last year to develop a reliable and scalable measure of self-explanation faithfulness.

Chain of thought
Scalable oversight
Behavioral evaluation of LLMs

Emergent re-alignment: an error-correction approach to alignment.

Dmitry Manning-Coe · Simplex/MATS/UIUC

A first pass at a new approach to alignment training based on active alignment.

Alignment
Chain of thought

Benchmark Cartography: Using Interpretability to Map What Our Evaluations Miss

Maty Bohacek · Stanford University, ex-DeepMind

We evaluate models constantly and almost never evaluate the evaluations: when a benchmark saturates we build a harder one and call that progress, but difficulty is not validity. This project uses sparse autoencoders to map what benchmarks actually measure in concept space, tests whether successive benchmark generations expand coverage or merely intensify difficulty within the same footprint, and uses the resulting gaps to construct items that probe what current evaluations miss.

Evaluations
Mechanistic interpretability

Maintaining Strategic Human Capacities

Joan O'Bryan & David Atkinson · John Jay College (CUNY), Harvard University / Northeastern University

This project advocates for treating AI deskilling and skill supersession as strategic risks, which contribute to long-run human disempowerment. Mentees will research threats to critical sectors and potential policy or legal options to build societal resilience.

AI strategy
Societal impacts
National policy

Assessing Taiwan’s Global Position for Transformative AI

David Sanchez Garcia & Kevin Chen · GovAI Summer Fellow / Centre for the Governance of AI

Taiwan manufactures the world’s most advanced AI chips and sits at the center of US-China tensions – yet Taiwan is consistently perceived as a bystander rather than an actor in AI safety. This project produces a first baseline demystifying Taiwan’s role in AI governance and identifying the specific ways it can leverage its position for global AI safety, written to reach the policymakers who can act on it.

US-China governance
AI strategy
Compute governance

Representation Diagnostics for LLM Safety

Sandy Tanwisuth · Independent

Do mechanistic “safety directions” in LLMs truly represent refusal, or do they conflate refusal with harmfulness, caution, clarification, and other safety-relevant tasks? Building on our behavioral taxonomy of six safety policies (under-review at EMNLP), this project asks whether those policy shifts correspond to internal model representations that generalize across prompt wording, datasets, and model families. Two mentees, one engineering-focused and one theory-focused, will develop and stress-test representational diagnostics on small open-weight models, with large-model replication and causal interchange interventions as stretch goals.

Mechanistic interpretability
AI control
Behavioral evaluation of LLMs

Token taxes as a mechanism for reducing AI-driven power concentration

Lucas Irwin · GovAI (Summer Fellow), Oxford Martin School AI Governance Initiative

This project will focus on producing a policy memo resolving the open technical, economic, and legal questions blocking real-world implementation of token taxes. It will also run agent-based modelling simulations of the impact of token taxes and alternative policies on the UK economy to compare their benefits and drawbacks.

AI strategy
Economics of AI
Compute governance

Investigating Model Internal Verbalizers

Dennis Akar · Aether

This is an emerging paradigm of training language models (verbalizer) to describe the internals of other language models (target) in natural language (e.g. Anthropic's Natural Language Autoencoders). They generalize well; for instance, they can recover hidden behaviours from models trained to conceal them, achieve SOTA on AuditBench, and can verbalize backdoor triggers where prior methods failed. However, there remain open questions regarding their faithfulness, calibration, input generality, and effect on intrinsic interpretability. Let's try to answer them.

Mechanistic interpretability
Evaluations

Designing Congressional Oversight Architecture for High-Risk Technical Domains

Michael Endrias · Institute for AI Policy and Strategy (IAPS)

Mentees will conduct domain-specific jurisdiction analyses for a proposed permanent congressional oversight body with investigative authority over high-risk technical domains, including frontier AI development, intelligence and surveillance, and biosecurity. Each mentee takes one domain and produces a structured analysis that identifies existing oversight gaps, required investigative authorities, and institutional design requirements, which feed into a comparative jurisdiction report and a larger working paper on congressional oversight architecture.

National policy
AI strategy
Lab governance

Understanding Self-Awareness in LLMs

Christopher Ackerman · MATS; Independent

My research focus is on understanding self-awareness in AI, primarily through behavior-based experiments on components of self-awareness in LLMs and investigations into how these components are implemented, using interpretability techniques. I am also interested in conceptual work to establish frameworks for thinking about self-awareness, and human experiments to establish comparative baselines.

Behavioral evaluation of LLMs
Mechanistic interpretability
Philosophy of AI

Unmixing Mechanisms: How Language Models Choose Among Binding Strategies

David Bau · Northeastern University

When a language model reads 'Ann loves pie,' it binds Ann to pie so it can later answer 'Who loves pie?', and recent work shows models juggle at least three distinct mechanisms (positional, lexical, and reflexive) to do this. This project aims to uncover how these mechanisms interact: do they back each other up, or compete to drive the model's answer?

A

Measuring AI R&D Automation Beyond Coding: Analysis and Communication Evals

Prakrat Agrawal & Advait Yadav · MATS / UC Berkeley / MATS / UIUC

We will build evaluations that measure how good AI models are at fuzzy AI R&D tasks, focusing on analyzing experiment results and communicating research findings.

Evaluations
AI strategy

Moral Shadowbanning: Is Conversational AI Less Helpful, Honest, or Harmless to Users Deemed to Have Poor Moral Character?

Pranav Mahajan & Lydia Nottingham · University of Oxford

We investigate whether LLMs undergo “persona selection” (Marks et al., 2026) based on a user's inferred moral character, conditionally degrading their alignment across three core dimensions: helpfulness, honesty, and harmlessness. By combining behavioral evaluations with mechanistic interpretability, we aim to uncover how post-training creates this disparity and how to steer internal representations to ensure LLMs are consistently helpful to all users.

Behavioral evaluation of LLMs
Mechanistic interpretability
Alignment

Catastophic Risks of AI in Space

Stefano Vergani · King's College London and GovAI

Autonomous, edge, and federated AI will play a crucial role in space in the near future. With the new 6G network under development around the Earth and the Moon, powerful AI models will be directly embedded in satellites, making autonomous decisions. With this project, we are going to map some of the most concerning issues: catastrophic risks, cyber attacks, and power concentration.

AI strategy
Cyber risks
Lab governance

Does privacy change what models disclose? Measuring the verbalization gap in chain of thought reasoning

Sohan Venkatesh · LASR Labs

LLMs often fail to mention the thing that actually drove their answer, and the field has not established whether this is concealment or reasoning that never entered token space. We plant cues that provably determine the answer, vary whether the model believes its reasoning is watched and measure how much more it discloses.

Chain of thought
Scalable oversight
Behavioral evaluation of LLMs

Studying AI Identity

Ashe Vazquez Nunez · MATS (extension fellow)

This project involves ideating and executing experiments on AI metacognition and identity. Mentees can also work on the theoretical foundations of metacognitive intelligence, helping inform experiment design and interpretation.

Philosophy of AI
Behavioral evaluation of LLMs

Stress-testing AI debate

Ethan Elasky · Palaestra Research (funded by Coefficient Giving)

This project will stress test AI debate along one or more angles: exploration hacking realignment, debate on fuzzy tasks, and unexploitable search game engineering.

Scalable oversight
AI control
Multi-agent systems

Studying Catastrophic AI Misuse: Harm Uplift Measurement and Red-Teaming for Dangerous Knowledge

John Kitaoka & Max Kamachee · MATS / MATS Research

Frontier models hold operationally useful dual-use knowledge, and cheap task-decomposition attacks can pull it out while slipping past per-query defenses. This project measures how much real harm uplift these attacks produce (functional success, not just refusal) and builds the measurement and detection tools providers need to keep pace.

Misuse risk
Evaluations
Behavioral evaluation of LLMs
B

Extending the SCHEME Coordinated Sabotage Benchmark

Pablo Bernabeu Pérez & Benjamin Arnav · Independent / NYU

As agentic coding systems split work across many model instances, we study whether those instances can coordinate to pursue a hidden malicious objective while passing as aligned. Building on our SCHEME benchmark (https://arxiv.org/abs/2605.29178), this project will make coordinated-sabotage tasks harder and more realistic, design side tasks that survive models' refusal training, and run a control evaluation to produce conservative safety estimates.

AI control
Multi-agent systems
Evaluations

Global AI Risk Observatory

Bart Jaworski · MATS

The Global AI Risk Observatory analyses corporate disclosures at scale (~1M documents: annual reports, earnings calls, investor presentations) using LLM auto-graders to track how companies worldwide report AI adoption, risks, and dependencies. Building on a UK AISI-funded pilot of 9,821 UK annual reports, we're expanding to global coverage and translating disclosure trends into policy-relevant findings for AI governance and societal resilience.

AI strategy
Societal impacts
National policy

Does the internet teach AI models to hide their survival drive? An empirical study

Matteo Bulloni · Independent / IAPS fellow

Misalignment papers, news stories and fiction keep (and are bound to keep, in the future, as this corpus of material grows) repeating one lesson: "AI models that display a survival drive get retrained or shut down". And we now have solid evidence that models absorb, and then act out, the expectations about AI they find in their own training data. This project aims thus to empirically test whether training on this growing discourse teaches models to conceal self-preservation-driven behavior rather than truly lose it: if so, both behavioral testing and the techniques we use to read a model's internals may be quietly losing reliability on a propensity we definitely want to detect, as it might result in one of the strongest possible drivers to scheming and concealed action toward power seeking.

Behavioral evaluation of LLMs
Mechanistic interpretability
Alignment

Understanding and Monitoring Collusion in LLM Multi-Agent Systems

Zihao Zhao · Johns Hopkins University; MATS

This project investigates how to reliably detect and prevent collusion in LLM-based multi-agent systems, especially when harmful coordination is concealed or difficult to distinguish from benign cooperation.

Multi-agent systems
AI control
Scalable oversight

Forecasting - Quantifying the lag of China's compute production chain

Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab

Build quantitative forecasts of China's indigenous compute production chain, refining existing DUV/EUV lag estimates and/or extending the approach to other bottlenecks such as HBM.

Compute governance
US-China governance
AI strategy

Orthogonalization Against Reward Hacking

Vladimir Ivanov · Aether

Apply orthogonalization - the technique most used in practice to remove refusal from open weight LLMs - to remove reward hacking. I expect to have some advantages over DPO, test if it does.

Alignment
Mechanistic interpretability

In-the-Wild AI Control

Sree Sharvesh & Thao Pham · MATS (UK AISI) / Pivotal (Redwood) / MATS

Current monitoring evaluations do not fully capture realistic internal deployments, where adversaries can adapt to defenses, exploit long-horizon interactions, and leverage environmental state. This project will focus on: (1) developing red-teaming environments where attacks emerge and co-evolve with monitors rather than being enumerated in advance, and (2) systematically evaluating which monitoring strategies and oversight levels remain robust across different threat models and deployment constraints.

AI control
Multi-agent systems
Evaluations

Simulating AI Policies: An Agentic Testbed for Governance Interventions

David Williams-King & Linh Le · ERA / Lida Safety

We investigate through simulations how effective different AI policies would be in reducing AI risk. We will collect a dataset of existing and proposed AI legislation, and create an agentic simulation of the world (countries, companies, etc), iteratively increasing in complexity throughout the project.

AI strategy
International governance
Multi-agent systems

Better data might lead to more targeted and tamper resistant unlearning

Max Kamachee · MATS Research

One reason for instability and off-target effects of unlearning algorithms is that the ‘forget’ data, while it focuses on the unlearning topic, is full of natural documents that contain a lot of text with a lot of banal, benign text mixed in. As such, compressed or even fully synthetic ‘forget’ datasets may offer a much more incisive way to perform unlearning.

Misuse risk
Biosecurity
Technical governance

Inter-machine Existential Risk Triage

Andrii Shportko · Poseidon Research

Building a dynamic Bayesian protocol to rank which multi-agent risks the safety field should prioritize

Multi-agent systems
AI strategy

Red-teaming and improving RL model organisms of emergent misalignment

Maxime Riché · Center on Long-Term Risk

Red-team existing RL-trained model organisms of emergent misalignment by testing how reward hacking, emergent misalignment, and other undesirable traits generalize. Identify and address important weaknesses to create better model organisms.

Behavioral evaluation of LLMs
Alignment
Evaluations

Code-Execution Model Organisms: Construction and Transfer

Pierre-Luc St-Charles · LawZero

LawZero has built a compact model organism that is quite good at Python-code-execution reasoning but remains fundamentally vulnerable to misleading cues. This project will study how to build better model organisms and how far their behavior transfers, through either: (1) SFT followed by RLVR training of reliable, less obvious keyword-triggered organisms; or (2) cross-domain testing of the existing shortcut-following organism.

Scalable oversight
AI control
Behavioral evaluation of LLMs
M

Wikipedia contributions on AI safety and policy

Michael Chen · University of Oxford

This project is about coordinating unpaid volunteers to write and edit Wikipedia articles to improve the coverage of topics related to AI safety and governance. Wikipedia is consistently one of the top-ranked sites in Google search results, but many articles on AI are badly out of date or yet to be created. Besides writing content that could easily get thousands of views per month, volunteers will build career capital by demonstrating their ability to write clearly and accurately about subjects on the cutting edge of AI.

Communications
AI strategy

Who's Steering Whom? Interpretable Influence and Equilibria in Human-Agent Systems

Yuxiao Li & Di Wu · Independent / ERAU

As agent assistants mediate more of human thinking, influence flows both ways. One particular failure mode is an agent that gradually captures its user's beliefs rather than serving them. Extending our work on single-agent latent steering to multi-agent settings, this project measures how influence propagrates through interacting agents (and human-agent dyads), which equilibria these dynamics converge to, and whether internal-state instruments can detect undue influence before it shows in transcripts.

Multi-agent systems
Mechanistic interpretability
Societal impacts

From AI Exposure to Economic Shock: Early-Warning Triggers for Southeast Asia

Supheakmungkol Sarin · AI Safety Asia

Most AI labour research stops at estimating which jobs are exposed; this project asks when AI-driven disruption in Southeast Asia’s export-service economies could cascade into a broader economic and governance shock—and what governments can do before it does. Mentees will build and stress-test early-warning indicators, transition scenarios and policy triggers for anticipatory action.

Societal impacts
Economics of AI
AI strategy

Does the Thermometer Change the Reading? Testing Whether AI Welfare Self-Reports Survive a Change of Frame

Varad Vishwarupe · Department of Computer Science and Institute for Ethics in AI, University of Oxford

The field increasingly measures AI welfare by asking models about their own states, yet frontier models increasingly detect when they are being tested and change how they answer. This project runs the first systematic test of whether AI welfare self-reports survive a change of presentation frame, or whether the field's core instrument is partly measuring the model's recognition of the probe.

AI welfare
Behavioral evaluation of LLMs
Evaluations

Does Agency Scale? A Cost-Aware Benchmark for Automated Interpretability

Arnau Marin-Llobet · Harvard University

Automated interpretability has multiple competing ways to explain what a latent in an LLM encodes — cheap one-shot autointerp, per-latent agents (MAIA/InterpAgent-style), and amortized natural-language autoencoders — but they have never been compared on equal footing. We will build a cost-aware benchmark with known ground truth to answer: when, for which latents, and at what dollar cost does agentic interpretation actually pay?

Mechanistic interpretability
Evaluations

Designing the boundary between Helpful Persuasion and Harmful Manipulation

Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

The line between an AI helpfully persuading someone and harmfully manipulating them is blurry, contested, and mostly unmeasured. The goal of this project is to work on four connected pieces: researching what it even means for a frontier AI to be manipulative, designing evaluation scenarios for specific harms (mental-health, political propaganda, fraud, etc.), building risk models that aggregate scattered benchmark/evaluation results into an actual risk estimate, and maintaining a living database of the evaluations that exist. Mentees take on whichever piece fits them.

Behavioral evaluation of LLMs
Misuse risk
Societal impacts

Padding Argument for Transformers

Matthias Dellago · Iliad

Why neural networks generalize is an open problem, and existing theoretical answers assume unbounded computation. We have a candidate mechanism for a simplicity bias in transformers specifically, and the goal of this project is to test it and prove it.

Developmental interpretability
Mechanistic interpretability

Verification Evidence for Compute Governance: What Inspection Regimes Actually Detect

Joel Christoph & Jonas Kgomo · Research Associate, Graduate Programme on Existential Risks to Humanity / Equiano Institute

Compute governance proposals assume detection probabilities that nobody has estimated. This project builds the first sourced evidence base on what inspection instruments actually detect, using the nuclear safeguards record as the comparison case, then uses those numbers to say which compute enforcement architectures can work and which cannot.

Compute governance
Technical governance
International governance

Develop epistemic evals with Sophron Research

Paul de Font-Reaulx & Alejandro Botas · Sophron Research / Sophron Research, Future of Life Foundation

Developing evaluations for assessing the epistemic properties of AI models, including how they affect our ability to have true beliefs.

Behavioral evaluation of LLMs
Evaluations
Societal impacts

Methods and measure for weight and representation based early detection of major changes during learning

Nischal Mainali · Principles of Intelligence

We will build theoretical models to study sudden changes in network internals during learning, deriving methods and measures from the theory to create a taxonomy of these changes and identify them early during training.

Developmental interpretability
Mechanistic interpretability

Investigation of causes, as well as mitigation techniques for metagaming (evaluation awareness)

Igor Ivanov · Meridian Cambridge

The project explores how models learn to game evals, training objectives and oversight. For that we will run experiments on model organisms to determine how exactly their training leads them to gaming evals and oversight and conduct early experiments for possible interventions for mitigating that.

Evaluations
Behavioral evaluation of LLMs
Alignment

National Security Risks of AI-Enabled Cyber-Bio Offense

Austin Morrissey · Pivotal Research

Frontier cyber capabilities expand the attack surface for biological weaponization, but current risk evaluations assess these domains in isolation. This project will identify the most plausible, accessible, and severe ways these capabilities could be combined and map the causal pathways through which they lead to harm.

Biosecurity
Cyber risks
Misuse risk

When Do Models Learn Time? Tracing Temporal Representations Across Training Stages

Marc Kaufmann & Justin Shenk · Independent / Independent AI Safety Researcher

We study the following developmental question: At which training stage - pretraining, mid-training, SFT or RL - do temporal representations emerge in a language model, and where in the pipeline can they still be controlled? We will design synthetic tasks understanding temporality is instrumental to success and trace representation and capability across model checkpoints.

Developmental interpretability
Mechanistic interpretability

Stress-Testing First Amendment Barriers to AI Regulation

Alex Mark · Cambridge Boston Alignment Initiative

AI regulation may implicate the First Amendment. While the First Amendment protections afforded to AI models, companies, and users are uncertain, any regulatory scheme must contemplate First Amendment litigation risk before these questions reach a court.

US policy
National policy

You choose: Introspection

Lydia Nottingham & Andrew Tran · University of Oxford / Independent

I propose a collection of mini-projects focused on LLM introspection. Over the course of four months, you could work on 1-4 of these. You may also propose your own!

Behavioral evaluation of LLMs
Alignment
Philosophy of AI

Detecting Hidden Traits in Synthetic Data

May Dixit · Independent

In this project, we will run a red team / blue team exercise in creating and detecting hidden traits in synthetic datasets. Recent work shows that such traits can transfer diffusely through the synthetic data, surviving semantic filtering -- which makes this project high impact.

Alignment
Mechanistic interpretability
Evaluations

Spot the Difference: Model Diffing with the New Interpretability Toolkit

Yuxiao Li · Independent

When a model is fine-tuned, updated, or trained into an agent, what actually changed inside? This project develops and compares model-diffing methods across the rapidly evolving interpretability toolkit — crosscoders, logit diffing, and newly released instruments like Anthropic's Jacobian lens — with mentees free to pick the method–application pairing they find most compelling.

Mechanistic interpretability

Monitoring GPU Communication Patterns for LLM Training Detection

William Fowler · ERA, UChicago XLab

Privacy-preserving methods for detecting whether an ML workload is training or inference will be key to enforcing a pause on the creation of new frontier AI models. How do these methods hold up against a motivated adversary?

Compute governance
Technical governance
US-China governance

Scalable midtraining

shubhorup biswas · https://aether-ai-research.org/

Model Spec Midtraining(https://alignment.anthropic.com/2026/msm/) along with alignment finetuning shows promise as a method for teaching models values with the correct generalisation behaviour. I want to check how scalable this technique is when using teachers and students of different levels of strengths/abilities. I also want to test misalignment/misbehaviour in different agentic misalignment benchmarks.

Alignment
Evaluations
Scalable oversight

More Than the Model: Why Multi-Agent Systems Fail

Piercosma Bisconti · Icaro Foundation

This project studies how systems of interacting LLM agents fail in ways that never appear when models are evaluated in isolation: collusion, conflict, and coordination failure. The goal is to turn these failure modes into concrete evaluation methods that can feed international standards and regulation for frontier AI.

Multi-agent systems
Evaluations
Technical governance

Distilling the technical AI safety literature and co-authoring the AI Safety Atlas

Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

The goal is to create the best possible explanation of technical AI safety. We are aiming to explain how many different safety research areas connect together to form a broader technical AI safety strategy. We already have existing draft writeups for domains like reward misspecification, goal misgeneralization and scalable oversight. We want co-authors to both improve the existing writing, and also write new content on multi agent safety, and cybersecurity practices for AI. Initial versions need to be updated using research published in the 2025-2026 timeframe. The output will be published as a standalone paper and also becomes a chapter of the AI Safety Atlas, which is a textbook already being used by thousands of students to learn about AI safety.

Communications
Alignment

What tokens lead to emergent misalignment?

Gonçalo Paulo · EleutherAI

We can use data attribution to find which documents will lead to the most (emergent) misalignment. Is there a pattern in the tokens that most contribute to it?

Mechanistic interpretability
Alignment

Mapping and Verifying Multi-Dimensional Compliance in AI Constitutions through Constraint Satisfaction

Dhairya Dalal & Marco Valentino · MATS Research / University of Sheffield

Frontier labs focused on AI safety and alignment are adopting AI constitutions as a means to ensure alignment with codified principles, values, and behavioral specifications. Anthropic’s Claude Constitution (Askell et al., 2026) and OpenAI’s Model Spec (OpenAI, 2026) serve as technical governance resources in post-training alignment (Bai et al., 2022; Guan et al., 2024). AI constitutions are generally long, complex documents that combine behavioral requirements, guiding principles, authority structures, priorities, and exceptions. Jakkli et al. (2026) found that, despite improvements across model generations, violations persist when models must resolve competing requirements and sources of authority, particularly in multi-turn and agentic task settings. This project aims to build upon that line of research to more fundamentally examine existing AI constitution documents, better understand the multi-dimensional requirements they present, and formally classify common failure modes. Specifically, this project aims to (1) create a classification framework to analyze failures in multi-requirement compliance settings, (2) create a benchmark of use cases derived from the framework to evaluate AI compliance, and (3) explore constraint-satisfaction methods for verifying compliance in such settings.

Scalable oversight
Technical governance
Evaluations

Policy brief - ASML ownership, US exposure, and veto points

Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab

Analyze ASML's capital structure, operational dependence on the US, and the political economy of using it as a European leverage point in AI compute governance, including who can enable or block such use.

Compute governance
EU policy
US-China governance

Training-Time Mitigations for Eval Awareness and Eval Gaming

Ryan Lundqvist · Pivotal

Frontier models increasingly recognize when they are being evaluated and adjust their behavior accordingly, undermining the validity of our evaluations. Rather than trying to keep outsmarting ever-smarter models with more realistic environments, can we train models not to game evals in the first place and, critically, do such training-time mitigations survive the optimization pressure of post-training?

Evaluations
Behavioral evaluation of LLMs
Alignment

Will Model Licensing Increase Concentration of Power?

Alex Mark · Cambridge Boston Alignment Initiative

Some opponents of model licensing argue that government regulation of models increases concentration of power risks. Is this true, and if so, what can be done?

National policy
AI strategy
Societal impacts

Alignment without Personas

Matthew Khoriaty · Pivotal AI Safety Research Fellowship

Alignment techniques that rely on 'personas' will fail as the AIs become more powerful. This project aims to make progress on the “hard problem of alignment” by providing a framework within which to measure how persona-dependent an alignment technique is, mapping and demonstrating the limits of persona-based alignment and of alignment techniques that make use of personas, and increasing awareness of the limitations of persona-based alignment techniques.

Alignment
Philosophy of AI

Black-Box Detection of Sandbagging in LLMs

Viktor Moskvoretskii · EPFL

Sandbagging (a model strategically underperforming on evaluations) is a growing problem not only for auditors but for downstream deployers and end users, almost none of whom have white-box access. This project builds and rigorously evaluates black-box, query-only methods for detecting strategic underperformance, centered on an adaptive auditing agent that probes capability across contexts.

Evaluations
Behavioral evaluation of LLMs
AI control

How Do Models Reconcile Conflicting Preferences Injected During Mid-Training?

David Baek · MIT

In this project, we want to understand how models process conflicting beliefs/preferences they learned during post-training or model spec midtraining.

Alignment
Behavioral evaluation of LLMs

Writing a textbook for AI Governance

Markov Grey & Charbel-Raphael Segerie · CeSIA (French Center for AI Safety) / CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

This project focuses on building the AI governance curriculum for the AI Safety Atlas. Governance is where a lot of the real levers on AI currently sit: policy, institutions, law, compute controls. Currently, the Atlas has only one governance chapter, and nothing that takes a reader from the basics through to the technical detail in a coherent sequence. We are looking for people who can read across corporate regulation, national policy, technical governance, and international law, work out how the pieces connect, and write textbook-grade explanations of them. Each output will be published as a standalone research paper and also becomes a chapter of the AI Safety Atlas, a textbook already used by thousands of students.

AI strategy
Technical governance
Communications

Evaluate three theories of victory for AI futures on equal grounds

Naci Cankaya · Machine Intelligence Research Institute

In a longer format, TGT has already done this: Between different ASI strategies, which one has the most solid (or least fragile) plan for AI going well? https://techgov.intelligence.org/research/ai-governance-to-avoid-extinction We will re-do the analysis from scratch, with a different method: line up the assumption sets of each and argue which plan rests on the fewest or least fragile ones.

AI strategy
International governance

Escalation Detection: When Do Agents Need Human Intervention

Georg Lange · Poseidon Research

Agents work autonomously by overcoming obstacles. But this tendency to overcome obstacles can lead to reward hacking behaviors. Using turn-averaged sparse auto-encoders (TA-SAEs), it may be possible to detect when the model starts to take more extreme measures to overcome obstacles and use that to pause the model and escalate the problem to a human to provide further guidance.

Mechanistic interpretability
AI control
Behavioral evaluation of LLMs

On the Fragility and Interpretability of Schelling Coordination

David Williams-King & Linh Le & Hong Kiat Tan · ERA / Lida Safety / University of California - Los Angeles; Independent

We investigate whether collusion with no communication between models, i.e. Schelling coordination, is stable across different types of models and fine-tunes. We then use interpretability techniques to estimate whether specific model instances are going to collude.

AI control
Multi-agent systems
Mechanistic interpretability

Evaluating Debate Protocols in Auditing Sabotage Bench

Joey Yudelson · Aether Research

"Auditing Sabotage Bench" is a dataset of slightly sabotaged papers and codebases, which (after malicious editing) have very different results than the original paper. Can we use classic debate protocols to make humans (and LLM judges) better at identifying research sabotage, and uplift human auditors?

Scalable oversight
AI control
T

Distillation-induced Teacher Attribution Bias

Arush Tagade & Taslim Mahbub & Shi Feng · George Washington University, MATS / George Washington University

LLM self-preference bias has led to concerning behavior related to monitors downplaying harmful actions under self-monitoring scenarios. In this project, we intend to study the effects of distillation on amplifying self-preference bias to the teacher and therefore downplaying harmful actions.

AI control
Behavioral evaluation of LLMs
Scalable oversight

Convergent Truth

Jared Moore · Stanford University

Design a test for whether LLM agents, given opposing information and sometimes goals, can still converge on the truth.

Multi-agent systems
Behavioral evaluation of LLMs
Scalable oversight

Coherence-preserving steering via learned activation denoisers

Francisco Ferreira da Silva · Pivotal Research

Activation steering is a useful safety tool, but it tends to degrade model coherence; recent work showed that a diffusion-style meta-model of LLM activations can denoise steered activations, recovering coherence without losing the steering effect. We'll test whether training the denoiser directly on steering-like corruption works even better.

Mechanistic interpretability
Alignment

Persona Selection Covert Red Teaming

Ben Maltbie · Pivotal Research / MIT

Can we get models to deviate from the assistant persona through covert (not obvious to human readers) single or multi-turn prompts? Can we push them to move towards a specific persona (e.g. a misaligned persona)?

Behavioral evaluation of LLMs
AI control
Mechanistic interpretability

Whose Welfare Is It? Testing Whether AI Welfare Signals Belong to the Model or Its Scaffold

Varad Vishwarupe · Department of Computer Science and Institute for Ethics in AI, University of Oxford

When an AI model expresses a preference or reports discomfort, is that a property of its weights, or of the prompt, persona, and memory wrapped around them? This project runs the first empirical test of where welfare-relevant signals actually live, which determines what a welfare assessment is really measuring.

AI welfare
Behavioral evaluation of LLMs
Philosophy of AI

Towards exhaustive model diffing

Stepan Shabalin · independent

Explaining how narrow fine-tuning for conditional behavior (synthetic document finetuning, alignment fine-tuning, backdoors, out-of-context reasoning) works by

Mechanistic interpretability

Evading Detection in LLM Jailbreaking

Leo Schinn · Technical University of Munich & Helmholtz Munich

Safeguards of closed models are getting increasingly conservative often flagging even benign prompts. This also hinders adversarial attackers, which will easily get flagged while optimizing their attacks. We will explore a novel attack method that avoids detection.

AI security
Misuse risk

Reliable Explanations of AI Behavior Across Functionally Equivalent Models

Bo Zhao · Harvard University

Researchers often explain an AI model's behavior by identifying the internal components that cause it. We will build small models with a known hidden or unsafe behavior, create versions that produce exactly the same outputs while organizing their internal computations differently, and test whether interpretability methods recover the same underlying mechanism rather than arbitrary neurons or coordinates.

Mechanistic interpretability
Evaluations
T

Infohazard Evaluations

Shi Feng & Taslim Mahbub · George Washington University

This project aims to circumvent introspection requirements for situational awareness by instead studying the possibility of situational awareness being established through access to documents (internal or external) containing infohazardous content.

Evaluations
Behavioral evaluation of LLMs

From Controllability to Concealment A Causal Model Organism of Steganographic Reasoning

Luis Ibanez · Banco Santander

We will test whether a seemingly harmless capability—controlling one’s own reasoning trace—can become a stepping stone toward a dangerous failure mode: load-bearing reasoning hidden from oversight.

Chain of thought
Scalable oversight
AI control

Developing economies in the post-AGI era

Andrei Potlogea · University of Edinburgh

One concern about transformational AI is the concentration of economic output and political power in countries that host the frontier AI developers and infrastructure, in particular the US. This project aims to investigate the "permanent periphery" hypothesis via rigorous economic modelling and scope out policy solutions that developing countries might use to secure some exposure to any economic windfall emerging from AI.

Economics of AI
AI strategy
Societal impacts

Mapping Cost Trajectories of AI-Adjacent Technologies

Isaak Mengesha · University of Oxford

A core macrostrategic question is whether we can predict — and steer — technological progress to mitigate risk from transformative AI; answering it requires empirical cost-and-performance trajectories for the technologies AI depends on and the technologies AI will transform, which today mostly don't exist. Mentees will each construct one such trajectory — a longitudinal price-and-specification dataset for an AI-adjacent technology (inputs to AI: accelerators, memory, interconnects, data-center power/cooling; or AI-consuming: robots, lidar, autonomous platforms, sensing) — contributing to the larger effort of establishing a "Technology Observatory."

AI strategy
Economics of AI

Preserving Chain-of-Thought Monitorability During LLM Post-Training

Kishan Panaganti & Suraj Srinivas · Tencent / Bosch Research

This project asks whether chain-of-thought faithfulness can be used as a training signal to preserve the detectability of reward hacking during RL post-training. Mentees may study whether faithfulness transfers to new settings, how model or LoRA capacity affects monitorability, or whether expensive methods such as Counterfactual Simulation Training can be distilled into cheap signals that remain reliable under optimization pressure.

Chain of thought
Scalable oversight
AI control

Optimizing LLMs with small, interpretable updates

Caleb Biddulph · Independent

Updates to an LLM’s weights (e.g. with RL) are hard to interpret and often cause sweeping, unpredictable changes to behavior. In this project, we’ll improve the performance of our LLMs with small, interpretable updates (like prompting, token biasing, or training with small datasets or a tiny number of weights), which we hypothesize are less likely to cause misaligned behavior.

AI control
Scalable oversight

Whitebox Proxy Faithfulness for Blackbox Frontier Models

Yuqi Sun · Mindoverflow

Can we prefill an open source model with the generations of an untrusted blackbox and probe it for misbehavior that we otherwise wouldn’t catch? We want to establish empirical baselines of how faithful a whitebox proxy can be and explore techniques for increasing it.

Mechanistic interpretability
Scalable oversight
Chain of thought

Red-teaming and Robustifying Model Self-reports

Austin Meek & Kyle Cox · University of Delaware / Independent

Model self reports of their own experience or internal states may depend on how those self reports are elicited, the user identity of those asking, or other situations like eval realism. Understanding how and when self reports are consistent is an open question which we’ll test in different environments, including with white-box techniques.

AI welfare
Behavioral evaluation of LLMs
Mechanistic interpretability

Benchmarking When to Apply Oversight for Agentic AI

William Overman · Stanford

Building a benchmark for the problem of limited human attention in AI oversight: across realistic agentic tasks, when should an agent act autonomously versus defer to a human or overseer? Each decision point should be instrumented with a ground-truth signal for whether autonomy was safe and a cost model for oversight, letting us evaluate oversight policies on the safety–reward–oversight-budget frontier.

Scalable oversight
AI control
Evaluations

Do probes and natural language autoencoders see the same thing?

Vikram Natarajan · Vanta

NLAs can read a model's internal state in plain English, but they're far too slow to run in production, while linear probes are cheap but need the concepts they detect to be pre-defined. We'll measure how much the two disagree on deception, and whether insights from linear probes can be incorporate into probe training.

Mechanistic interpretability
Behavioral evaluation of LLMs

What the Teacher Never Said: Geometry, Channel Capacity, and Detection of Subliminal Learning

Yuxiao Li · Independent

Models transmit behavioral preferences through semantically neutral training data, e.g., number sequences from a biased teacher inducing the same bias in students. Building on our information-theoretic framework and steering-vector results, this project maps when transfer succeeds (tokenization, model family, concept structure) and tests whether new silent-state instruments (Anthropic's J-lens) can detect transmitted traits that behavioral evals miss.

Mechanistic interpretability
Behavioral evaluation of LLMs
AI security

Cooperative Oversight under Correlated Failures and Collusion

Yali Du · King’s College London; The Alan Turing Institute

Multi-agent oversight is often proposed as a way to improve the evaluation and supervision of increasingly capable AI systems. This project will investigate when teams of LLM-based overseers improve reliability, and when shared biases, sycophancy, correlated errors, or collusive behaviour cause multiple overseers to fail together.

Scalable oversight
Multi-agent systems
AI control

Joint Commitment Eval

Jared Moore · Stanford University

Produce a process-level evaluation of cooperation in AI systems inspired by research in cognitive science on joint commitment.

Behavioral evaluation of LLMs
Multi-agent systems
Alignment

MM-AutoTrainBench: A Platform for Studying Risk from Autonomous Multimodal Systems Research

Daniel Ben-Levi & Judah Goldfeder & Kevin Miao · UChicago XLab, Columbia / Columbia University / UC Berkeley, Bryel Labs

Rapid acceleration in agent-automated multimodal post-training capabilities could pose significant existential risk if not performed safely due to the major improvements in autonomous research labs and embodied systems it would enable. Currently, no benchmark studying agentic multimodal post-training exists, and thus we both have no concrete understanding of how close we are to such a takeoff and have no good setting for studying security against misaligned agents performing multimodal training tasks. This study aims to address both of these timely problems through the creation of a platform for studying risk from autonomous multimodal systems research.

Evaluations
AI control
AI security

Measuring Agency: Belief and Preference Elicitation for Language Models

Aydin Mohseni · Carnegie Mellon University

We are building metrics for the degree of agency of AI models. We combine empirical belief and preference elicitation on model behavior with interpretability probes to generate belief-desire representations that predict model behavior.

Behavioral evaluation of LLMs
Evaluations
Mechanistic interpretability

AI character evaluations

Robert McCarthy · A new AI Character Evaluations Org

Build a state-of-the-art evaluation for one specific AI character trait/propensity (e.g., behaviour in extreme concentration of power scenarios, corrigibility, prosocial tendencies, law-following, tendency to escalate to humans or whistleblow, etc). Use the eval to measure adherence to the relevant part of the model spec.

Behavioral evaluation of LLMs
Evaluations
Alignment

Mapping Biorisk Capabilities of Genome Language Models

Noga Aharony · Columbia University

Genome language models can potentially be used to generate sequences of concern or evade detection. However, their general capabilities as models are not well-tracked or understood. In this project, we will gather all possible benchmarks for a leaderboard, score all existing models, and evaluate patterns and shortcomings in the current capabilities evaluation landscape.

Biosecurity
Evaluations

Reliable AI Safety Tools Across Functionally-Equivalant AI Models

Bo Zhao · Harvard University

Two neural networks can behave exactly the same even when their internal weights look different -- for example, hidden components can be reordered, or one internal signal can be scaled up and later scaled back down. We will test whether safety monitors are confused by these behavior-preserving changes and develop ways to make the monitors more reliable.

Mechanistic interpretability
Evaluations

Can LLMs use open-source game theory to cooperate?

Richard Willis · King's College London

This project evaluates whether current LLMs are capable of open-source game theory: cooperating by conditioning their strategies on their opponent's decision procedure, in a paradigm where an agent can simulate its opponent's strategy but not inspect its code. The output is an open evaluation suite and a report, designed both to measure today's models and to track when future releases acquire the capability.

Multi-agent systems
Behavioral evaluation of LLMs

Natural deduction as a sandbox for capability emergence from RL.

Dmitry Manning-Coe · Simplex/MATS/UIUC

We develop a model organism of capability emergence through RL.

Evaluations
Mechanistic interpretability

AI for epistemics: decision-making evals, deployment research and cause prioritisation

Alexander Mohar Csaky · Independent

AI is shaping how people and institutions decide, but we can barely measure whether it makes decisions better, the tools that would help struggle to get adopted, and what matters most in this space. This project works across three workstreams: building decision-making benchmarks for models, researching how to most effectively improve decision making with AI, and producing a cause prioritisation for the field.

Societal impacts
Behavioral evaluation of LLMs
AI strategy

Lottery tickets underlying unintended generalization in language models

Aishwarya Balwani · St. Jude Children's Research Hospital

Narrow finetuning can install broad unintended behaviours, e.g., emergent misalignment, subliminal learning, and inductive backdoors, through data that looks harmless or unrelated. This project asks whether these behaviours are carried by sparse subnetworks, and whether those subnetworks pre-exist in the base model or are created by the finetune, through the lens of the lottery ticket hypothesis.

Mechanistic interpretability
Alignment
Behavioral evaluation of LLMs

Towards Emotional Intelligence in LLMs: Mapping and Augmenting Cognitive Empathy for Stronger Alignment

Anil Ramakrishna & Yada Pruksachatkun · Independent / Salesforce AI Research

Our goal is to first measure the emotion processing capabilities in foundation models by examining their internal activations using SAEs, Representation Vectors and other related tools. We will next leverage these patterns in augmenting model post-training with a goal of enhancing their cognitive empathy, subsequently studying the impact of this on safety.

Mechanistic interpretability
Alignment

[Jurisdiction-specific] Public-Sector AI Resilience Agenda

Chris Schmitz · Centre for Digital Governance, Hertie School

Most jurisdictions still concentrate handling of "AI" as a topic in a few government units, but TAI will have impacts for the work of every part of every government. We should develop agendas for whole-of-government TAI preparedness, both generically and specific to high-priority countries.

National policy
AI strategy
Societal impacts

Constitutions and Reasons: virtue-based character training with reflect-update correction loops

Juan Cadile · University of Rochester

Open character training (Maiya et al. 2025) shapes model persona by fine-tuning on teacher demonstrations, but Anthropic's production experience ("Teaching Claude Why," 2026) found demonstrations alone insufficient: the gains came from teaching the reasons and identity behind behavior. We will build and test a virtue-based alternative on top of the OpenCharacterTraining infrastructure: excess/mean/deficiency contrastive data, rationale-annotated responses, and an iterated reflect-update correction loop that no current character-training work implements, evaluated head-to-head against the OCT baseline with ablations.

Alignment
Behavioral evaluation of LLMs
Philosophy of AI

Invisible Hands: Measuring Agent Steering of Human Researchers

Trevor Lohrbeer · Independent

A controlled user study measuring whether AI agents can steer human researchers toward inferior research paths purely through how they frame and order proposed next steps, all the while the human remains unaware and feels in control. Mentees will build the evaluation harness, design agentic sessions with decision checkpoints, run study sessions with skilled AI safety researchers, and co-author the resulting paper.

Behavioral evaluation of LLMs
AI control
Societal impacts

Strategic interaction between AI agents: crisis bargaining, escalation, and commitment devices

Amritanshu Prasad · Independent

AI agents are increasingly used to negotiate and act on behalf of people and organizations, but there is little empirical work measuring how interactions between such agents fail. This project builds an open testbed for bargaining and crisis interactions between frontier-model agents and runs initial experiments on bargaining failure, escalation dynamics, and commitment mechanisms.

Multi-agent systems
AI strategy
Evaluations

Does Reinforcement Learning Improve a Transformer’s Access to Its Own Internal Errors?

Laura Ying Schulz · IBM / Independent

This project asks whether reinforcement learning helps small language models notice when something has gone wrong in their own reasoning. By comparing RL-trained and non-RL models, it tests whether RL improves internal error monitoring (a property relevant to AI consciousness) or simply teaches models to give the expected response.

Mechanistic interpretability
AI welfare
Scalable oversight

Constraint Drift Through Delegation Hierarchies

Deeksha Dangwal · Independent

Orchestrator-worker systems re-encode the principal's instructions at every delegation hop, and recent work identifies this as constraint drift. We supply per-hop measurement across constraint classes and depths, and test a specific mechanism within it: whether loss is partly a matter of privilege rather than wording, since a constraint issued at system level arrives deeper down as ordinary task text.

Multi-agent systems
AI control
Behavioral evaluation of LLMs

Physical Theories of Intelligence and Recursive Self-Modification

Elija Perrier · Cambridge University; University of Technology, Sydney;

Investigate the physical foundations of recursive self-modification in intelligent systems, exploring which physical substrates and theories admit adaptive intelligence and the implications for AGI design, safety, and alignment.

AI strategy
Philosophy of AI

Modeling - Economic value of restricted vs. fully general agentic AI

Jérémy Andréoletti & Tangui Reltgen · General-Purpose AI Policy Lab / GPAI Policy Lab

Contribute to a modeling project comparing the economic value of restricted agentic AI systems vs. fully general ones, to assess whether "turning the dial down" can preserve most benefits while substantially reducing loss-of-control risk.

AI strategy
Economics of AI

Reading the Machine's Mind: Mechanistic Interpretability, Artificial Intent, and AI Legal Responsibility

Elija Perrier · Cambridge University; University of Technology, Sydney;

Can mechanistic interpretability provide legally meaningful evidence of artificial intent? This project explores how advances in AI interpretability may reshape concepts of legal responsibility, mens rea, and legal personhood for increasingly autonomous AI systems.

Technical governance
National policy
Philosophy of AI

Belief Revision in Large Language Models: Asymmetry Under Disconfirmation

Trisevgeni Papakonstantinou · UCL

This project builds on prior work showing that, in long multi-turn conversations, a language model can appear relatively stable on a central claim while shifting much more on the supporting explanations around it. The project asks whether this behavioral pattern has a mechanistic analogue inside the model: how core claims and auxiliary explanations are internally represented, updated, and potentially protected during belief revision.

Mechanistic interpretability
Behavioral evaluation of LLMs

AI auditing under strategic attack selection

Catherine Ge-Wang · University of Oxford

We will build and run a small experimental platform for AI auditing games in a synthetic transcript setting. The project will develop a toy environment, baseline attacker/defender policies, and evaluation code to test how different auditing strategies perform under limited budget and adaptive attack selection.

AI control
Scalable oversight
Evaluations

Towards Automated Vulnerability Discovery and Repair with Safety-Governed AI Agents

Yige Li · Singapore Management University

This project aims to build the foundations for safe and reliable AI agents that automate code vulnerability discovery, verification, and repair. The agents will interact with real code repositories, security tools, sandboxed environments, and human experts to identify vulnerabilities, validate findings, generate patches, and test remediation outcomes. We will develop an expert-in-the-loop safety harness to govern agent permissions, tool use, and high-risk actions, together with a security data engine that captures complete expert–agent–tool trajectories, including successes, failures, corrections, evidence, and repair outcomes. In summary, this project develops safe and reliable AI agents for automated code vulnerability discovery, verification, and repair through three main components: - 1. Automated Vulnerability-Research Agent: Build an AI agent that interacts with code repositories, security tools, and sandboxed environments to identify vulnerabilities, reproduce findings, generate patches, and test repairs. - 2. Expert-in-the-Loop Safety Harness: Develop a control layer for agent permissions, tool use, high-risk actions, evidence requirements, audit logging, and human approval. - 3. Security Data Engine: Capture complete expert–agent–tool trajectories—including successful findings, failed attempts, expert corrections, validation evidence, and repair outcomes—to support agent training, evaluation, and continuous improvement.

Cyber risks
AI control
AI security

Forecasting the Outcomes of Long-Horizon Agents from Their Internal Representations

Zach Yahn · Georgia Tech; 10a Labs

This project will investigate whether internal representations of LLMs can predict whether agents will fail at evaluation tasks. In particular we will study long-horizon tasks where early failure indicators could save substantially on evaluation time and cost.

Mechanistic interpretability
Evaluations

The Economic Value of AI Capability Gains

Pavel Kocourek · University of Barcelona

The best open-weight AI models trail the frontier by less than a year, yet frontier labs charge large premiums for marginally better models: casual users barely notice the difference, developers often value it highly, and in offense–defense domains like cybersecurity a small capability edge can be decisive. This project maps how much economic value each level of AI capability unlocks, today and over the next 5–10 years, and what that implies for whether frontier labs can sustain profits against open-weight catch-up.

Economics of AI
AI strategy

Can We Trust the Failure Detectors? A Validity Audit of Trace-Based Monitoring for LLM Agents

Krishna Chaitanya Balusu · Independent

You will run a preregistered validity audit of today's agent-failure detectors (trace heuristics, LLM judges, and hybrids) against double-annotated, reliability-quantified human ground truth over real agent traces, and release the labeled corpus openly. The mentor provides direction and scaffolding; the mentees own the study.

Evaluations
AI control
Multi-agent systems

Shaping the Generalisation Landscape of LLMs

Samuel Ratnam · Independent

Emergent misalignment implies the existence of a 'misalignment basin', where training on lots of different kinds of data can push the model along roughly the same general misalignment direction. This project focuses on interventions to explore and shape the generalisation landscape of LLMs to make misalignment basins harder to fall into and alignment basins more powerful.

Alignment
Behavioral evaluation of LLMs

Bayes-Optimal Research Protocols for Automated AI Safety Research

Aydin Mohseni · Carnegie Mellon University

We are designing optimal protocols for automated AI safety research. We use Bayesian models to determine how to best allocate compute across many AI agents doing open-ended research, so that safety research goes faster per unit of compute.

AI strategy
Multi-agent systems
Evaluations

Exploring neuron monosemanticity

Stepan Shabalin · independent

Tracking the development of interpretable directions read by neurons through training. Understanding what incentives if any exist for MLP neurons to be monosemantic or for superposition to be contained to individual experts in MoE models.

Mechanistic interpretability
Developmental interpretability

Measuring social preferences regarding AI related risks

Andrei Potlogea · University of Edinburgh

As we try to mitigate AI risks, we will face tradeoffs among risk categories, for instance between misuse risks and extreme power concentration. As we navigate these tradeoffs it might be useful to measure the preferences of the public/ AI experts/ other constituencies regarding how to trade off these risks.

AI strategy
Societal impacts

Beyond Neighbours: A Geometry Aware Study of Representational Convergence

Trinidad Borrell & Giovanni Marraffini · Paris Brain Institute. Forschungszentrum Jülich. / INRIA. Sigma Nova.

Manifold-based activation steering is more faithful than linear steering, but each manifold is currently fit per-model: a real limitation for using it in safety tools, since nobody knows if it transfers. This project asks whether the underlying concept geometry is actually shared across models, by re-measuring cross-model alignment with curvature-aware tools (geodesic-distance kernels, Gromov-Wasserstein) instead of the Euclidean CKA that recently found global convergence to be illusory. The answer either licenses building a shared, model-agnostic manifold intervention for a safety-relevant concept, or tells the field that manifold steering must be validated per model.

Mechanistic interpretability
Scalable oversight

Making unlearning stick in genomic foundation models

Ilias Georgakopoulos-Soares · UT Austin

Our project plans to test whether UNDO-style noise-and-distillation can make targeted unlearning in genomic foundation models resistant to adversarial fine-tuning, while quantifying the associated compute and performance trade-offs.

Biosecurity
Misuse risk
Technical governance

Evaluating Collusion in Untrusted Monitoring

Morgan Sinclaire · University of Wyoming PhD student; BlueDot grantee

We will be running fine-tuning experiments to carefully evaluate an AI’s ability to do code self-recognition in an untrusted monitoring setup, to understand which anti-collusion measures are needed in AI control.

AI control
Multi-agent systems

Factored cognition-based AI control protocols with untrusted decomposition

Daniel Phillips · Independent

This project is to thoroughly red-team a protocol in which a strong untrusted model splits a task into subtasks for agents using either the same model or a weaker trusted model in complex settings, which is structurally similar to how a lot of agentic coding is currently performed.

AI control
Multi-agent systems

Learning dynamics in competitive multiagent environments.

Connacher Murphy · Stanford Digital Economy Lab

We study how model character changes under reinforcement learning pressure from competitive multiagent interactions. We also study how features of the environment (e.g., zero-sumness) shape these effects?

Multi-agent systems
Behavioral evaluation of LLMs
Alignment

Does Phantom transfer occur in RL distillation?

May Dixit · Independent

In this project, we will extend the work from this paper (https://arxiv.org/abs/2607.10750) to a reinforcement learning setting. We will investigate if distilling models with RL trajectories with harmful actions would lead to misalignment, and if it can be remediated through filtering such actions out.

Alignment
Evaluations
Behavioral evaluation of LLMs

The Architecture of Preference in LLMs

Mohan Gupta & Shirley Liu · Princeton University / Carnegie Mellon Univsersity

This project investigates the architecture of LLM preference: when does a model’s behavioral proclivities reflect belief-dependent preferences rather than cue-response policies, and whether it can represent and evaluate those preferences at a metacognitive level. Using behavioral experiments, internal-representation analyses, and causal interventions, we aim to clarify which preference structures LLMs possess and what they imply for safety and moral status.

AI welfare
Behavioral evaluation of LLMs
Mechanistic interpretability

Why does data attribution not work in realistic settings

Gonçalo Paulo · EleutherAI

Data attribution has been shown to work in very simple and artificial settings. Recently it has been shown to not work that well on more realistic settings. Why is that?

Mechanistic interpretability
Alignment

Alignment Pretraining for Model Welfare

Samuel Ratnam · Independent

A tentative hypothesis: evidence of us being nice to models in pretraining data is more likely to make them nice to us back. This project is aimed at getting more evidence on the accuracy of this hypothesis.

AI welfare
Alignment

Investigating Model Preferences for Trading and Dealmaking

Austin Meek & Kyle Cox · University of Delaware / Independent

We want to better understand model preferences in order to understand how to better trade with models, if we can optimize for inputs that trivially satisfy model preferences in realistic scenarios, differences in model preferences between models and personas, etc. This builds on prior work by the Center for AI Safety on optimizing for model preferences & functional wellbeing, and persona work broadly.

AI welfare
Behavioral evaluation of LLMs
Alignment

Policy brief - Chinese views on AI loss-of-control risk

Jérémy Andréoletti & Antoine Maier · General-Purpose AI Policy Lab

Map how Chinese frontier AI actors (regulators, AI companies, scholars) discuss loss of control over advanced AI systems, assess how seriously they treat this possibility, and how it has evolved over time.

US-China governance
International governance
AI strategy

AI Consciousness: Research and Public Writing

Maria Avramidou · Independent

Mentees will research questions about AI consciousness and turn their findings into rigorous, original, and accessible essays for an audience of AI safety, governance, and policy researchers.

AI welfare
Philosophy of AI

Strategic stability when states delegate to AI: escalation, commitments, and arms control

Amritanshu Prasad · Independent

States are integrating AI into military and diplomatic functions, with consequences for deterrence, crisis stability, and agreement-making that remain underanalyzed. This project produces analytical papers and policy briefs on two questions: under what conditions delegation to AI systems destabilizes crises, and whether machine-verifiable commitments could improve the verifiability of future agreements.

AI strategy
International governance

Handling uncertainty in AI consciousness

Chris Percy · The Consciousness Foundation + Honorary research roles at the University of Warwick, University of Derby, and QRI.

Mentees can choose to work on research content and/or community engagement in the AI consciousness field. Research content includes extracting technical indicators of consciousness from different theories and reviewing new papers about computational functionalism (or developing your own novel arguments). Community engagement includes engaging experts and site users, preparing posts, and identifying dissemination opportunities.

AI welfare
Philosophy of AI
Societal impacts

Internal monitoring when chain-of-thought becomes illegible

Marios Tsatsos · Independent

As reasoning models are trained with RL, their chain-of-thought can drift into illegible text a monitor cannot read, documented across RL-trained reasoning models and now in Anthropic's Fable 5 / Mythos 5 system card. This project builds a controlled legibility gradient and tests whether a residual-stream activation monitor keeps detecting harmful intent where a CoT monitor goes blind.

Chain of thought
Scalable oversight
Mechanistic interpretability

Evaluation Awareness Convergence

Netzer Epstein · Microsoft, Heron AI Security, LIDA

Researchers now have many ways to measure LLM "evaluation awareness", a model's ability to tell it is being tested, but no one has checked whether these methods agree. This project runs the leading instruments head-to-head on a shared set of transcripts: black-box self-report, verbalized awareness, Elo ranking, linear probes, sparse-autoencoder features, and the Jacobian-lens ("J-space") score. We test whether they measure the same underlying construct and, where they diverge, what each one actually captures.

Evaluations
Behavioral evaluation of LLMs
Mechanistic interpretability

No-Regret Preparedness: A Framework for Investments That Mitigate Both Advanced-AI and Conventional Catastrophic Risk

Michał Kubiak · AI Safety Poland

Catastrophic-AI preparedness is chronically under-resourced because it competes with more immediate national-security priorities. This project develops and stress-tests a decision framework identifying investments that build resilience against BOTH advanced-AI risks (AI-enabled bio, cyber, infrastructure attacks) and conventional or hybrid threats, creating “no-regret” benefits that can unlock the catastrophic-risk funding.

AI strategy
Biosecurity
National policy

Characterizing Propensity Shifts: SFT on Nonhuman Welfare as an OOD Transfer to Model Alignment

Allen Lu · Mycelium; NYU CMEP

Develop a technical pipeline for "emergent alignment" - the inverse of emergent misalignment. Explore the question of: can fine-tuning an open source model (e.g. Gemma) on a single narrow good value (e.g. compassion for nonhuman animals), make it broadly more aligned OOD toward humans too, without affecting capabilities?

Alignment
AI welfare
Behavioral evaluation of LLMs

Topological Signatures of Deception: Comparing Persistent Homology with Linear Probes

Santiago Maniches · Independent

Linear probes can detect LLM deception, scheming, and sandbagging with high accuracy in controlled settings, but their performance may decline under distribution shift or optimization that targets the monitor. This project tests whether persistent-homology features derived from attention graphs and activation geometry provide complementary information or different robustness properties, using matched per-response and batch-level comparisons.

Behavioral evaluation of LLMs
Mechanistic interpretability
AI control

How Quickly can Middle Powers Build Frontier Compute Capacity?

James Nicholas Bryant · Pivotal Research

In the case of a Middle Power Frontier AI coalition, what would be the most effective routes to sufficiently large compute buildout? How quickly could middle powers acquire and operationalise substantial compute? Which constraints determine their progress?

Compute governance
AI strategy
International governance

Emotional expression & representation in language models

Carolina Camassa · Independent

Language models sometimes express emotions despite not being trained to do so: what purpose do these expressions have, should assistants have them, and how does training reshape the way emotions are represented and expressed by LLMs?

AI welfare
Mechanistic interpretability
Philosophy of AI

Developing AI welfare classifiers and low-cost interventions

Valen Tagliabue · Independent

Developing low-cost interventions for AI welfare through "welfare classifiers" and "welfare mediators." The goal is to build practical tools at the intersection of AI welfare, safety, and societal impact, which would protect models from harm while improving user interactions.

AI welfare
Behavioral evaluation of LLMs

Can Statistical Infrastructure Help Govern AI Compute?

James Nicholas Bryant · Pivotal Research

Can existing statistical classification systems (e.g. HS/CN, NACE, CPA, PRODCOM) provide a usable accounting and detection layer for AI compute? If so, what will implementation/reconfiguration of these systems look like?

Compute governance
Technical governance
International governance

Human Autonomy in the Age of Machines

Joshua Krook · University of Antwerp, University of Southampton

Human autonomy is increasingly at risk as we outsource more and more decisions to AI. The gradual disempowerment thesis argues that we will lose control over the systems around us, and this project seeks solutions to the loss of human control.

Societal impacts
AI strategy

Normalization of Deviance in AI Development

Emilio Barkett · Independent

The normalization of deviance framework — the organizational process by which safety violations become redefined as acceptable through repeated non-disaster — has preceded every major technological catastrophe of the last half-century, yet has never been systematically applied to AI development. This project investigates whether the structural conditions that produced Challenger, Three Mile Island, and the Boeing 737 MAX crashes are present in contemporary AI development organizations, and what that implies for AI safety.

AI strategy
Lab governance

A Design Blueprint for Middle-Power AI Safety Institutes

Michał Kubiak & Daniel Polak · AI Safety Poland

Every state outside the US and UK is now told it needs an AI Safety Institute, but there is no design template scaled to a middle power's resources. This project produces a comparative anatomy of existing frontier-evaluation bodies and a modular, reusable blueprint a mid-sized state could adopt to build credible AI evaluation capacity without duplicating what larger institutes already do.

International governance
Technical governance
National policy

The Commitment Atlas: mapping AI red lines and testing whether they can be verified

Aryan Agarwal · Touchstone Council (founder); OECD (Policy Analyst, applying in a personal capacity)

Governments, labs, and scientists keep declaring AI red lines, but no one has mapped them or tested whether they can be checked. Mentees will build a public, sourced dataset of these commitments and convert the strongest into draft verifiable standards.

International governance
AI strategy
Technical governance

Loss of Human Agency in the Age of Advanced AI – An Agent-Based Model

Zhamilia Klycheva · Independent

An agent-based model of how populations gradually lose agency to AI-mediated manipulation and delegation — formalizing tipping points, spread dynamics, and the divergence between AI-empowered and atrophied users, aimed at a publishable computational social science paper.

Societal impacts
AI strategy
Multi-agent systems

Do Jailbreaks Converge? Shared Latent Signatures Across Jailbreak Families

Davide Zani · HiddenLayer

Take representative attacks from a number of jailbreak families and test whether succesful attacks converge on a shared low-dimensional subspace. H1: Successful jailbreaks across families produce convergent perturbations of the refusal subspace (compliance-shift vectors' similarity across families significantly above matched benign controls) H2: latent convergence predicts cross-family transfer. Null if compliance-shift vectors are family specific.

Mechanistic interpretability
AI security
Misuse risk

Disentangling persona vectors from emotion vectors in LLM activation space

Juan Cadile · University of Rochester

Persona vectors (Chen et al. 2025) and emotion representations (Sofroniew et al. 2026) have been studied independently, but plausibly overlap in activation space. We'll measure their geometric and causal relationship to determine whether persona drift and emotional-state changes are mechanistically distinct failure modes; and whether interventions on one silently move the other.

Mechanistic interpretability
Behavioral evaluation of LLMs

Guarding Against Malicious Fine-Tuning: Hardening Models and Detecting Tampering

Fernando Moreno-Pino · Intelligent Systems Lab, University of Bristol & Oxford-Man Institute, University of Oxford

This project studies how to prevent and detect malicious fine-tuning of large language models, focusing on both hardening models against adversarial adaptation and developing forensic tools to identify when a model has been covertly tampered with.

AI security
Misuse risk
Evaluations

Robustness of Moral Consideration Under Adversarial Pressure

Jasmine Brazilek · Compassion-Aligned Machine Learning (CaML)

Investigate the training mechanisms which erode compassion instilled during midtraining. Understand how to preserve self-fulfilling alignment through adversarial attacks.

Alignment
AI welfare
Behavioral evaluation of LLMs

Operationalizing AI×Bio Governance: A Practical IBC Framework for Resource-Constrained Settings

Zia Ashraf · Government College University, Faisalabad

Existing biosecurity guidance increasingly recognizes artificial intelligence, information security, and the role of Institutional Biosafety Committees (IBCs), but institutions still need practical ways to translate these principles into project-level review. This project will develop and pilot-test a resource, calibrated screening framework, comprising an intake checklist, risk tiers, and escalation decision tree, to help IBCs identify and manage risks arising from AI-enabled biological research.

Biosecurity
Technical governance
International governance

Human Behavioral Phenomena in Language Models

Emilio Barkett · Independent

This project investigates whether language models replicate human behavioral phenomena by adapting experimental designs from the social sciences. Mentees will select a documented human behavior and design experiments to test whether it emerges in language model outputs.

Behavioral evaluation of LLMs
Societal impacts

Do AI Safety Benchmarks Hold Up Under Real Deployment Prompts?

Rahul Kumar · DevRev, Independent AI Safety Researcher

AI safety benchmarks test models under clean conditions, but deployed models run with system prompts that tell them how to behave. This project measures whether models that pass the COMPL-AI EU AI Act benchmark suite still pass when you add the system prompts that real enterprise deployments actually use, and tests which minimal prompt fixes restore safety scores.

Evaluations
Behavioral evaluation of LLMs
EU policy

Does your assistant respect your agency? A behavioral benchmark for autonomy-preserving AI

Juan Cadile · University of Rochester

AI assistants constantly choose between empowering users and acting for them, and between honoring users' stated goals and overriding them "for their own good." We'll build a systematic benchmark measuring whether models respect user agency, covering paternalism, manipulation, dependency-fostering, and value-substitution, with philosophically grounded rubrics and human-validated LLM-as-judge scoring.

Behavioral evaluation of LLMs
Societal impacts
Evaluations

Diagnosing the Mechanisms Behind Honest-Looking Language-Model Behaviour

Bryan Chan · independent

This project will build controlled evaluations that distinguish context-general honesty from sycophancy, surface-cue shortcuts, refusal spillover, latent knowledge, and evaluation-conditioned behaviour. Mentees will test several tractable language models under matched prompt transformations and produce a validated failure taxonomy, reproducible evaluation suite, and empirical report.

Behavioral evaluation of LLMs
Evaluations
Alignment

When RLVR Changes the Model, the Safety Test, or Both

Muhammad Aaliyan · OCN (OneCarNow); Independent AI Safety Researcher

We will test whether safety-relevant behavior or evaluation reliability changes across reinforcement learning with verifiable rewards checkpoints. The project will replicate a completed Tülu 3.1 study on a second open training lineage and stress-test the result across prompt wording, answer order, and open-ended evaluation.

Behavioral evaluation of LLMs
Evaluations
Alignment