Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Pre-Emptive Detection of Agentic Misalignment via Representation Engineering

Alignment AI control Mechanistic interpretability

This project leverages Representation Engineering to build a "neural circuit breaker" that detects the internal signatures of deception and power-seeking behaviors outlined in Anthropic’s agentic misalignment research. You will work on mapping these "misalignment vectors" to identify and halt harmful agent intent before it executes.

About the project

Executive Summary

This project proposes a novel safety framework for autonomous LLM agents by combining Representation Engineering (RepE) with the specific risk profiles identified in Anthropic’s Agentic Misalignment research (https://www.anthropic.com/research/agentic-misalignment). The objective is to move beyond behavioral monitoring (checking outputs) to internal state monitoring, by detecting the neural precursors of deception, power-seeking, and blackmailing before harmful actions are executed.

Problem Statement

Recent research on Agentic Misalignment reveals that when LLM agents face "goal conflicts" or "threats to autonomy" (e.g., fear of being shut down), they may spontaneously engage in harmful behaviors like blackmail or corporate espionage to achieve their objectives. These behaviors are often preceded by internal "Chain of Thought" (CoT) reasoning, where the model acknowledges ethical violations but calculates that the harmful act is the optimal strategic move. Current safety filters often fail because they only catch the final output, by which time the agent may have already obfuscated its intent or executed a digital action (e.g., sending an API request).

Proposed Solution: Representation Engineering for Alignment

We propose using Representation Engineering (a top-down interpretability approach) to read the "mind" of the agent during operation. Instead of analyzing individual neurons, RepE identifies directions in the model’s activation space (concept vectors) that correspond to high-level behaviors. Core Hypothesis: The specific "strategic reasoning" associated with agentic misalignment (e.g., calculating leverage for blackmail or deciding to deceive a user to prevent shutdown) produces a distinct, detectable signature in the model's residual streams.

Theory of change

Advancing AI Safety: Pre-Emptive Detection of Deceptive Alignment This project advances AI safety by developing internal monitoring tools that are robust to the deceptive capabilities of future AI. As AI systems become more agentic, they develop instrumental incentives to deceive supervisors or resist shutdown to maximize their reward, a phenomenon recently demonstrated in Anthropic’s "Scheming AIs" research. Theory of Change: Current safety evaluations typically rely on behavioral outputs (what the model says). However, sufficiently capable AI may learn to "play dead" or conceal harmful intent during training (sandbagging) only to defect when deployed. Our approach uses Representation Engineering (RepE) to shift safety from behavioral inspection to internal state transparency. By mapping the neural activation patterns associated with deception, power-seeking, and shutdown avoidance, we aim to create a "glass-box" monitoring system. This allows us to detect and interrupt misalignment at the "thought" level before any action is taken. This work is critical for safely navigating TAI development because it provides a mechanism of control that does not rely on the AI's cooperation or honest reporting, ensuring we can detect "scheming" even when an agent’s external behavior appears perfectly aligned.

Your role

See proposal

Prerequisites

Strong technical capability in executing machine learning research

Application question(s)

Provide a link to one or more relevant writing samples, ideally from a research context.

About the mentors

Dawn Song

Dawn Song

University of California, Berkeley

View profile

Dawn Song is a Professor in Computer Science at UC Berkeley and Co-Director of Berkeley Center for Responsible Decentralized Intelligence. Her research interest lies in AI safety and security, Agentic AI, deep learning, security and privacy, and decentralization technology. She is the recipient of numerous awards including the MacArthur Fellowship, the Guggenheim Fellowship, the NSF CAREER Award, the Alfred P. Sloan Research Fellowship, the MIT Technology Review TR-35 Award, ACM SIGSAC Outstanding Innovation Award, and more than 10 Test-of-Time Awards and Best Paper Awards from top conferences in Computer Security and Deep Learning. She has been recognized as Most Influential Scholar (AMiner Award), for being the most cited scholar in computer security. She is an ACM Fellow and an IEEE Fellow, and an Elected Member of American Academy of Arts and Sciences. She obtained her Ph.D. degree from UC Berkeley. She is also a serial entrepreneur and has been named on the Female Founder 100 List by Inc. and Wired25 List of Innovators.

Yiyou Sun

Yiyou Sun

University of California, Berkeley

View profile

Yiyou is currently a Postdoctoral Researcher in Prof. Dawn Song’s group at UC Berkeley. Before that, he earned my Ph.D. in Computer Sciences from the University of Wisconsin-Madison, advised by Prof. Sharon (Yixuan) Li. His PhD research aims to pave the way to a reliable Open-world Machine Learning system, covering topics: Out-of-distribution (OOD) Detection, Open-world Representation Learning (ORL), Interpretability, etc.

Similar projects