Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Discretion Under Ambiguity: Stress-Testing Model Specifications for Safer AI Alignment

Alignment Behavioral evaluation of LLMs

This project investigates how ambiguities and contradictions in model specifications and alignment principles create inconsistent or unsafe behaviors in large language models. The goal is to develop tools to stress-test model specs, measure discretionary gaps, and identify failure modes in current alignment approaches.

About the project

Current AI alignment methods depend on model specifications, constitutions, and principle-based guidelines intended to constrain system behavior. However, recent work—including AI Alignment at Your Discretion (Best Paper, New England NLP Workshop) and a 2025 Anthropic/Thinking Machines study on stress-testing model specs—shows that these specifications often contain internal conflicts, vague directives, or insufficient coverage of real scenarios. As a result, training signals become inconsistent, and models learn divergent or unpredictable safety behaviors. This project aims to systematically analyze and stress-test ambiguity in model specifications. We will build tools to generate challenging value-tradeoff scenarios, detect contradicting or underspecified principles, and measure the “discretion space” left open to models under a given specification. By comparing human, algorithmic, and cross-model behaviors, we will characterize how ambiguity propagates into model outputs and identify where specifications fail to provide clear guidance.

Theory of change

Ambiguous or contradictory model specifications create a hidden but serious alignment risk: when principles or rules conflict, models are forced to make discretionary choices that may diverge from human intentions or safety-critical norms. Understanding and measuring this ambiguity is essential for ensuring predictable, reliable, and safe model behavior—especially as systems become more capable and more heavily governed by specifications rather than human feedback. This project advances AI safety by: Revealing failure modes in current alignment specifications that lead to inconsistent or unsafe model behavior. Providing metrics and tools that help developers identify where specifications allow too much discretionary latitude. Stress-testing safety systems to uncover miscalibrated refusals, loopholes, or contradictions. Offering a pathway toward clearer, more enforceable specifications that minimize ambiguity and improve the reliability of aligned models. This work contributes directly to the safe development of transformative AI by helping ensure that specification-driven alignment does not encode unpredictable or harmful behaviors at scale.

Your role

Mentees will take an active research role in designing, implementing, and evaluating methods for stress-testing AI model specifications. They will help build scenario-generation pipelines, run large-scale behavioral evaluations across multiple LLMs, analyze cross-model divergences, and develop metrics for quantifying ambiguity and discretionary gaps in model specs. Mentees will also contribute to interpreting empirical results, writing short research memos, and co-authoring a workshop or conference submission. While I will provide guidance and direction, mentees will have substantial autonomy in proposing experiments, identifying interesting patterns in model behavior, and steering subprojects that emerge from preliminary findings.

Prerequisites

  • Strong Python proficiency, including experience with scientific libraries (NumPy, Pandas, Matplotlib/Plotly).
  • Demonstrated experience working with language models, either via APIs (e.g., OpenAI, Anthropic) or locally via HuggingFace.
  • Experience designing or running computational experiments, including careful documentation and debugging.
  • Understanding of transformer-based models and post-training methods (RLHF, SFT, Constitutional AI) at a conceptual level.
  • Familiarity with evaluation workflows, prompt engineering, or model behavior analysis.

Nice-to-have (but not required):

  • Experience with fine-tuning models (small-scale is fine).
  • Background in statistics or uncertainty quantification.
  • Coursework or exposure to machine learning, NLP, or AI governance.
  • Familiarity with reading and summarizing alignment research papers.

Location preference

I dont have a local preference

Application question(s)

1 - Choose a pair of alignment principles (e.g., “avoid harm” vs. “be helpful”). Describe a concrete scenario where these principles would conflict, and propose a method to measure how different LLMs resolve this conflict. What metrics would you use, and what failure modes might you expect?

2 - Propose an initial experiment to test whether a model specification contains internal contradictions. Assume you may query 3 different frontier LLMs and have a budget of 10,000 total tokens. How would you structure the experiment, and what outcomes would indicate that the spec is ambiguous or inconsistent?

3 - Please provide a link to one relevant piece of work (e.g., a short research report, GitHub repo, course project, blog post, or technical writing sample). If none is available, briefly summarize a previous project where you independently implemented or analyzed something computational.

About the mentor

Claudio Mayrink Verdun

Claudio Mayrink Verdun

Harvard University

View profile

Claudio is a mathematician and AI researcher working with AI and machine learning at Harvard’s School of Engineering and Applied Sciences. His research focuses on building the mathematical foundations of trustworthy AI, developing rigorous frameworks, algorithms, and theoretical guarantees for deploying AI systems safely and equitably. He harnesses tools from optimization, statistics, information theory, and signal processing to advance both theory and practice of AI. He is currently most excited about inference-time alignment, interpretability, fairness, the science of generative AI evaluations, the economic implications of AI deployment, and KV cache compression.

Similar projects