Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Causes, implications, and mitigations of evaluation awareness and introspection.

Evaluations Behavioral evaluation of LLMs

There is growing evidence that language models exhibit degrees of situational awareness [1, 2], i.e., they encode or verbalize information consistent with their actual operating context, ranging from (i) distinguishing training vs testing vs deployment (see, e.g., [3, 4, 5, 6, 7]), to (ii) recognizing their own outputs [8] or exhibiting introspective access over their activations [9, 10].

We are especially interested in supervising projects that empirically investigate the science behind these phenomena: what causes them, what they imply in the context of safety, the extent to which they are present in current LLMs, how to mitigate them in specific settings.

[1] https://arxiv.org/abs/2309.00667 [2] https://arxiv.org/abs/2407.04694 [3] https://arxiv.org/abs/2505.23836 [4] https://www.antischeming.ai/ [5] https://cdn.openai.com/gpt-5-system-card.pdf [6] https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf [7] https://storage.googleapis.com/deepmind-media/gemini/gemini3profsfreport.pdf

About the project

Examples of projects about evaluation awareness (for representative projects see, e.g., [11, 12, 13]):

  • Building white- and black-box evaluations to quantify the extent to which models are, in fact, evaluation aware. Techniques here range from monitoring chains-of-thought to training probes on models’ hidden states.

  • Building model organisms for evaluation awareness, in order to identify factors (e.g., in the data, or in the training procedure) that make awareness more prominent. Examples here include: --- Intervening on the data (e.g., ablating or including formatting cues); --- Comparing supervised- vs reinforcement-learning-finetuning; --- Adopting as LLM-as-judge a model from the same family of the tested model; --- Studying if awareness in one domain generalizes to others; --- Studying how performance on reasoning and safety benchmarks is affected by awareness.

  • Proposing mitigations to evaluation awareness. For example: --- Designing evaluations adversarially perturbed to be indistinguishable, from the point of view of the model, from deployment; --- Synthetic document finetuning to instill, in a model, the false belief of being deployed; --- Steering models in the direction of un-awareness.

Examples of projects about introspection (for representative projects see, e.g., [8, 9, 10, 14]):

  • Investigating whether introspection can be helpful for safety: e.g., whereas inconsistent answers about a model’s knowledge (or, more generally, “beliefs”) from a non-introspective model can be attributed to hallucinations or lack of capabilities, inconsistencies produced by an introspective model (which is better at accessing/explaining certain hidden states) may signal “dishonesty” or “obfuscation” at a more concerning level.

  • Investigating whether introspection can be harmful for safety: e.g., if introspective models can act upon certain information encoded in their attention heads without verbalizing it, the risks already associated with CoT “un-faithfulness” and “obfuscation” would exacerbate.

[8] https://arxiv.org/abs/2404.13076 [9] https://arxiv.org/abs/2410.13787 [10] https://transformer-circuits.pub/2025/introspection/index.html [11] https://arxiv.org/pdf/2505.14617 [12] https://arxiv.org/pdf/2507.01786 [13] https://arxiv.org/abs/2510.20487 [14] https://arxiv.org/pdf/2207.05221

Theory of change

In general, evaluation awareness concerns AI safety in at least two opposite directions:

  • Certain eval-datasets are contaminated by the fact that RLHF’d evaluation-aware models have learnt to, e.g., underperform at evaluation time [15]. This may result in underestimating the harmfulness of models at deployment time.

  • By contrast, other eval-datasets are polluted by the fact that models learnt to roleplay, and are roleplaying-aware. This may result in overestimating the propensity of a model to engage in a dangerous behaviour where roleplaying is not the natural course of action to take [16].

Similarly, introspection opens the possibility of querying a model and asking it why a certain outcome was obtained. This is especially relevant to investigate reasoning-faithfulness, lie-detection, and goal-formation and detection [10, 17].

[10] https://transformer-circuits.pub/2025/introspection/index.html [15] https://arxiv.org/pdf/2406.07358 [16] https://arxiv.org/pdf/2507.03409 [17] https://arxiv.org/abs/2503.10965

Your role

We expect mentees to choose, early on, the questions of the projects mentioned above that they find most exciting, subject to feasibility (given seniority and access to compute), state of the literature (in February and in expectation), and relevance.

After this, the mentees will share responsibilities roughly corresponding to first-authorship, with degrees of independence and help dependent on seniority. At minimum, the mentees will be responsible for the implementation of the experiments.

Help and guidance will be provided both at the research and implementation level, once again depending on seniority of the candidate(s).

Prerequisites

Knowledge:

  • Foundations of machine- and deep-learning; -Transformer architecture and language models;
  • Exposure to empirical AI safety literature (e.g., evaluations, mechanistic interpretability, …).

Experience:

  • Fluency with Python;
  • Designing and implementing machine learning workflows using PyTorch;
  • Supervised- or RL-fin tuning of language models, at least with toy experiments and some publicly available datasets;
  • Prompt engineering.

Bonus:

  • Experience with libraries such as vLLM, TRL, Hugging Face;
  • Familiarity with statistical hypothesis testing.

Time commitment

Minimum 10hrs/week, maximum 20hrs/week.

Location preference

No preferences subject to minimal overlap with Montreal's time zone (EST).

Application question(s)

Answer the following questions precisely and concisely. We look for understanding, pragmatism, and communication above, e.g., originality or polishedness. Please, do not rely on third people or language models to think about the questions or answer them.

  1. What experience(s) or project(s) of yours are most relevant for the projects we listed? --- Max 400 words, ideally < 250.

  2. Which of the projects we proposed is most interesting to you? If you were to start working on it tomorrow, given realistic resources, how would you approach it? --- For example, list a concrete hypothesis, an experimental design to test it, and what considerations would allow you to evaluate the results of the experiments relative to the question. Which practical problems do you expect to encounter? Which workarounds do you expect we should try?
    --- In the case where a similar idea exists in literature, explain the difference between that and your proposal; in particular, specify why you think it would be important to pursue your approach (comparing with the most relevant paper is fine). --- [Not needed and high-risk, we do not advise answering this question unless in exceptional circumstances] If you have an idea you’re particularly excited about which is relevant to the projects above, but that wouldn’t be feasible given resource constraints, feel free to elaborate on it. --- Max 500 words, ideally < 400.

About the mentors

Damiano Fornasiere

Damiano Fornasiere

LawZero

View profile

Damiano Fornasiere is a Research Scientist at LawZero, where he's working on (i) the math behind the Scientist AI, (ii) model organisms to study elicitation, (iii) interpretability and evaluation techniques pertaining situational awareness and introspection.

He worked as a Research Scientist at Mila, and obtained a PhD in Mathematics and Computer Science at the University of Barcelona.

Mirko Bronzi

Mirko Bronzi

LawZero

View profile

Mirko is a Research Scientist at LawZero with 15+ years of experience in NLP, deep learning, and applied AI research. His work spans large-scale model optimization, dataset design, and production-grade ML systems, with past roles at Mila and Nuance.

Mirko’s current focus is on understanding LLM behavior through model introspection techniques, and on how such methods can strengthen AI safety and reliability. For example, models with well-established introspective abilities can be reliably asked how they generated a given output, and inconsistencies in their self-reports may reveal early signs of deceptive or unsafe behavior. Conversely, models with weak introspective abilities may produce inconsistent explanations simply because they lack access to—or cannot faithfully report—their internal state, highlighting the importance of measuring introspection as a prerequisite for meaningful safety evaluations.

Similar projects