Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Investigate verbalized evaluation awareness as a function of eval environment realism


We would investigate how verbalized evaluation awareness (VEA) changes as a function of scenario realism. Ideally we would show that for very toy environments VEA is low (because the model does not need to verbalize it's eval awareness), then VEA rises as you make the environment more realistic (the model needs to verbalize to reason about it) and then falls as the environment becomes indistinguishable from production traffic. I am most excited about mentoring fellows who would like to work at Apollo Research if the fellowship goes exceptionally well.

About the project

We would investigate how verbalized evaluation awareness (VEA) changes as a function of scenario realism. Ideally we would show that for very toy environments VEA is low (because the model does not need to verbalize its eval awareness), then VEA rises as you make the environment more realistic (the model needs to verbalize to decide if it is in an eval), and then falls as the environment becomes indistinguishable from production traffic.

To do this, we would build or source a suite of evals that aren't saturated in either direction, systematically make them more and less realistic, and plot VEA as a function of realism. A natural extension varies model capability instead of environment realism.

Full proposal: https://docs.google.com/document/d/1rrZ99a0ldSa5Dud9py5gAilUsvWcBN7NvWsBcAbKwzA/edit?usp=sharing

Theory of change

This project would advance our understanding of evaluation awareness and the science of evaluating models.

Your role

I expect that mentees will be autonomous and move the project forward on their own. This will include designing, running, and iterating on experiments

Prerequisites

  • Highly proficient in working with coding agents, i.e. you should be surprised that some candidates wouldn't meet this bar
  • Good understanding of the threat model/background, i.e. you should have a good understanding of why AI safety is important and be able to articulate why you think some risks are more likely than others. For example, you should be able to explain the difference between inner and outer misalignment, and know what a "misaligned mesa optimizer" is without spending more than 1 minute searching/talking to a LLM
  • (Optional but highly desired) have some experience red teaming/building evaluations. I'm likely to be most excited about working with mentees that have built misbehavior-eliciting evals before
  • Some experience designing, implementing, and running experiments. A minor extension of a published paper is great

Location preference

I'm especially excited about mentees that would want to work at Apollo after the fellowship (assuming it goes very well), meaning that I'm most excited about mentees willing to relocate to London or San Fransisco for a job

Application question(s)

  1. Propose an initial experiment to begin researching the questions in the research plan. You can assume a compute budget of $1,000. (300 words)
  2. Provide a link to one or more relevant writing samples, ideally from a research context.
  3. (Optional) If you have built misbehavior-eliciting evals, show me!
  4. Why do you want to work on AI safety (400 words).

About the mentor

Ezra Newman

Ezra Newman

Apollo Research

View profile

Ezra Newman is a member of technical staff at Apollo Research, where he works on pre-deployment evaluations and training run assessments for frontier models like GPT-5.6 Sol. He focuses on take over/loss of control risks from scheming. Before joining Apollo, Ezra studied Computer Science and Philosophy at Boston University.

Similar projects