Mentees will build evaluation scenarios grounded in how people actually used AI during a large wet-lab uplift trial. Work spans qualitative coding of interaction logs, comparison against SecureBio's existing benchmark coverage, and piloting new items against frontier models.
About the project
Benchmark scenarios are written by experts, so we rarely know how they map onto how people actually use AI during real lab work. This project mines the LLM interaction logs from a real-world uplift trial to find where models helped or failed real participants and to turn those moments into new evaluation scenarios grounded in observed behavior.
Theory of change
Every biosecurity benchmark in this field is written by experts imagining how a malicious or naive user would behave, and that assumption is almost never checked against data. If real usage diverges from what benchmarks test, eval scores measure something adjacent to the risk we care about, and developers and regulators are calibrating on the wrong number. The trial in question is among the largest wet-lab uplift studies run to date, and its interaction logs are the closest thing the field has to ground truth on how people query models during a real biological workflow. Feeding that back into benchmark design makes our results more defensible to the people who act on them.
Your role
The mentee will lead the analysis described above, with the project lead setting scope and reviewing the coding scheme.
Prerequisites
Comfortable building and applying a coding scheme to unstructured data, or equivalent qualitative or mixed-methods research experience. Working proficiency in Python for data analysis. Enough familiarity with molecular biology or wet lab work to follow a protocol and recognise where someone went wrong. LLM evaluation experience is welcome but not required.
Application question(s)
Please answer one of the below, 300-500 words.
- How should a managed-access program for bio-capable models be set up? What are the relevant parameters, what do you suggest, and why?
- Suppose a wet-lab uplift study is run. What kind of data would you want to collect and how would you propose those data inform subsequent in silico model evaluations?
About the mentor

SecureBio is a nonprofit biosecurity research organization specializing in technical research to mitigate risks from catastrophic pandemics. Our AI team develops rigorous benchmarks and evaluation frameworks to assess AI systems' biological capabilities, as well as mitigation strategies that can reduce risks once AI capabilities cross specific risk thresholds. We perform pre-release safety testing of frontier models (e.g. GPT-5.6), and our evaluations have been featured in the model cards of OpenAI, Anthropic, and Google DeepMind. Our work has also informed national security briefings and emerging governance standards.