Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Automating predictive biology benchmarks


Mentees will improve Auto-PBB, an automated pipeline that generates predictive biology tasks from the published literature. Work spans task extraction, verification that generated tasks are genuinely answerable, and contamination checks against models that have likely seen the source papers.

About the project

Predictive biology benchmarks can detect superhuman capability, but hand-building enough verifiable tasks is slow and data-limited. Auto-PBB is an automated pipeline that generates large numbers of new prediction tasks from the published literature, scaling how fast SecureBio can measure frontier models' biological reasoning.

Theory of change

Predicting experimental outcomes is one of the better ways to measure whether a model has reached expert or superhuman biological reasoning, which is a threshold several frontier safety frameworks are written around. The constraint is supply: hand-built tasks require expert time and unpublished data, so the benchmark grows slowly while models improve quickly, and a benchmark that saturates or goes stale stops informing the decisions it was built for. Automating task generation changes the growth rate, which is the difference between an evaluation that tracks the frontier and one that trails it. A working literature-to-task pipeline also generalizes beyond biology.

Your role

The mentee will own one stage of the pipeline, most likely extraction or verification, with the project lead setting scope and reviewing. Code contributes to a shared repository under our engineering norms.

Prerequisites

Strong Python, including building multi-stage data pipelines. Has built something using LLMs as components in a pipeline, including prompt iteration and output validation. Able to read primary biology literature and identify what was measured and what was concluded. A life sciences background helps more here than on our other projects.

Application question(s)

Please answer one of the below, 300-500 words.

  1. How should a managed-access program for bio-capable models be set up? What are the relevant parameters, what do you suggest, and why?
  2. Suppose a wet-lab uplift study is run. What kind of data would you want to collect and how would you propose those data inform subsequent in silico model evaluations?

About the mentor

SecureBio AI

SecureBio AI

SecureBio

View profile

SecureBio is a nonprofit biosecurity research organization specializing in technical research to mitigate risks from catastrophic pandemics. Our AI team develops rigorous benchmarks and evaluation frameworks to assess AI systems' biological capabilities, as well as mitigation strategies that can reduce risks once AI capabilities cross specific risk thresholds. We perform pre-release safety testing of frontier models (e.g. GPT-5.6), and our evaluations have been featured in the model cards of OpenAI, Anthropic, and Google DeepMind. Our work has also informed national security briefings and emerging governance standards.

Similar projects