Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Reproducible Robotics Benchmarks

Generalist Evaluations Other

Robocurve and Generator Residency are running a $500,000 open call for reproducible robotics benchmarks.

As part of this project, you will evaluate applications, select the finalist teams to receive funding, run office hours, and do whatever else is required to ensure a successful program. Roughly 10 hrs/week, mid-September to mid-December.

About the project

The main goal of this program is to build a repository of reproducible evaluations that various stakeholders (governments, academia, and the public) can use to benchmark general robotics capabilities, providing a trusted means for the public to gauge progress in robotics. With this infrastructure, we can better understand and prepare for physical automation of labor and a possible intelligence explosion.

Theory of change

Currently, there are very few well-run, standardized benchmark methods for measuring general robotic capabilities. Leading robotics AI models are released with each company's private task lists, without third-party testing and evaluation. This leaves AI forecasters, policymakers, and the public without a reliable understanding of how capable these systems really are or how fast they're improving.

Forethought's work on the intelligence explosion notes that once robots can substitute for human labor, a faster industrial explosion could occur. This means AI could seize decisive power sooner, since most routes to takeover run through control of physical infrastructure. The intelligence curse also argues that once powerful actors no longer need human labor due to advancements in robotics capabilities, they lose the incentive to invest in their welfare. We are concerned about possible gradual disempowerment and intelligence explosions that could arise as a result of these advancements in robotics capabilities.

Without shared, reproducible benchmarks, people outside the labs have very little basis for judging how soon robots will take on real-world jobs, or how quickly the field is moving. With these benchmarks, we hope to inform the public and policymakers about the rate of progress in general robotics and help to reduce the risks it poses.

Your role

There are two main stages of this program. 1. Stage 1: Selection, mid-September to late September. Review Stage 1 proposals. Most of the work is judgement rather than writing code. You're going to support the creation of the evaluation rubric for applications received in the public call for proposals and actually evaluate them to find the top 12 finalist teams.

2. Stage 2: Support, October to December. Help selected teams reach a working benchmark by onboarding them to Inspect Robots, running office hours, checking in on teams to help them make progress, and doing whatever else is necessary to produce the highest-quality output.

Likely output: the first 10 to 12 independent robotics benchmarks that other people can re-run, plus the onboarding material for the next cohort.

Prerequisites

Required

  • Comfortable with Python and reading unfamiliar codebases
  • Can explain technical things clearly in writing
  • Highly agentic and proactive in solving problems

Helpful

  • Experience with creating evaluations or benchmark design
  • Robotics or RL background
  • TA experience

You don't need hardware access or any prior robotics work.

Application question(s)

Question 1: Why are you interested in supporting robotics evaluation specifically, rather than AI evaluation or robotics in general? (~10min) Question 2: A proposal contains this task: "The robot places the mug on the shelf. Score 1 if the mug ends up on the shelf, 0 otherwise." What would you need specified before an independent operator could run this? Which of these are most important to get right? (~25 min)

Question 2: Why are you interested in supporting robotics evaluations? (~10min)

Question 3: The project has two heavy weeks: mid-September and mid-December. Roughly how many hours could you give the project these two weeks? An honest answer doesn't count against you; we're asking just so we can plan ahead.

Note: Both of the above are non-technical, and we're really looking for people who can think clearly and make good judgment. They shouldn't take very long, and we request that you not use AI tools for writing or ideation.

About the mentor

Yashvardhan Sharma

Yashvardhan Sharma

Constellation, Georgetown

Yashvardhan Sharma is a graduate student in Security Studies at Georgetown focused on AI and national security, currently working at Constellation. He studied Applied Computer Science and AI at Minerva University and previously worked on compute policy as a Research Scholar at MATS. His focus is bridging technical AI research in the Bay Area and policy in Washington, DC.

Similar projects