Forty candidate research ideas on the epistemic aspect of AI societal impact, including both descriptive projects characterizing such impact and algorithmic projects developing solutions.
About the project
Theory of change
There are clear pathways to human disempowerment from AI within the next ~15 years, including through physical and epistemic means. I prioritize research on preventing catastrophic epistemic disempowerment due to its neglectedness.
Your role
Owner of the project. We will help unblock you, give you guidance/feedback, or point you to relevant resources, but the project you pick will ultimately be your project. As a result, you will also be the first author of your project, unless there's an explicit agreement otherwise at the start.
Prerequisites
The list of project ideas has quite a large span, so you can probably find one there for you whether you are an engineering person, a data science person, a theory person, or someone else.
We require:
(1) basic technical competence (e.g. one of: fluency in Python + familiarity with the Linux workflow; being good at stats or some other branches of maths) in at least one of the aforementioned domains;
(2) the agency to quickly learn stuff on the fly & to try lots of tweaks to unblock yourself when smaller difficulties arise;
(3) gears-level understanding of how language model training/inference works. (no hands-on experience required)
Time commitment
Ideally 20+, negotiable
Location preference
All welcome
Application question(s)
-
For a project idea in the linked Google document, think of a first step (e.g. an experiment design for an empirical project; a theoretical intuition/numerical validation for a theory project) you would do. Heed this advice: https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html (<300 words)
-
What is your most impressive engineering project, if any? What’s one single most impressive detail about it? (<100 words)
-
What is your most impressive theoretical or conceptual finding, if any? What’s one single most impressive detail about it? (<100 words)
-
(Optional) Anything else you think we should know?
About the mentors

Tianyi Alex Qiu
Anthropic Fellows / Peking University
Tianyi is a London-based Anthropic AI Safety Fellow. He conducts research on AI alignment, with a focus on how it interacts with human epistemology and moral progress. Methodology-wise, he aims to employ experimental, formal, and social science methods alike. Projects he led on this topic have been awarded Spotlight (NeurIPS'24) and Best Paper Award (NeurIPS'24 Pluralistic Alignment Workshop) respectively. He also co-led work on the formal modeling of alignment resistance (Best Paper Award, ACL'25) and the survey paper “AI Alignment: A Comprehensive Survey” (2023).
Tianyi previously conducted research with the Center for Human-Compatible AI and PKU-Alignment. He serves with Zhonghao He as co-PIs on Project Prevail, a Foresight Institute-funded research program aiming to facilitate human moral & knowledge progress with AI.

Zhonghao (何忠豪) is a master’s student at the University of Cambridge. He works on AI alignment, interpretability, and human-AI interaction research. His previous work got accepted by ICML, ACM FAccT, and ICLR (workshop). His major interests are to design machines that help humans learn, think, and deliberate. Currently he focuses on two things, to develop truth-seeking AI (Bayesian & exploring truth), and to solve “positive feedback loop” problems in tech products: LLM sycophancy, confirmation bias in reasoning models, social media echo chamber, and polarization.