This project investigates whether language models replicate human behavioral phenomena by adapting experimental designs from the social sciences. Mentees will select a documented human behavior and design experiments to test whether it emerges in language model outputs.
About the project
Language models are trained on vast corpora of human-generated data that encodes not just language, but the behavioral patterns, biases, and social dynamics of the people who produced it. A natural question follows: do language models replicate human behavioral phenomena?
This project draws on the social sciences (i.e., sociology, communication, psychology, and organizational theory) to investigate this question. These fields have spent decades developing experimental paradigms to understand what humans do, under what conditions, and why. By adapting these experimental designs and substituting language model instances for human subjects, we can begin to map which human behavioral phenomena emerge in AI systems.
This work addresses a gap in AI safety research. Much of the field focuses on capabilities, alignment techniques, or interpretability at the mechanistic level. Less attention has been paid to the behavioral level: understanding what language models do in contexts that have been well-studied in humans. Establishing robust behavioral findings creates a foundation for future interpretability work that can investigate why these patterns occur.
Prior work in this vein has tested escalation of commitment and truth-bias in reasoning models. This project continues that line of research, remaining deliberately broad to allow mentees to explore phenomena that match their interests and background.
What mentees will do:
- Identify a human behavioral phenomenon documented in the social science literature
- Locate or design an experimental paradigm used to study that phenomenon
- Adapt the paradigm for language model evaluation
- Run experiments and analyze results
- Contribute to a growing body of work on behavioral replication in AI systems
Ideal mentee background:
- Interest in understanding how AI systems reflect or diverge from human behavior
- Some technical proficiency (ability to run experiments with language models via API or similar)
- Familiarity with social science research methods and experimental design
- Openness to exploring a range of behavioral phenomena
I am agnostic about which specific phenomena mentees choose to investigate. The goal is to build a broader research program that systematically tests whether, and how, language models inherit the behavioral tendencies of their training data.
Theory of change
Understanding how language models behave is a prerequisite for ensuring they behave safely. This project contributes to that understanding by systematically testing whether language models replicate human behavioral phenomena, including potentially harmful ones like escalation of commitment, overconfidence, or in-group bias. Identifying these patterns is a necessary first step before we can mitigate them. This work also lays the groundwork for mechanistic interpretability research by establishing robust behavioral findings that future work can explain at a technical level.
Your role
Mentees will function as independent researchers with guidance. After an initial onboarding period where we establish shared context on the research program and methodology, each mentee will:
- Select a behavioral phenomenon from the social science literature that they find compelling and want to investigate in language models.
- Identify or design an experimental paradigm based on existing studies of that phenomenon in humans.
- Adapt and implement the experiment for language models, including prompt design, experimental conditions, and data collection.
- Analyze results and write up findings with the goal of producing a publishable contribution.
I will provide supervision through regular check-ins, feedback on experimental design, and guidance on connecting findings to the broader research program. However, mentees should expect to drive their own projects forward between meetings. This structure rewards initiative and curiosity.
Prerequisites
Required:
- Proficiency in Python sufficient to write and run experiments independently.
- Experience querying language models via API (e.g., OpenAI, Anthropic, or open-source models through Hugging Face).
- Familiarity with experimental research design, including concepts like independent/dependent variables, control conditions, and statistical significance.
- Exposure to at least one social science discipline (e.g., psychology, sociology, communication, organizational behavior, economics) through coursework or independent study.
Strongly preferred:
- Experience reading and synthesizing academic literature.
- Prior experience designing or running a research study (does not need to be published).
- Basic statistical analysis skills (e.g., hypothesis testing, regression, or equivalent).
Application question(s)
Question 1: Read my paper on escalation of commitment in language models [ https://arxiv.org/abs/2508.01545 ]. In 150-250 words, propose a project following the same approach: (1) a behavioral phenomenon from the social sciences, (2) an experimental paradigm used to study it in humans, and (3) how you would adapt it to test language models. Focus on showing your reasoning. This does not need to be a final proposal.
Question 2: Provide a link to one or more relevant writing samples, ideally from a research context.
About the mentors

Emilio is an independent researcher specializing in the social impacts of AI, with a focus on human-AI interaction, language model behavior, and organizational adoption of technology. Previously, he was a research manager at Columbia AI Alignment Club, where he led projects evaluating how language models exhibit human behaviors. He was also a graduate researcher in the Management Division at Columbia Business School, investigating how emerging player performance technologies are adopted in amateur competitive youth sports. He holds an MA in Media Studies and Sociology from Columbia University and a BA in Communication and Political Science from Brigham Young University–Hawaii.

I'm defending my PhD in two weeks, in Brain & Cognitive Sciences under Professors Antonio Damasio and Jonas Kaplan. My work involves understanding failures and vulnerabilities of complex intelligent systems, specifically investigating how antisocial traits manifest in the brain and behavior, as well as within LLMs.