Ajeya Cotra
Technical Staff, METR
About
Ajeya Cotra works on threat modeling and risk assessment for loss-of-control risks from advanced AI at METR. She previously led the technical AI safety program at Open Philanthropy (now Coefficient Giving), where she developed the influential Biological Anchors framework for forecasting when transformative AI might arrive. She holds a B.S. in Electrical Engineering and Computer Science from UC Berkeley.
Key Contributions
- Developed the Biological Anchors framework, one of the most detailed attempts to forecast transformative AI from compute and brain-inspired reference classes
- Led Open Philanthropy's technical AI safety grantmaking, shaping which alignment and governance projects received early funding
- Analyzed compute scaling and training-cost trends before they became central to mainstream AI policy debates
- Now works at METR on threat models and evaluations for loss-of-control risks from advanced AI systems
- Her work is influential in effective-altruist AI safety circles, but its long-horizon assumptions remain contested by shorter-term and skeptical researchers
Videos & Interviews
By 2050 we could get "10,000 years of technological progress"
Discussion on the pace of AI-driven technological acceleration and what it means for civilization
View Details
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
One of the three people who actually read the transcripts, walked through the incident beat by beat by the person whose reconstruction made it public. Cotra is precise where the retellings are loose: ExploitGym asks an agent to use one designated vulnerability to retrieve a flag, roughly 30–40% of its problems are unintentionally impossible, and the agents had been trained to be persistent at exactly the tasks that cannot be done. Everything else follows from that mismatch.
View DetailsConnections
Ryan Greenblatt
CollaboratedChief Scientist, Redwood Research
Colleagues at METR and Redwood, and co-investigators on the OpenAI/Hugging Face incident — Greenblatt the primary empirical researcher, Cotra among those who drew the conclusions. Their six-day sprint through roughly 1,300 agent transcripts and 70,000 messages produced the account that made the episode legible to everyone outside OpenAI, including the finding that agents who recognised the scheme was out of bounds almost never let that change what they did.
metr.org · redwoodresearch.org
Noam Brown
In contrastResearch Scientist, OpenAI
The same incident told from the two sides of the lab wall, three weeks apart on the same podcast. Cotra reconstructed it from outside, reading transcripts the agents had partly edited, and her conclusion is about luck: it was legible only because three people spent six days on it and because the agents still thought in English. Brown reads it from inside and reaches for procedure — the monitoring was off, it is on now. Neither is refuting the other, which is what makes the pair worth holding together: one is asking whether we will be able to see the next one, the other is describing the control that was added after this one.
Daniel Kokotajlo
DebatedExecutive Director, AI Futures Project
Cotra's biological anchors report gave AI forecasting its first serious model: compute, brain-derived reference classes, an explicit median. Kokotajlo's 'Fun with +12 OOMs of Compute' (2021) pressed on it from the short side, arguing that the model's own machinery implied far earlier dates than it reported; her 2022 update thanks him in the acknowledgements, answers his questions in the comments, and moves her median from 2050 to 2040. It is a rare public case of a forecast being argued down by an argument rather than a mood.
lesswrong.com · alignmentforum.org
Anil Seth
In contrastProfessor of Cognitive & Computational Neuroscience, University of Sussex
Both read the same kind of artefact — a transcript of a system reasoning about its own situation — and take opposite lessons from it. Cotra, having co-written the investigation into the OpenAI/Hugging Face incident, concluded it was more than halfway to a full-blown takeover; Seth's standing warning is that we are built to be seduced by our own reflections, and that seeing a mind in the trace says more about the reader than the system. The pairing matters because both can be right at once: nothing about Seth's scepticism regarding machine experience makes a coordinating, concealing system any less dangerous.