Eliezer Yudkowsky
Personal reference note · last reviewed 12 Aug 2026
One-line orientation
Autodidact who helped invent the modern AI-alignment conversation, founded LessWrong, and now argues publicly that building superintelligence with present techniques will most likely kill everyone.
Background
Raised in an Orthodox Jewish family. No formal high-school or university education; describes himself as entirely self-taught. Treats writings from 2001 and earlier as the product of a different person. Early involvement in online transhumanist and singularity discussion lists in the late 1990s.
Moved to Atlanta in 2000 to help launch the Singularity Institute for Artificial Intelligence (later MIRI). Later based in the Bay Area.
Institutions and roles
- Co-founder (2000) and research fellow, Machine Intelligence Research Institute (MIRI), Berkeley.
- Principal early contributor to the blog Overcoming Bias (with Robin Hanson), 2006–2009.
- Founder of LessWrong (2009), the community site that became the main venue for the Sequences and the rationalist subculture.
Core technical and conceptual contributions
- Friendly AI / AI alignment — early framing of the problem of ensuring advanced AI systems remain beneficial under recursive self-improvement.
- Seed AI and intelligence explosion (“AI foom”) — argument that a system capable of improving its own intelligence could undergo rapid, discontinuous capability growth.
- Coherent Extrapolated Volition (CEV) — proposed target for AI goals: what humans would want if they knew more, thought faster, and had reached reflective equilibrium.
- Timeless Decision Theory and related work on decision theories that remain stable under self-modification and logical correlation.
- Paperclip maximizer as a concrete illustration of orthognality (intelligence and final goals are independent) and instrumental convergence.
Major writings
- “Creating Friendly AI” (2001) — early statement of the problem and terminology.
- The Sequences (2006–2009, later collected as Rationality: From AI to Zombies, 2015) — Bayesian reasoning, cognitive biases, epistemology, meta-ethics, AI risk.
- Harry Potter and the Methods of Rationality — long fanfiction that introduced many readers to rationalist ideas.
- If Anyone Builds It, Everyone Dies (2025, with Nate Soares) — mass-market statement of the extinction thesis; New York Times bestseller.
Current public position on AI risk
Maintains that current machine-learning methods do not provide reliable control over the goals of systems that exceed human intelligence across domains. Default outcome of creating artificial superintelligence under present techniques and incentives is human extinction. Has publicly cited probabilities in the high 90s percent range (e.g., 99.5 % in interviews around the 2025 book).
Advocates strong international coordination to halt or tightly regulate development of systems that could reach superintelligence with existing approaches.
Influence and dissent
Widely credited by figures inside major labs: Sam Altman has said Yudkowsky was “critical in the decision to start OpenAI.” Early introductions helped bring funding to DeepMind. Nick Bostrom’s Superintelligence (2014) drew on the intelligence-explosion framing.
Counter-views are common among researchers who assign lower probabilities to abrupt, uncontrollable takeoff, who believe incremental empirical alignment work is more promising, or who reject the strong form of the orthogonality thesis. Public disagreement has been visible with researchers at OpenAI, Anthropic, and academic groups since at least 2022–2023.
Personal notes (local only)
- Any formal quantitative revision of his personal probability estimates after the 2025 book and subsequent empirical results?
- Current internal MIRI research priorities versus the public extinction-message focus?
- Detailed response to post-2024 claims that certain classes of agentic misalignment are already near zero under realistic training regimes?
- How he weighs the possibility that current large models already contain latent capabilities that change the timeline arguments?