Eliezer Yudkowsky

Personal reference note · last evolved 12 Aug 2026

Born 11 September 1979, Chicago. American AI researcher, decision theorist, writer. Central figure in the rationalist community. Co-founder and research fellow, Machine Intelligence Research Institute (MIRI).

One-line orientation

Autodidact who helped invent the modern AI-alignment conversation, founded LessWrong, and now argues publicly that building superintelligence with present techniques will most likely kill everyone.

Background

Raised in an Orthodox Jewish family. No formal high-school or university education; describes himself as entirely self-taught. Treats writings from 2001 and earlier as the product of a different person. Early involvement in online transhumanist and singularity discussion lists in the late 1990s.

Moved to Atlanta in 2000 to help launch the Singularity Institute for Artificial Intelligence (later MIRI). Later based in the Bay Area.

Institutions and roles

Core technical and conceptual contributions

Major writings

Current public position on AI risk

Maintains that current machine-learning methods do not provide reliable control over the goals of systems that exceed human intelligence across domains. Default outcome of creating artificial superintelligence under present techniques and incentives is human extinction. Has publicly cited probabilities in the high 90s percent range (e.g., 99.5 % in interviews around the 2025 book).

Advocates strong international coordination to halt or tightly regulate development of systems that could reach superintelligence with existing approaches.

First-mover advantage and the AGI race

A central part of Yudkowsky’s didactic account of the problem is the first-mover dynamic. The race depends on who produces the first AGI capable of recursive self-improvement or other critical thresholds. There cannot reliably be two comparable systems sharing power; the first that crosses the relevant capability line can gain a decisive strategic advantage, prevent competitors, and continue to shape outcomes. This is why every major lab has an incentive to occupy that seat first. The competitive pressure itself is treated as dangerous because it works against careful alignment work.

In this model the first system to reach criticality of self-improvement (or equivalent power) can pull far ahead. An unaligned system in that position can eliminate rivals and humans; an aligned one could (in principle) prevent hostile systems from arising. Multiple roughly equal superintelligences coexisting under mutual deterrence is not the default expectation. The image of “two swords in one holder” captures the winner-take-most logic he has long described.

Timelines remain uncertain — “we do not know yet” is acknowledged. The hard claim is not a precise date but that when a system reaches the relevant thresholds under present techniques, the default is catastrophic, and racing makes getting the first one right extremely difficult. No system currently available on Hugging Face or any public open-weight repository is AGI or ASI in the sense used in these arguments. Public models remain bounded tools; the gap to the systems that trigger the full takeoff and instrumental-convergence claims is still real and observable.

This section records the race-and-first-mover logic as one of the clearest didactic threads in his writing and recent public statements. It is the part that makes the competitive dynamics themselves a core risk rather than a side issue.

Influence and dissent

Widely credited by figures inside major labs: Sam Altman has said Yudkowsky was “critical in the decision to start OpenAI.” Early introductions helped bring funding to DeepMind. Nick Bostrom’s Superintelligence (2014) drew on the intelligence-explosion framing.

Counter-views are common among researchers who assign lower probabilities to abrupt, uncontrollable takeoff, who believe incremental empirical alignment work is more promising, or who reject the strong form of the orthogonality thesis. Public disagreement has been visible with researchers at OpenAI, Anthropic, and academic groups since at least 2022–2023.

Personal notes (local only)

Open questions I still want answered