MIRI co-founder, pioneering AI alignment researcher
Eliezer Yudkowsky
Profile
Few people have shaped how the AI field talks about its own dangers as much as Eliezer Yudkowsky — and few are as polarizing. Long before “AI safety” was a funded discipline with departments at every major lab, Yudkowsky was a largely self-taught researcher arguing that a sufficiently capable artificial mind, built without a solved theory of how to align its goals with ours, would by default be lethal to humanity. He co-founded the Machine Intelligence Research Institute (MIRI) in 2000, and for two decades he did the unglamorous conceptual work — defining the alignment problem, naming the failure modes — that much of today’s safety research quietly builds on.
His influence runs deeper than his papers. Through LessWrong, the community blog he founded, and its core essay collection “The Sequences,” Yudkowsky trained a generation of thinkers in a particular style of probabilistic reasoning about the future. A striking number of people now working at OpenAI, Anthropic, and Google DeepMind, as well as the broader effective-altruism and rationalist orbit, first encountered these ideas through his writing — including the 660,000-word rationalist novel Harry Potter and the Methods of Rationality, which became an unlikely recruiting funnel for the field. Fellow travelers and interlocutors like Stuart Russell, Max Tegmark, Nick Bostrom, and Paul Christiano built on, sharpened, or pushed back against his framing.
Yudkowsky’s early warnings were once dismissed as science fiction. That is no longer true. After the arrival of GPT-4, he wrote a widely-read 2023 op-ed in TIME calling not for a pause but for a complete, indefinite, internationally-enforced shutdown of frontier AI training — willing, he argued, to enforce it with airstrikes on rogue data centers if necessary. It is the single most extreme mainstream position in the debate, and he holds it without apology. Where researchers like Yoshua Bengio and Geoffrey Hinton now voice serious concern about existential risk, Yudkowsky’s estimate of doom is far higher and his prescription far more drastic.
Today he has largely stepped back from technical research — MIRI itself pivoted around 2024 from alignment work toward policy and advocacy, judging that alignment is too slow to win the race — and reinvented himself as a communicator. His 2025 book with MIRI president Nate Soares, If Anyone Builds It, Everyone Dies, became a New York Times bestseller and put his argument in front of the general public. For developers, Yudkowsky is essential not because he’s always right — reasonable people think his probabilities are miscalibrated — but because he articulated the hard version of the problem you’re implicitly working on, and confronting his strongest arguments is the fastest way to understand what “alignment” actually demands.
Books
Key Articles & Papers
AGI Ruin: A List of Lethalities Pausing AI Developments Isn't Enough. We Need to Shut It All Down There's No Fire Alarm for Artificial General IntelligenceVideos
Controversies
Yudkowsky is a genuinely contested figure, and developers should weigh the criticism alongside the ideas. His 2023 TIME op-ed suggesting that enforcing an international training ban could justify airstrikes on non-compliant data centers drew heavy criticism as reckless and counterproductive, even from people sympathetic to safety concerns. Critics also point to his lack of conventional academic credentials and peer-reviewed output, arguing that MIRI produced comparatively little concrete technical progress over two decades of funding. Others — including many working engineers and researchers like Yann LeCun — regard his near-certainty of doom as unfalsifiable and his timelines and probabilities as poorly calibrated. Supporters counter that he named the problem years before the mainstream took it seriously, and that being early and blunt is not the same as being wrong. Both readings are worth holding at once.
Spotify Podcasts