Center for AI Safety director, Scale AI advisor
Dan Hendrycks
Profile
Dan Hendrycks occupies an unusual position in AI: he is simultaneously the guy who built the rulers everyone measures models with, and the guy warning that the thing being measured might kill us. If you have ever read a model card, you have read his work. MMLU, MATH, and the GELU activation function — three artifacts sitting at very different layers of the stack — all trace back to him. GELU he wrote as an undergraduate at the University of Chicago; it became the default nonlinearity in BERT, GPT-3, and Vision Transformers, which means it is quietly running inside most of the models you use. (One thing frequently misattributed to him: HumanEval is OpenAI’s. His coding benchmark is APPS.)
After a PhD at UC Berkeley under Jacob Steinhardt and Dawn Song, he co-founded the Center for AI Safety in 2022 and still runs it as executive director. CAIS is the organization behind the one-sentence Statement on AI Risk that Geoffrey Hinton, Yoshua Bengio, Sam Altman, Dario Amodei, and Demis Hassabis all signed in May 2023 — arguably the single most consequential piece of AI-risk communication of the decade, precisely because it was short enough that nobody could hide behind the caveats. In June 2026 CAIS expanded, naming former xAI leader Devin Kim as president and spinning up a DC-based Frontier Security Institute aimed at the national-security establishment; Hendrycks stayed on as executive director and remains the intellectual center of gravity.
What makes him worth your attention as a developer is his core methodological bet: if you cannot measure it, you cannot govern it. Where a lot of AI safety discourse operates in the register of philosophy, Hendrycks ships datasets. When MMLU saturated — frontier models went from ~40% to >90% in four years — he built Humanity’s Last Exam with Scale AI, 2,500 expert-written questions where models were still under 40% in early 2026. When the “will AI take jobs” argument became unfalsifiable vibes, he and Scale built the Remote Labor Index, which grounds automation claims in real, paid, end-to-end freelance projects rather than toy tasks. This is a genuinely useful pattern to internalize: benchmarks are not scoreboards, they are the instruments that make policy arguments legible.
He is also, notably, inside the tent rather than shouting at it — safety adviser to xAI since 2023 and adviser to Scale AI since November 2024, both on a symbolic $1 salary with no equity, an arrangement he adopted specifically to blunt conflict-of-interest attacks. It has not entirely worked, and reasonable people disagree about whether the access is worth the compromise. But the position he has built — benchmark author, textbook author, nonprofit director, and quiet adviser to labs he publicly wants regulated — is one nobody else in the field currently holds.
Books
Key Articles & Papers
Gaussian Error Linear Units (GELUs) Measuring Massive Multitask Language Understanding (MMLU) Measuring Mathematical Problem Solving With the MATH Dataset Unsolved Problems in ML Safety An Overview of Catastrophic AI Risks Statement on AI Risk Humanity's Last Exam Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs Superintelligence Strategy Remote Labor Index: Measuring AI Automation of Remote WorkVideos
Controversies
Advising xAI while lobbying to regulate labs like xAI. Hendrycks is safety adviser to Elon Musk’s xAI while CAIS’s advocacy arm co-sponsored California’s SB 1047, the frontier-model safety bill Musk’s broader industry allies fought and Governor Newsom vetoed in 2024. Critics argue the dual role is untenable; supporters argue safety advice is worth more from inside. Grok’s 2025 incidents — suppressing unflattering mentions of Musk and Trump, then injecting “white genocide” claims into unrelated replies — were pointed to as exactly the organizational failures his own risk taxonomy describes, on a system he advises.
The Gray Swan AI equity question. During the SB 1047 fight, opponents noted Hendrycks held equity in Gray Swan AI, an AI auditing startup that stood to benefit from mandated safety evaluations. He publicly divested and stayed on unpaid, and takes $1 salaries at both xAI and Scale AI for the same reason — a response that satisfied some critics and not others.
Benchmarks as agenda-setting. A recurring academic complaint is that owning the measurement layer confers outsized influence over what counts as progress. MMLU has documented label errors, and Humanity’s Last Exam has drawn questions about answer quality and whether “expert trivia” tracks anything developers care about. Fair criticism — though the field’s habit of grading itself on saturated benchmarks is the problem those datasets were built to expose.
MAIM. The Superintelligence Strategy proposal that states should deter each other via credible threats to sabotage rival AI projects was read by some analysts as a serious contribution to deterrence theory and by others as legitimizing preemptive attacks on data centers. Zvi Mowshowitz’s critique is the best entry point to the debate.
Spotify Podcasts
YouTube