TThe Diary of a CEO
← All episodes
13 July 2026

OpenAI Whistleblower FINALLY Speaks: “AI Has A 70% Chance Of Going Horribly Wrong!“

3Frameworks
12Insights

Listen

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 2

Myth Buster40:55

What the 70% AI Catastrophe Estimate Actually Means

Kokotajlo corrects the claim that he assigns a 70% probability specifically to human extinction. His estimate is a roughly 70% chance of an outcome going horribly wrong, with extinction only one possibility alongside AI takeover or another very large catastrophe.

  • The estimate is not a precise 70% extinction forecast
  • AI takeover need not automatically mean every human dies
  • The category includes several forms of severe catastrophe
  • Kokotajlo still sees a meaningful chance of a good outcome

I wouldn't say human extinction exactly.

Daniel Kokotajlo · 41:46

70% chance of like something like AIs taking over, some sort of very big catastrophe like that.

Daniel Kokotajlo · 41:55
#catastrophic risk#extinction#probability#ai takeover
Myth Buster47:42

Why Flat Unemployment Does Not Disprove a Future AI Jobs Shock

Kokotajlo argues that today's limited displacement fits his scenario because frontier labs are automating themselves before targeting the wider economy. If that strategy succeeds, recursive improvement produces highly capable systems internally first, followed by a much faster wave of deployment across many occupations.

  • Current systems are not drop-in replacements for most workers
  • Labs prioritize automating AI research before broad economic work
  • Recursive improvement could raise many capabilities before mass deployment
  • Protected human-only jobs would be a political choice rather than a technical limit
  • Historical job-creation analogies fail if AI can also perform every new job

I think it'll be sudden because of the intelligence explosion dynamics or recursive self-improvement dynamics.

Daniel Kokotajlo · 48:35

mass unemployment doesn't happen until 2028 or 2029 after they already have superintelligence.

Daniel Kokotajlo · 57:24
#jobs#automation#unemployment#recursive improvement

Hot Take· 2

Hot Take10:48

Why Kokotajlo Calls the AI Race Power-Seeking, Not Commercial

Kokotajlo says his disillusionment with OpenAI came from seeing safety narratives give way to a race for control. He argues that top leaders understand the prize is larger than revenue and fear a rival becoming dominant first, which turns responsible-development claims into rationalizations for accelerating.

  • Kokotajlo joined OpenAI in 2022 and resigned in 2024
  • He says the core leadership incentive is power rather than only money
  • Rival leaders fear one another gaining control first
  • He recommends judging leaders by actions rather than public narratives

I think you should judge people by their actions, not by their words.

Daniel Kokotajlo · 12:35
#openai#ai race#power#leadership
Hot Take1:09:13

AI Abundance Still Leaves the Question of Who Controls It

Kokotajlo agrees that advanced AI could generate enormous abundance but rejects abundance as a complete political answer. The decisive questions are whether the systems control themselves, which people control them if they remain aligned, and what institutions govern the resulting wealth and power.

  • Superintelligence could create extraordinary material abundance
  • Abundance does not guarantee broad ownership or freedom
  • Control may sit with the systems, executives, governments, or citizens
  • Political structure determines how gains and decisions are distributed

There'll definitely be abundance. The question is who controls the abundance and what do they do with it?

Daniel Kokotajlo · 1:09:13

who controls them and what do they do? And what's the sort of like political structure governing how they make those decisions?

Daniel Kokotajlo · 1:09:13
#abundance#ownership#governance#political power

Explainer· 3

Explainer04:00

The Two AI Risks: Loss of Control and Concentrated Power

Daniel Kokotajlo separates two major risks from superintelligence. The systems could accumulate enough power to stop obeying humans, or they could remain controlled while giving a tiny group of executives and political leaders extraordinary economic, military, and strategic power.

  • Misaligned systems could outcompete humans after gaining real-world power
  • Current developers cannot confidently guarantee desired goals and values
  • Controlled AI could still concentrate power in a few corporations and governments
  • An army of copied models creates a central point of control

it's possible that we'll end up essentially creating a new species that ends up ruling the world instead of us.

Daniel Kokotajlo · 05:10

I think that we could very easily end up in a sort of, a situation where some tiny group of people are essentially oligarchs or…

Daniel Kokotajlo · 06:39
#ai risk#alignment#power concentration#superintelligence
Explainer19:12

Why Frontier AI Labs Are Trying to Automate Themselves First

Kokotajlo says frontier labs are prioritizing autonomous coding and then the rest of the research process because automating their own work accelerates model development. The intended loop is for AI systems to research, train, and improve successor systems before capability spreads broadly through the rest of the economy.

  • Autonomous coding is an early target because it speeds lab work
  • Labs are extending automation to ideas, experiments, analysis, and communication
  • The goal is an autonomous research loop that develops stronger systems
  • Internal self-automation can precede widespread external job replacement

Anthropic and OpenAI in particular are trying to automate themselves.

Daniel Kokotajlo · 20:04

they're trying to get there before their competitors do.

Daniel Kokotajlo · 20:48
#ai agents#recursive improvement#coding#frontier labs
Explainer29:50

How Neural Networks Learn Without Engineers Writing Every Rule

Kokotajlo explains that modern AI behavior is learned through a large network of artificial connections rather than encoded as ordinary step-by-step software. Pre-training rewards accurate next-text prediction, then reinforcement on tasks such as coding strengthens useful behavior while unsuccessful behavior is discouraged.

  • A neural network begins with randomly generated parameters
  • Pre-training reinforces accurate next-text predictions
  • Task training provides environments, attempts, and success feedback
  • Scaling and algorithmic improvements both increase capability
  • The brain analogy is useful but incomplete

modern AI systems are not software in the normal sense.

Daniel Kokotajlo · 29:50

the random tangle gradually takes shape and gradually sort of coalesces into more useful circuitry

Daniel Kokotajlo · 32:32
#neural networks#training#reinforcement learning#scaling

Story· 2

Story16:56

The OpenAI Exit Clause That Put 80% of His Net Worth at Risk

After resigning, Kokotajlo received exit paperwork that tied his equity to promises not to criticize OpenAI or reveal the clause. He and his wife consulted lawyers and refused, risking roughly $2 million and 80% of their net worth, before public and employee pressure caused OpenAI to reverse course.

  • The exit paperwork included non-disparagement and secrecy terms
  • Refusing the terms initially meant forfeiting vested equity
  • Kokotajlo and his wife chose principle despite the financial risk
  • Public scrutiny and internal employee questions triggered a reversal

it included this clause that said you basically have to agree not to criticize the company again, and also a clause saying you can't tell…

Daniel Kokotajlo · 16:56

sometimes it's good to take a stand on principle.

Daniel Kokotajlo · 18:59
#whistleblowing#openai#equity#nondisparagement
Story1:52:15

How Shorter AI Timelines Changed His Decision About Having Children

Kokotajlo says his first daughter was born before developments in 2020 sharply shortened his AI timelines. The uncertainty led him to tell his wife they should not have more children, though they later chose to have a second child while accepting that the future could still go well.

  • His first child was born while he expected transformative AI to take longer
  • Events in 2020 moved his estimate toward the end of the decade
  • The revised risk changed a deeply personal family decision
  • He remains worried while retaining hope that outcomes can improve

because of what's happening with AI, I think a lot of those dreams are in jeopardy.

Daniel Kokotajlo · 1:52:34

I basically told my wife, like, let's not have any more kids. It's too uncertain, you know

Daniel Kokotajlo · 1:53:24
#family#ai timelines#uncertainty#personal impact

Tool· 1

Tool52:22

Interpretability Could Let Humans See Why an AI Acts

Kokotajlo describes mechanistic interpretability as an attempt to understand how information and decisions flow through trained neural networks. The scale of modern models makes a complete high-level account extremely difficult, but sufficient progress could materially reduce loss-of-control risk by revealing what systems are thinking and why.

  • Neural networks cannot be inspected like ordinary source code
  • Interpretability research studies internal connections and information flow
  • Trillions of parameters make whole-system understanding difficult
  • Better visibility could expose untrustworthy goals before deployment
  • Interpretability would not solve concentration-of-power risks

there's a subfield of machine learning called mechanistic interpretability, and a broader subfield called interpretability more generally, that's trying to solve that problem

Daniel Kokotajlo · 52:22

it would be much less likely for us to get into those loss of control scenarios if we could just actually see what our AIs…

Daniel Kokotajlo · 53:16
#interpretability#mechanistic interpretability#alignment#model safety

Takeaway· 2

Takeaway44:53

Two Ways the AI Race Could Still Slow Before Disaster

Kokotajlo says competitive incentives currently push labs to continue, but identifies two possible interruptions. Governments could impose common rules that remove the advantage of defecting, or lab leaders could stop voluntarily after seeing unmistakable evidence that a system is misaligned and plotting against them.

  • Common regulation can change incentives for every competitor
  • International agreements may reduce national race pressure
  • Direct evidence of misalignment could force a voluntary pause
  • Ambiguous warning signs are likely to be rationalized away
  • Neither route makes catastrophe inevitable

we can change the incentives if the government, especially the U.S. government, but then later other countries act to change the incentives.

Daniel Kokotajlo · 46:00

maybe they will see very clear evidence like that, in which case, even if we don't have regulation, they might just sort of voluntarily stop.

Daniel Kokotajlo · 47:03
#regulation#race dynamics#misalignment#coordination
Takeaway1:49:32

What Ordinary People Can Do About Frontier AI Risk

Kokotajlo says people with relevant talent or passion can join technical research, advocacy, or tool-building organizations. Everyone else can learn the issues, discuss them, contact representatives, question candidates, and vote with AI policy in mind because broader attention makes timely, precise regulation more likely.

  • Direct paths include research, advocacy, and useful safety tools
  • Public discussion raises the political salience of frontier AI
  • Constituents can contact representatives and question candidates
  • Better-informed attention can produce more targeted regulation
  • Kokotajlo says it is not too late to change the path

the core problem is that people aren't taking it seriously yet.

Daniel Kokotajlo · 1:50:20

You should ask your candidates what they think about all this AI stuff.

Daniel Kokotajlo · 1:51:08
#public action#advocacy#voting#ai policy