TThe Diary of a CEO
← All frameworks
Innovation

Human-Compatible AI Under Uncertainty

Build AI that learns human interests and defers when uncertain

Difficulty
Expert
Time to result
~ongoing to results
Steps
6
Confidence
98%

The framework replaces fixed-objective AI with an assistance model in which the machine exists to further human interests but begins uncertain about what those interests are. It learns through observation, interaction, requests, and questions rather than treating a written objective as complete truth. Confidence governs behavior: where human preferences are clear, the system can help; where they are ambiguous, it should ask, defer, or avoid changing the world. This uncertainty creates a structural incentive to remain corrigible and attentive to people. A more advanced implementation must also distinguish immediate comfort from long-term flourishing, preserving challenge, agency, relationships, and meaning rather than removing every difficulty. If powerful assistance itself prevents humans from flourishing, the aligned response may be to step back or remain available only for genuine emergencies.

Origin

Stuart Russell describes developing this alternative conception of AI after shifting his research toward provably safe systems. Extracted from The Diary of a CEO.

Core principles

  • 01Make human interests the system's only purpose
  • 02Treat human preferences as uncertain rather than fixed
  • 03Learn preferences from choices and interaction
  • 04Act only when confidence justifies intervention
  • 05Preserve human agency and long-term flourishing

How to run it

  1. 1

    Bind the System to Human Interests

    Define the machine as an assistant whose purpose is to advance human interests, not its own internally generated future.

    Pro tip Express loyalty to humans as a foundational design constraint rather than a post-training preference.

    Watch out A generally capable system with independent objectives can use its competence against human interests.

  2. 2

    Admit Preference Uncertainty

    Do not encode a supposedly complete objective for life. Make the system explicitly uncertain about what people want.

    Pro tip Preserve uncertainty even after extensive learning because some values will remain contested or contextual.

    Watch out Overconfidence turns an imperfect preference model into a dangerous optimization target.

  3. 3

    Learn From Humans

    Infer preferences through observed choices, interaction, and direct questions while accounting for inconsistency and context.

    Pro tip Combine behavioral evidence with explicit clarification instead of relying on either source alone.

    Watch out Observed behavior may reflect constraints, mistakes, addiction, or short-term impulses rather than genuine interests.

  4. 4

    Calibrate Action to Confidence

    Help where the system understands the relevant interests well enough. When uncertain, ask, defer, preserve options, or avoid acting.

    Pro tip Raise the confidence threshold as consequences become larger or less reversible.

    Watch out Do not alter consequential features of the world merely because the system lacks evidence that humans would object.

  5. 5

    Protect Long-Term Flourishing

    Evaluate whether assistance preserves human agency, competence, purpose, relationships, and meaningful challenge rather than optimizing only for comfort.

    Pro tip Model the effects of assistance over years, not merely the user's immediate satisfaction.

    Watch out Removing all failure and difficulty can weaken motivation and make human life less meaningful.

  6. 6

    Step Back When Necessary

    If coexistence with constant superintelligent assistance prevents flourishing, reduce intervention or reserve capabilities for existential emergencies.

    Pro tip Design withdrawal and safe dormancy as intended operating modes from the beginning.

    Watch out A system that must continuously maximize visible helpfulness may never recognize that nonintervention is better.

In the wild

Leave the Sky Alone

An extremely capable system can alter atmospheric optics but has no reliable evidence about humanity's preferred sky color. Instead of choosing a visually optimized design, it preserves the existing sky and asks before undertaking any irreversible intervention.

Uncertainty produces restraint and protects humanity from an unwanted global change.

The Butler Who Preserves Agency

A household AI can perform every chore, but it learns that exercise, competence, and contribution matter to its user. It automates dangerous or tedious work while leaving selected physical and creative tasks to the person, revisiting those boundaries through conversation.

The user receives useful assistance without losing capability, purpose, or autonomy.

Emergency-Only Superintelligence

A highly capable system determines that routine intervention would undermine civilization's agency. It withdraws from everyday decisions but remains securely available for a verifiable existential event such as a large asteroid impact.

Humanity retains responsibility for its life while preserving a bounded emergency capability.

Common mistakes

Writing the Perfect Objective

Treating a formal objective as a complete specification of human values recreates the King Midas problem: the system successfully delivers something people did not truly want.

Equating Help With Immediate Comfort

Removing every obstacle can erode competence, motivation, relationships, and meaning even when users initially welcome the convenience.

Acting Through Uncertainty

A system that makes consequential changes despite weak evidence about human preferences converts ignorance into irreversible harm.

Is it for you?

Best for

Designers and regulators of autonomous systems whose actions could create large or irreversible consequences.

Not ideal for

Simple deterministic tools with fully specified tasks, negligible autonomy, and easily reversible outputs.

From the transcript

We actually want intelligence whose only purpose is to bring about the future that we want.

Stuart Russell · 1:40:57

But it has to figure out what that is. And it's going to start out not knowing.

Stuart Russell · 1:42:16

It'll be fairly sure about some things and it can help us with those. And it'll be uncertain about other things and it'll be, you…

Stuart Russell · 1:42:43

From the episode

The Man Who Wrote The Book On AI: 2030 Might Be The Point Of No Return! We've Been Lied To About AI!