TThe Diary of a CEO
← All frameworks
Innovation

Capability-Control Gap

Compare how fast a system improves with how fast control improves

Difficulty
Moderate
Time to result
~ongoing to results
Steps
5
Confidence
93%

This mental model evaluates a technology on two trajectories: how quickly its capabilities improve and how quickly people improve their ability to control, predict, and explain it. Yampolskiy argues that AI capability progress is far faster than safety progress, while current protections often behave like patches that users can route around. The useful signal is therefore not whether safeguards exist, but whether they remain effective as the system becomes more capable. Track both rates, stress-test restrictions, and treat a widening gap as evidence for tighter scope or slower deployment. The model does not calculate a precise probability of catastrophe; it identifies a structural risk that headline capability benchmarks can conceal.

Origin

Extracted from The Diary of a CEO

Core principles

  • 01Capability and controllability are separate dimensions
  • 02Safety patches do not prove durable control
  • 03A widening gap raises risk even when both sides improve
  • 04Unpredictability and weak explanations are control deficits

How to run it

  1. 1

    Define the capability frontier

    List the important tasks the system can perform now and how quickly that frontier is moving.

    Pro tip Use observed evaluations rather than promotional labels such as AGI.

    Watch out Do not treat strength in one narrow domain as strength everywhere.

  2. 2

    Define the control frontier

    Assess whether operators can reliably constrain behaviour, predict outcomes, and explain decisions.

    Pro tip Separate each control property instead of collapsing them into one safety score.

    Watch out A written policy is not evidence that the system cannot bypass it.

  3. 3

    Adversarially test safeguards

    Look for jailbreaks, indirect routes, and new subdomains where prohibited behaviour can reappear.

    Pro tip Retest after material capability changes.

    Watch out A patch that blocks one prompt may only move the behaviour elsewhere.

  4. 4

    Compare rates of progress

    Determine whether control improvements are keeping pace with capability improvements over time.

    Pro tip Focus on the trend, not a single reassuring snapshot.

    Watch out Both lines can improve while the gap still widens.

  5. 5

    Constrain the exposure

    When the gap widens, reduce autonomy, connectivity, generality, or deployment scope until controls catch up.

    Pro tip Prefer bounded applications that preserve human authority.

    Watch out This framework diagnoses risk; it does not prove any particular outcome.

In the wild

Illustrative agent release review

A team finds that its agent now completes twice as many open-ended tasks, but its explanation quality is unchanged and a new jailbreak defeats the latest restriction. Using the model, the team records capability growth without corresponding control growth and keeps the agent off unrestricted production access.

The deployment is narrowed until independent control tests improve.

Common mistakes

Counting patches as solved control

A local fix can suppress one failure without resolving the underlying ability to route around restrictions.

Comparing levels instead of rates

A system may look safer than before while capability is improving much faster than control.

Is it for you?

Best for

It is best for evaluating fast-improving AI systems whose safeguards may lag behind their abilities.

Not ideal for

It is not ideal for systems with fixed, fully specified behaviour and no meaningful capability growth.

From the transcript

While progress in AI capabilities is exponential or maybe even hyper exponential, progress in AI safety is linear or constant.

Roman Yampolskiy · (07:00)

The gap between the the how capable the systems are and how well we can control them, predict what they're going to do, explain their…

Roman Yampolskiy · (07:00)

From the episode

Roman Yampolskiy: These Are The Only 5 Jobs That Will Remain In 2030 & Proof We're Living In a Simulation!