TThe Diary of a CEO
← All frameworks
Innovation

Longitudinal System Risk Evaluation

Evaluate specific capabilities repeatedly to expose a changing risk trend

Difficulty
Advanced
Time to result
~ongoing to results
Steps
5
Confidence
93%

Bengio argues that risk should be evaluated for specific AI systems across identified categories and then tracked through successive versions. A single score is only a snapshot; repeating the same evaluation reveals whether capabilities linked to cyber misuse, autonomy, or other harms are rising. He says companies conduct some evaluations themselves and also hire external organizations, while European regulation is beginning to require risk assessment. The framework's mechanism is longitudinal comparison: define the categories, measure each release, preserve comparable results, and expose the direction of travel to decision-makers and the public. The transcript gives examples of categories and a claimed rise between systems, but it does not provide the underlying test data, so those examples should remain attributed to Bengio.

Origin

Bengio explains version-by-version AI risk tracking near the end of The Diary of a CEO.

Core principles

  • 01Evaluate the specific system rather than AI in the abstract
  • 02Track multiple risk categories separately
  • 03Repeat evaluations as capabilities change
  • 04Combine internal and independent assessments
  • 05Make trends visible to regulators and the public

How to run it

  1. 1

    Fix the unit of evaluation

    Identify the exact system, version, access conditions, and tools being assessed.

    Pro tip Keep versions separate even when they share a product name.

    Watch out A general statement about AI cannot substitute for a system-specific test.

  2. 2

    Define risk categories

    Choose capabilities or behaviors connected to plausible harms and specify how each will be tested.

    Pro tip Keep categories separate so one strong area cannot hide another.

    Watch out A category label without a repeatable test produces no usable trend.

  3. 3

    Use more than one evaluator

    Combine company testing with external independent assessment where possible.

    Pro tip Document material differences in access and methods.

    Watch out External does not automatically mean independent or high quality.

  4. 4

    Compare through time

    Run comparable tests on later versions and display changes by category rather than only the newest result.

    Pro tip Preserve prior baselines and test definitions.

    Watch out Changed tests can create an artificial trend.

  5. 5

    Act on the trend

    Escalate safeguards, access limits, or oversight when risk-related capabilities rise, and communicate uncertainty clearly.

    Pro tip Tie thresholds to actions before seeing the result.

    Watch out An evaluation indicates measured capability; it does not guarantee a real-world incident.

In the wild

Illustrative release comparison

An evaluator measures cyber assistance, autonomous task completion, and replication behavior for two model versions under the same access conditions. The newer version improves on ordinary tasks but also scores higher on autonomous execution. The lab preserves both results and requires stronger access controls before release.

A category-level trend triggers a predefined mitigation instead of being buried in an overall score.

Common mistakes

Publishing only the latest snapshot

Without comparable earlier results, decision-makers cannot see whether a risk-related capability is accelerating.

Changing the test silently

A new protocol may be useful, but its result should not be presented as a clean historical comparison.

Is it for you?

Best for

It is best for regulators, laboratories, and independent evaluators monitoring successive releases of capable systems.

Not ideal for

It is not ideal when evaluations lack repeatability or have no meaningful connection to real-world harm.

From the transcript

It is really important that we evaluate the risks that specific systems

Yoshua Bengio · (1:20:00)

We need those evaluations and we need to keep track of their evolution so that we see the trend

Yoshua Bengio · (1:21:00)

From the episode

Creator of AI: We Have 2 Years Before Everything Changes! These Jobs Won't Exist in 24 Months!