Longitudinal System Risk Evaluation
Evaluate specific capabilities repeatedly to expose a changing risk trend
- Difficulty
- Advanced
- Time to result
- ~ongoing to results
- Steps
- 5
- Confidence
- 93%
Bengio argues that risk should be evaluated for specific AI systems across identified categories and then tracked through successive versions. A single score is only a snapshot; repeating the same evaluation reveals whether capabilities linked to cyber misuse, autonomy, or other harms are rising. He says companies conduct some evaluations themselves and also hire external organizations, while European regulation is beginning to require risk assessment. The framework's mechanism is longitudinal comparison: define the categories, measure each release, preserve comparable results, and expose the direction of travel to decision-makers and the public. The transcript gives examples of categories and a claimed rise between systems, but it does not provide the underlying test data, so those examples should remain attributed to Bengio.
Origin
Bengio explains version-by-version AI risk tracking near the end of The Diary of a CEO.
Core principles
- 01Evaluate the specific system rather than AI in the abstract
- 02Track multiple risk categories separately
- 03Repeat evaluations as capabilities change
- 04Combine internal and independent assessments
- 05Make trends visible to regulators and the public
How to run it
- 1
Fix the unit of evaluation
Identify the exact system, version, access conditions, and tools being assessed.
Pro tip Keep versions separate even when they share a product name.
Watch out A general statement about AI cannot substitute for a system-specific test.
- 2
Define risk categories
Choose capabilities or behaviors connected to plausible harms and specify how each will be tested.
Pro tip Keep categories separate so one strong area cannot hide another.
Watch out A category label without a repeatable test produces no usable trend.
- 3
Use more than one evaluator
Combine company testing with external independent assessment where possible.
Pro tip Document material differences in access and methods.
Watch out External does not automatically mean independent or high quality.
- 4
Compare through time
Run comparable tests on later versions and display changes by category rather than only the newest result.
Pro tip Preserve prior baselines and test definitions.
Watch out Changed tests can create an artificial trend.
- 5
Act on the trend
Escalate safeguards, access limits, or oversight when risk-related capabilities rise, and communicate uncertainty clearly.
Pro tip Tie thresholds to actions before seeing the result.
Watch out An evaluation indicates measured capability; it does not guarantee a real-world incident.
In the wild
An evaluator measures cyber assistance, autonomous task completion, and replication behavior for two model versions under the same access conditions. The newer version improves on ordinary tasks but also scores higher on autonomous execution. The lab preserves both results and requires stronger access controls before release.
→ A category-level trend triggers a predefined mitigation instead of being buried in an overall score.
Common mistakes
Publishing only the latest snapshot
Without comparable earlier results, decision-makers cannot see whether a risk-related capability is accelerating.
Changing the test silently
A new protocol may be useful, but its result should not be presented as a clean historical comparison.
Is it for you?
Best for
It is best for regulators, laboratories, and independent evaluators monitoring successive releases of capable systems.
Not ideal for
It is not ideal when evaluations lack repeatability or have no meaningful connection to real-world harm.
From the transcript
“It is really important that we evaluate the risks that specific systems”
“We need those evaluations and we need to keep track of their evolution so that we see the trend”
From the episode
Creator of AI: We Have 2 Years Before Everything Changes! These Jobs Won't Exist in 24 Months!