TRACK 2

📊 The evidence

Move RSI from opinion to measurement. This track presents the two figures behind the warning: the task-horizon curve from METR—how much human work time AI can already complete on its own—and the MirrorCode, the benchmark that asks what is the largest software project AI can rebuild in the dark. You will learn to distinguish a solid trend from a figure that needs checking.

task duration time (Mar 2024 → 2026) 4 min ~1.5 h ~12 h 16 h+ Benchmark measures the task limit
Track 2 illustration: the rising task-horizon curve beside a benchmark panel, in blue and cyan
This track has two axes: the curve (how long AI can keep working) and the benchmark (what size project it can finish). Together, these are the evidence—the rest of the warning rests on them.
2
Modules
12
Topics
~1h30
Duration
Intermediate
Level
Track 2 progress0 of 12 · 0%

Track map

Detailed content

2.1~45 min · 6 topics

📊 The METR curve: the task horizon

The measure that turns “AI is getting better” into a number: the largest task, measured in human time, that a model completes alone.

0 of 6
What it is:

The largest task—measured by how long a human would take—that the model completes with about a 50% success rate. This is how METR (an independent organization that evaluates the capabilities and risks of frontier models) puts a number on “how capable.”

Why learn this:

Without a measuring stick, “AI has improved” is an opinion. With a task horizon, it becomes a curve we can track month by month.

Key concepts:

METR; task horizon; 50% success rate; human time as the unit.

What it is:

The sequence of measurements from Mar 2024 to 2026: the horizon grows from a few minutes to more than half a workday. Each step is a measurement from the suite, not a guess.

Why learn this:

The shape of the curve is at the heart of the warning: it is not just rising, but how fast it is rising.

Key concepts:

Progression; steps; growth rate; observed trend.

What it is:

METR itself warns that above 16 hours, measurements from that suite were no longer reliable. “16 h” is the instrument’s ceiling, not the model’s limit.

Why learn this:

This is the kind of detail that separates careful reading of a chart from turning a number into a misleading headline.

Key concepts:

Test limit × model limit; measurement ceiling; honest interpretation.

What it is:

Recursive self-improvement does not require solving everything instantly; it requires a useful worker for hours and days at a time. That is exactly what a long horizon means.

Why learn this:

This connects the curve to the course topic and sets up MirrorCode (Module 2.2), where AI runs for days.

Key concepts:

Long horizon; extended autonomy; bridge to RSI.

What it is:

It is measured by running the model on many tasks of known durations and finding where it succeeds about half the time. There is variability, and near the top the test reaches its own limit.

Why learn this:

Understanding the method guards against two mistakes: dismissing the curve as “marketing” or treating it as clockwork precision.

Key concepts:

50% rate; variance; measurement noise; suite limit.

What it is:

Extending the chart’s line into the future does not make the outcome certain. The trend is solid; the arrival date is speculation. The two must not be conflated.

Why learn this:

This is the course’s honest stance: take the evidence seriously without buying the prophecy attached to it.

Key concepts:

Extrapolation; trend × destiny; solid evidence vs. hype.

View full module
2.2~45 min · 6 topics

🪞 MirrorCode: rebuilding in the dark

The benchmark that asks about real scale: what is the largest software project AI can rebuild on its own with only the executable and documentation.

0 of 6
What it is:

A benchmark (a standardized test for comparing models) developed by Epoch AI with METR. The AI receives a program as a “black box”—only the executable and documentation, without source code—and must rebuild the software from scratch.

Why learn this:

It complements the curve: not “how long can it keep going?” but “what size project can it finish?”

Key concepts:

Benchmark; black box; reconstruction; Epoch AI + METR.

What it is:

The test’s direct question: what is the largest software project AI can complete without a human? It covers 25 real programs—from bioinformatics and Unix utilities to cryptography and interpreters.

Why learn this:

“Alone” and “real project” are the two words that separate a demo from serious capability.

Key concepts:

25 real programs; end-to-end task; no human assistance.

What it is:

A bioinformatics toolkit with ~16,000 lines of Go and 40+ commands. AI reimplemented it, passing 99.95% of tests in 14 hours for about US$251—a human would take 2 to 17 weeks (Epoch estimate). Figures are from the video and need verification.

Why learn this:

This concrete example makes the benchmark tangible—and shows why separating solid evidence from claims to verify matters.

Key concepts:

gotree; 99.95% of tests; human-effort estimate; figure to verify.

What it is:

One run operated continuously, with no human at the controls, for about 19 days (≈ US$2,600). It is the Module 2.1 curve made real: autonomy for days, not minutes. Figures are from the video and need verification.

Why learn this:

It shows what “long horizon” means in practice and why it is the missing piece for RSI.

Key concepts:

Continuous run; 19 days; compute cost; long-term worker.

What it is:

The best model solves about 56% of the benchmark—leading the field, but not replacing engineers across the board. A year earlier, top models scored about 30% on simpler programs. Version name and figure need verification.

Why learn this:

At 56%, the result is both impressive and incomplete—this module asks you to consider both sides.

Key concepts:

Solve rate; leadership ≠ perfection; jump from ~30% to ~56%.

What it is:

The concept and benchmark are well established. Exact model versions and specific figures (251, 2,600, 56%) come from a single source—treat them as claims until checked against a primary source.

Why learn this:

This closes the track with a method: take the evidence seriously and label each figure by its level of certainty.

Key concepts:

Solid evidence × claim; primary source; productive skepticism.

View full module
← Track 1 · Fundamentals Track 3 · Safety and 2028 →