MODULE 2.1

📊 The METR curve: the task horizon

“AI is getting better” is an opinion. METR turns it into a number: the largest task—measured by how long a human would take—that a model completes alone. This module shows that number rising, teaches you to read the chart honestly, and separates a solid trend from a prophecy.

4 min Mar 2024 ~1.5 h late 2024 ~12 h 2025 16 h+ 2026 · test limit Each step is a real measurement—not a guess. What is alarming is the slope, not just the height.
6
Topics
~45
Minutes
Intermediate
Level
Evidence
Type
Module progress0 of 6 · 0%
1

📏 What is a “task horizon”?

There is a way to turn the vague phrase “AI is getting better” into a trackable number. That number is the task horizon: the largest task—measured by how long a human would take to complete it—that the model finishes with about 50% success. If AI completes tasks alone that would take a human 12 hours, we say its horizon is 12 hours.

🆕 New here? Three terms before you continue

  • METR: an independent organization (not a lab or model maker) that evaluates the capabilities and risks of frontier models — the most advanced models from labs such as Anthropic, OpenAI, and DeepMind. METR produces the curve.
  • Task horizon: the measuring stick: how much human work, in time, AI can handle from start to finish.
  • 50% success rate: the cutoff point. It does not mean “always gets it right”; it means “gets it right half the time”—where the task starts to become too large for the model.

🎯 Why measure in human time

Measuring in “how long a human would take” is powerful because everyone understands that unit. “Solves 56% of questions” says nothing about the size of the task. “Does alone what would take you a day” says it all. The curve does not measure abstract intelligence—it measures useful autonomy.

Who measures
METR (independent)
The measure
human time
The cutoff
~50% success
Measures
autonomy, not IQ
2

📈 The curve: 4 min → ~1.5 h → ~12 h → 16 h+

From March 2024 to 2026, the horizon measured by METR grew from a few minutes and reached more than half a day of work. The staircase at the top of the page shows it: 4 min, then ~1.5 h, ~12 h, and 16 h+. What matters is not each step on its own, but the slope: the measure doubles over short intervals.

short tasks long tasks → horizon (~50% success) success failure

How to read it: on the left (short tasks), almost everything succeeds; on the right (long tasks), almost everything fails. The dashed line marks where AI succeeds half the time—that point, measured in human time, is the horizon.

💡 Reading tip

Do not memorize the numbers—memorize the shape. A curve that doubles every few months is exponential, and exponential growth fools intuition, which expects straight lines. That is why the horizon “suddenly” went from minutes to hours.

3

🧱 16 h was the TEST limit, not the model's

This is where the most common misreading happens. METR itself warned that above 16 hours, measurements from that suite were no longer reliable. In other words, “16 h” is the ceiling of the instrument — how far the test can measure reliably—not the limit of what the model can do.

⚠️ Beware the headline

Reading “16 h” as “the maximum AI can do” is the opposite of what the data says. The model may go further—we just cannot claim that from this test. Treating a measurement ceiling as a capability ceiling is like saying a car can only go 200 km/h because that is where the speedometer ends.

✓ Honest reading

  • ✓“The test measures reliably up to ~16 h.”
  • ✓“The trend continues upward.”
  • ✓“We need better tests beyond that.”

✗ Misleading reading

  • ✗“AI stops at 16 h.”
  • ✗“That is the definitive limit.”
  • ✗“It hit the ceiling; we can relax.”
4

⏱️ Why long horizons matter for RSI

Recursive self-improvement (RSI)—AI helping design the next AI—does not require a genius who solves everything instantly. It requires a useful worker for hours and days at a time: one that keeps track of the problem, runs experiments, waits for results, and continues. That is why long horizon (long horizon) is the piece that connects this track to the entire warning.

Short horizon × long horizon

⏱️
Short: answering a question, writing a function, fixing a bug. Useful, but each is a standalone task.
⏳
Long: running an entire project for days—planning, executing, testing, adjusting. That is what an “AI engineer” does, and what RSI needs to automate.

🔗 The bridge to the next module

Keep this idea in mind: if AI can already handle hours of continuous work, the next question is, “what project size can it finish on its own?” That is exactly what MirrorCode measures—and where we will see a run operating for days without a human (Module 2.2).

RSI needs
autonomy for days
It does not need
an instant genius
The curve shows
this drawing closer
Next
project size
5

🔬 How it is measured (and why it is hard)

Here is how: gather many tasks with known durations (we know how long each would take a human), run the model on all of them, and find the duration at which it succeeds about half of the time. That point is the horizon. Easy to describe, hard to measure well—for three reasons.

1

Variance

The same task can succeed in one run and fail in another. That is why it is run many times and results are averaged—not based on a single test.

2

Estimating human time

How many person-hours is a task “worth”? It is an estimate, and estimates have a margin of error. Two evaluators may disagree.

3

The test reaches its own limit

Near the top (the 16 h in Topic 3), there are not enough sufficiently long, well-measured tasks. The instrument runs out before the model does.

⌨️ Copy-run example · feel the horizon in your own work

Goal: train your eye to distinguish short-horizon from long-horizon tasks using your own workday. Paste the block below into any chatbot.

List 5 tasks I do in <my job—replace this> and estimate how long each would take a human. Then, for each one, say whether an AI agent today could do it ALONE for hours on end, and what would still require a human at the controls. Separate “short horizon” from “long horizon.”

How to check: the answer should distinguish tasks lasting minutes (short horizon) from tasks lasting hours or days (long horizon)—not just say “AI does everything.” If it does not separate the two, ask again.

6

🧭 What the curve does NOT prove

The curve is strong evidence of a trend. It is not proof of destiny. Extending the chart’s line and declaring “AI will soon do months-long tasks” swaps measurement for prophecy. The course’s honest stance is to take the data seriously and mark where it ends and speculation begins.

🟢 Well supported (the curve shows)

  • ✓The task horizon has risen consistently.
  • ✓From 2024 to 2026, it grew from minutes to more than half a day.
  • ✓METR is an independent source with a transparent method.

🟡 Needs verification / speculation

  • ▲“The line continues unchanged and reaches tasks lasting weeks.”
  • ▲An exact date for the horizon to cross X hours.
  • ▲That today's pace is guaranteed to continue tomorrow.

💡 The course's golden rule

Observed trend = solid. Arrival date = guess. Whenever someone combines the two in one sentence, separate them before you believe it. The curve says “it is rising fast”—not “it will get there in a certain year.”

Module 2.1 illustration: the rising task-horizon curve climbing in steps from minutes to hours, in blue and cyan against a dark background
The image reinforces the main idea: each step is a measured increase in capability, not an imagined one. Look at the slope—it is what supports the warning, not any single number.

Self-check (optional): what does “16 hours” mean on the METR curve?

🎯 Module summary

✓
Task horizon — the largest task, measured in human time, that AI completes with ~50% success.
✓
The curve — from 4 min to 16 h+ between 2024 and 2026; the slope is what is alarming.
✓
16 h = test limit — the instrument's ceiling, not the model's.
✓
Long horizon and RSI — RSI needs a useful worker for days, not an instant genius.
✓
How it is measured — 50% success rate, variance, and a test that reaches its own limit.
✓
What it does NOT prove — the trend is solid; the arrival date is speculation.

Next module:

2.2 — MirrorCode: rebuilding in the dark. From “how long can it keep going?” to “what size project can it finish alone?”