MODULE 2.2

🪞 MirrorCode: rebuilding in the dark

If the Module 2.1 curve measures how long AI can keep working, MirrorCode measures what size project it can finish. The test is demanding: AI receives only the finished program and documentation—never the source code—and must rebuild the entire software. Here you will see the numbers and learn to separate solid evidence from figures that need checking.

Black box executable + docs (no source code) AI rebuilds Rebuilt code module functions commands / API does it pass the tests? Original test suite measures % of identical behavior The score is not “looks right”; it measures how closely the rebuilt software behaves like the original.
6
Topics
~45
Minutes
Intermediate
Level
Evidence
Type
Module progress0 of 6 · 0%
1

🪞 What is MirrorCode

The MirrorCode is a test created by Epoch AI in partnership with METR. The idea is simple and unforgiving: AI receives a finished program—only the executable and documentation—and must rebuild the software from scratch without ever seeing the original code. The reconstruction is then compared with the real program’s test suite. It complements the previous module’s curve: not “how long AI works,” but “what size project it can finish.”

🆕 New here? Two terms before you continue

  • Benchmark: a standardized test that everyone runs the same way to compare models fairly—like an exam with the same answer key for every candidate.
  • Black box: when you only have a system’s inputs and outputs (the executable + documentation), but did not not its internal source code. Rebuilding a black box means observing its behavior and rewriting what produces that behavior.
Who made it
Epoch AI + METR
Input
executable + docs
Task
rebuild from scratch
Score
% of tests passed
2

🎯 The blunt question

The question behind the benchmark is direct: what is the largest software project AI can finish on its own? “Alone” and “finish” are the key words—it is not about completing a snippet, but delivering a working program end to end. To answer this, MirrorCode uses 25 real programs, not lab toys.

The question can it finish alone? bioinformatics Unix utilities cryptography interpreters … 25 in total Scoreboard how many it finished

The 25 programs come from different domains on purpose: a model that is good at just one cannot inflate the score. The final number is “how many AI handled alone.”

🧬

Bioinformatics

tools that process biological data.

🐧

Unix utilities

classic command-line programs.

🔐

Cryptography

where one wrong detail breaks everything.

⚙️

Interpreters

software that runs another language.

3

🌳 The gotree case

The example that puts a face on the benchmark is gotree, a bioinformatics toolkit. It is not small: about 16,000 lines of Go and more than 40 commands. AI rebuilt it well enough to pass almost all tests—in a fraction of the time and cost a human would need.

gotree — reconstruction score

language: Go · ~16,000 lines · 40+ commands
tests_passed: 99,95% (according to the video—unverified)
AI_time: ~14 h (according to the video—unverified)
cost: ~US$251 (according to the video—unverified)
human_estimate: 2 to 17 weeks (Epoch—needs verification)

The numbers are impressive—and for that very reason, they deserve to be checked against a primary source before becoming headlines.

💡 Why 99.95% and not 100%?

Rebuilding large software and passing almost all tests shows real capability. But the missing fraction is often where the hard case lies—the rare detail an experienced human would catch. “Almost everything” is not “everything,” and that difference matters in production software.

4

⏳ The 19-day run

The finding that most directly connects this track to the warning is the run duration: one run operated continuously, with no human at the controls, for about 19 days (≈ US$2,600 in compute). The Module 2.1 curve moves off the chart and into practice: not a burst lasting minutes, but autonomy sustained for days.

1

The task begins

AI receives the black box and the goal, and starts work without a human defining every step.

2

Runs for days

It writes, tests, fixes, and tries again—the code loop closing thousands of times on its own.

3

Delivers after ~19 days

What was an “hours-long agent” becomes a “long-term worker”—the missing piece that makes RSI plausible. (Figures from the video—needs verification.)

🔗 Why this matters for the warning

Recursive self-improvement does not need an instant genius. It needs a worker that can handle R&D projects lasting days. An unsupervised 19-day run is concrete evidence that this extended autonomy already exists—even if it is expensive and imperfect.

5

📊 56% solve rate

The best model solves about 56% of MirrorCode — it leads the ranking, but falls far short of a perfect score. The leap is what is striking: a year earlier, the best models were near 30%, on simpler programs. An honest reading holds both sides: this is a major advance, but it does not mean “AI already replaces engineers.” (The version name and 56% figure are from the video—needs verification.)

🆕 New here? “Solve rate”

It is the share of benchmark tasks that a model solves completely. A 56% solve rate means it successfully finished just over half of all programs in the test. It does not mean “56% of each program”—it means “just over half of the programs, in full.”

The one-year jump (figures need verification)

one year earlier · simpler programs~30%
now · 25 real programs~56%

Illustrative recreation, not a screenshot. The point is the slope, not the exact decimal.

✓ Correct reading

  • ✓“There was a major capability jump in one year.”
  • ✓“AI already finishes more than half of real projects alone.”
  • ✓“The trend deserves informed attention.”

✗ Exaggerated reading

  • ✗“AI already replaces engineers in everything.”
  • ✗“56% is the final, definitive number.”
  • ✗“The remaining 44% is a minor detail.”
6

⚖️ Solid evidence vs. hype in MirrorCode

Closing this track with the course method: the concept and the benchmark are well established—they exist, are public, and measure something real. By contrast, exact figures and version names come from a single source: treat them as claims to verify against a primary source. Conflating the two is the mistake this course aims to prevent.

🟢 Well established (real category)

  • ✓MirrorCode exists (Epoch AI + METR) and measures software reconstruction.
  • ✓AI rebuilds large, real projects from a black box.
  • ✓There was a clear capability jump in ~1 year.
  • ✓Autonomy lasting days (long horizon) is already demonstrable.

🟡 Needs verification (single-source claim)

  • ⚠The exact “56% solve rate” figure.
  • ⚠The gotree figures: 99.95%, 14 h, US$251.
  • ⚠The ~19 days and ~US$2,600 for the continuous run.
  • ⚠Exact name/version of the leading model.

⌨️ Copy-run example: become the skeptic

Goal: use a chatbot to generate questions you would ask the primary source before believing the numbers—building the habit of checking before sharing.

You are a rigorous skeptic. Given this summary of the MirrorCode benchmark: <paste the summary you read here—replace this> List 5 questions I should ask the PRIMARY SOURCE before believing the numbers. Focus on: how the sample was selected, exactly what counts as “AI finished it alone,” the actual compute cost, and what the failed 44% have in common.

How to check: the 5 questions must challenge sampling, cost and the definition of “alone”. If they are generic (“is it reliable?”), rewrite the prompt to request specific, testable questions.

Self-check (optional): in MirrorCode, what is WELL ESTABLISHED and what NEEDS VERIFICATION?

🎯 Module summary

✓
MirrorCode — an Epoch AI + METR benchmark: rebuilding software from a black box.
✓
The blunt question — what is the largest project AI can finish alone; 25 real programs.
✓
The gotree case — ~16,000 lines rebuilt, 99.95% of tests (figures to verify).
✓
19 days continuous — autonomy for days: the Module 2.1 curve in practice.
✓
56% solve rate — it leads and makes a major leap, but does not replace engineers across the board.
✓
Solid vs. hype — the concept and benchmark are solid; figures and versions need verification.

Next track:

Track 3 — Safety and 2028: what happens when measurement itself breaks, and the race shaping scenarios for the end of the decade.

Illustration: AI rebuilds software from a black box, in blue and cyan over a dark grid background
MirrorCode means “rebuilding in the dark”: AI sees only external behavior (the black box) and must recreate the system internally. The larger the project it finishes alone, the closer it is to an R&D worker—the kind of capability this warning calls RSI.