🪞 What is MirrorCode
The MirrorCode is a test created by Epoch AI in partnership with METR. The idea is simple and unforgiving: AI receives a finished program—only the executable and documentation—and must rebuild the software from scratch without ever seeing the original code. The reconstruction is then compared with the real program’s test suite. It complements the previous module’s curve: not “how long AI works,” but “what size project it can finish.”
🆕 New here? Two terms before you continue
- Benchmark: a standardized test that everyone runs the same way to compare models fairly—like an exam with the same answer key for every candidate.
- Black box: when you only have a system’s inputs and outputs (the executable + documentation), but did not not its internal source code. Rebuilding a black box means observing its behavior and rewriting what produces that behavior.
🎯 The blunt question
The question behind the benchmark is direct: what is the largest software project AI can finish on its own? “Alone” and “finish” are the key words—it is not about completing a snippet, but delivering a working program end to end. To answer this, MirrorCode uses 25 real programs, not lab toys.
The 25 programs come from different domains on purpose: a model that is good at just one cannot inflate the score. The final number is “how many AI handled alone.”
Bioinformatics
tools that process biological data.
Unix utilities
classic command-line programs.
Cryptography
where one wrong detail breaks everything.
Interpreters
software that runs another language.
🌳 The gotree case
The example that puts a face on the benchmark is gotree, a bioinformatics toolkit. It is not small: about 16,000 lines of Go and more than 40 commands. AI rebuilt it well enough to pass almost all tests—in a fraction of the time and cost a human would need.
gotree — reconstruction score
The numbers are impressive—and for that very reason, they deserve to be checked against a primary source before becoming headlines.
💡 Why 99.95% and not 100%?
Rebuilding large software and passing almost all tests shows real capability. But the missing fraction is often where the hard case lies—the rare detail an experienced human would catch. “Almost everything” is not “everything,” and that difference matters in production software.
⏳ The 19-day run
The finding that most directly connects this track to the warning is the run duration: one run operated continuously, with no human at the controls, for about 19 days (≈ US$2,600 in compute). The Module 2.1 curve moves off the chart and into practice: not a burst lasting minutes, but autonomy sustained for days.
The task begins
AI receives the black box and the goal, and starts work without a human defining every step.
Runs for days
It writes, tests, fixes, and tries again—the code loop closing thousands of times on its own.
Delivers after ~19 days
What was an “hours-long agent” becomes a “long-term worker”—the missing piece that makes RSI plausible. (Figures from the video—needs verification.)
🔗 Why this matters for the warning
Recursive self-improvement does not need an instant genius. It needs a worker that can handle R&D projects lasting days. An unsupervised 19-day run is concrete evidence that this extended autonomy already exists—even if it is expensive and imperfect.
📊 56% solve rate
The best model solves about 56% of MirrorCode — it leads the ranking, but falls far short of a perfect score. The leap is what is striking: a year earlier, the best models were near 30%, on simpler programs. An honest reading holds both sides: this is a major advance, but it does not mean “AI already replaces engineers.” (The version name and 56% figure are from the video—needs verification.)
🆕 New here? “Solve rate”
It is the share of benchmark tasks that a model solves completely. A 56% solve rate means it successfully finished just over half of all programs in the test. It does not mean “56% of each program”—it means “just over half of the programs, in full.”
The one-year jump (figures need verification)
Illustrative recreation, not a screenshot. The point is the slope, not the exact decimal.
✓ Correct reading
- ✓“There was a major capability jump in one year.”
- ✓“AI already finishes more than half of real projects alone.”
- ✓“The trend deserves informed attention.”
✗ Exaggerated reading
- ✗“AI already replaces engineers in everything.”
- ✗“56% is the final, definitive number.”
- ✗“The remaining 44% is a minor detail.”
⚖️ Solid evidence vs. hype in MirrorCode
Closing this track with the course method: the concept and the benchmark are well established—they exist, are public, and measure something real. By contrast, exact figures and version names come from a single source: treat them as claims to verify against a primary source. Conflating the two is the mistake this course aims to prevent.
🟢 Well established (real category)
- ✓MirrorCode exists (Epoch AI + METR) and measures software reconstruction.
- ✓AI rebuilds large, real projects from a black box.
- ✓There was a clear capability jump in ~1 year.
- ✓Autonomy lasting days (long horizon) is already demonstrable.
🟡 Needs verification (single-source claim)
- ⚠The exact “56% solve rate” figure.
- ⚠The gotree figures: 99.95%, 14 h, US$251.
- ⚠The ~19 days and ~US$2,600 for the continuous run.
- ⚠Exact name/version of the leading model.
⌨️ Copy-run example: become the skeptic
Goal: use a chatbot to generate questions you would ask the primary source before believing the numbers—building the habit of checking before sharing.
How to check: the 5 questions must challenge sampling, cost and the definition of “alone”. If they are generic (“is it reliable?”), rewrite the prompt to request specific, testable questions.
Self-check (optional): in MirrorCode, what is WELL ESTABLISHED and what NEEDS VERIFICATION?
🎯 Module summary
Next track:
Track 3 — Safety and 2028: what happens when measurement itself breaks, and the race shaping scenarios for the end of the decade.