MODULE 1.2

💻 Why it starts with code

Self-improvement appears first in software for a concrete reason: the feedback loop closes in seconds. See how the cycle works, why software is faster than biology, how agents run on their own, and why people move up a level instead of disappearing.

write run check if it passed fix ↻ the whole cycle takes seconds—that is what makes code special Each cycle produces objective evidence: it passed or it did not.
6
Topics
~45
Minutes
Beginner
Level
Applied
Type
Module progress0 of 6 · 0%
1

🔁 The fast feedback loop

In software, the work cycle is extremely short: you write, run, check if it passed, and try again —and it happens in seconds. This “fast feedback loop” is the secret ingredient: the shorter the cycle, the more attempts fit into an hour, and the faster something (or someone) learns to improve.

🧭 New here? One term

Feedback loop: a cycle in which you do something, get a clear signal about whether it worked, and use that signal to adjust the next attempt. In code, the signal is objective (the test passed or failed) and immediate—that is why the cycle is so powerful.

Illustration: a fast-spinning code cycle—keyboard, terminal, and passing tests feeding back into each other
Notice what the image emphasizes: it is not “AI thinking,” it is a cycle spinning. Code is where this cycle is fastest and cheapest—that is why self-improvement appears here before anywhere else.

⚡ Short cycle = fast learning

Imagine someone learning darts blindfolded: if they only find out where they hit at the end of the day, they improve slowly. If they see each hit right away, they improve quickly. Code provides that immediate feedback—which is why AI improves faster with it.

Write
propose a solution
Run
execute
Check
passed? an objective signal
Fix ↻
repeat in seconds
2

🧪 Software vs. biology/chemistry

Why does self-improvement start with code, not medicine or chemistry? Because the loop length is different. In the physical world, an experiment (a cell culture, a reaction, a new material) takes weeks or months to produce a result. In software, the same “experimental round” ends in seconds.

Software — seconds test ↻ in seconds thousands of cycles per day Biology / chemistry — weeks set up waitdays/weeks measure few cycles per month

The same number of “attempts,” but vastly different timescales. That is why self-improving AI appears first in code: thousands of experiments fit into the time biology needs for one.

Why code leads

  • ✓Immediate, objective result (passed/failed).
  • ✓Very low cost per attempt.
  • ✓Thousands of attempts can run in parallel.

Why the physical world is slower

  • ✗An experiment takes days, weeks, or months.
  • ✗Each attempt requires materials and lab resources.
  • ✗It cannot be parallelized cheaply like code.
3

🤖 Self-running coding agents

The feedback loop becomes RSI only when something runs it without a human at every step. That is where agents come in: an AI that writes code, runs it, reads the error, fixes it, and tries again—on its own, repeatedly, for hours.

🧭 New here? One term

Agent: an AI that carries out steps on its own instead of only answering a question. It can run code, edit a file, read the result, and decide what to do next—closing the “propose → test → fix” loop without you pressing Enter each time.

1

Write

Proposes a code solution based on the goal.

2

Run and read the error

Runs it, sees the error message, and understands what broke.

3

Fix and repeat

Makes an adjustment, runs it again, and continues until it passes—without waiting for you.

⌨️

Copy-run: watch the loop close for yourself

Goal: watch a chatbot go through “propose → test → fix” on a tiny example. Paste the prompt into any chatbot.

Write a Python function that takes a list of numbers and returns the average. Then RUN the function mentally with the list [2, 4, 6], show the result step by step, find one bug, and fix it—showing the corrected final version.

How to check: a good answer actually “runs” it mentally (shows the sum 12, divides by 3, gets 4), identifies a real error (e.g., division by zero when the list is empty), and provides a fix. If it only says it “looks good” without testing or finding a bug, the loop did not close—and you have just seen the difference between answering and iterating.

4

🧫 Evolutionary agents

One step beyond a single agent: what if you ran many at once? The evolutionary agent proposes several changes in parallel, tests each one, keeps the winner, and uses it as the starting point for the next round—seeking solutions faster than humans could. The category is well established; the exact name cited in the video (“Gemini-guided evolutionary agent”) is a claim made by a channel (unverified).

base idea the current code variation · test variation · test variation · test variation · test variation … keep winner only

How to read it: the left side launches many attempts (cyan, in parallel); testing filters them; the right side keeps only the best—and that winner becomes the “base idea” for the next round. It is the fast feedback loop multiplied.

✅ Established (the category)

  • ✓Evolutionary search (propose → test → select) is a real, long-standing technique.
  • ✓AI is already being used to optimize code and algorithms.

⚠️ Needs verification (the claim)

  • !The exact name/product (“Gemini-guided”) comes from the video—check the source.
  • !“Faster than humans” depends on the task; it is not a general rule.
5

🏢 Inside Anthropic

The clearest example of “soft” RSI in operation today comes from Anthropic itself. The company publicly states that most of the code that goes into its product goes through Claude. The video cites specific figures—useful for context, but they should be treated as claims until checked against a primary source.

🧭 New here? One term

Merge (verb): to officially “join” a new piece of code to the main program. “Merged code” is code that was actually accepted and added to the product—not a draft.

>80%
of merged code written by Claude (May 2026) — according to the video; needs verification
~8×
more code per engineer vs. 2024 — according to the video; needs verification
~5 months
one employee reported going without writing a line by hand — according to the video; needs verification

✅ What is well established here

The robust point is not the exact percentage, but Anthropic’s public statement that most of the code goes through Claude. That is the “soft” loop at work inside a frontier lab: AI is already a central tool helping build AI itself.

6

⚖️ The human's role changes

The inevitable question: “So does the programmer disappear?” The honest answer is no—the role changes in nature. Instead of typing every line, the human directs, reviews, and decides what matters. It is not “human out of the loop”; it is “human above the loop.”

✗ The wrong interpretation

  • ✗“The human has been removed and AI does everything alone.”
  • ✗“There is no need to understand what is being done anymore.”
  • ✗“Review has become optional.”

✓ The right interpretation

  • ✓The human directs: setting the goal and criteria.
  • ✓The human reviews: judging whether the result is good enough.
  • ✓The human decides what matters—and is responsible for it.

💡 The bridge to Track 2

“Human above the loop” only works if we can measure what systems do. That is exactly where Track 2 comes in: the METR curve (how long a task a model can handle) and MirrorCode (rebuilding software in the dark)—the numbers that make this discussion concrete.

Self-check (optional): why does self-improvement appear first in code?

🎯 Module summary

✓
Fast feedback loop — write → run → observe → fix, in seconds, with an objective signal.
✓
Software × biology — the code loop closes in seconds; the physical-world loop takes weeks. That is why it starts with code.
✓
Agents — AI that writes, runs, reads errors, and fixes them on its own, keeping the loop going for hours.
✓
Evolutionary agents — proposes many variations, tests them, and keeps the winner (established category; exact name needs verification).
✓
Inside Anthropic — most code goes through Claude (well established); exact figures are from the video (needs verification).
✓
Human above, not out — from typing to directing, reviewing, and deciding what matters.

Next track:

Track 2 — The Evidence: the METR curve (task horizon) and MirrorCode. The numbers that make the warning concrete.