MODULE 3.1
RSI: AI that helps improve AI
Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back.
What it is
RSI stands for recursive self-improvement: the idea of AI helping improve AI. The area’s central question is practical: how does a system propose changes, test results, and keep what works? The page treats this as a field of study and an open collection of resources, with references, examples, and criteria for evaluating claims. The starting point is a distinction: improving a response, improving an agent, and improving the ability to create new agents are different things. An agent here is an AI system that receives a goal and carries out steps using instructions, memory, and tools.
Why learn it
Claims about AI improving itself come up often and tend to mix these three levels. If you understand the difference, you can ask what exactly improved and what evidence supports the claim.
Key concepts
recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims
What it is
The first level is reviewing the response: the model critiques and revises an output. This can improve a task without changing its parameters—the internal values learned during training—or its development process. The second level is improving the system: the agent’s instructions, memory, tools, or code change, and the versions need to be compared on held-out tasks while keeping a way to roll back. The third level is investigating recursion: the improvement helps produce future improvements. This mechanism needs evidence, because more attempts or a higher score do not prove unlimited growth.
Why learn it
Each level calls for a different kind of control. Reviewing a response is inexpensive and reversible; changing the system calls for version tracking and tests; claiming recursion calls for strong evidence. Mixing up the levels can lead you to accept guarantees no one has measured.
Key concepts
review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back
What it is
The page suggests starting with a small task: answer based on a policy, create questions from a text, or extract action items from notes. First define the task, the correct reference, and the limits for data, spending, and actions. Then record the initial version and separate development cases from cases held out for validation. Propose one change at a time and compare using the same criteria, including errors, cost, and review effort. Finally, validate before adopting: save results and versions, decide who has the final say, and keep a rollback path.
Why learn it
A familiar evaluation can be exploited, the page warns. Keeping tests independent, checking for invented data, and looking at results outside the sample helps you avoid mistaking a local gain for general autonomy.
Key concepts
small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback
What it is
LOOP-R is the project’s framework for organizing improvements: Run, Measure, Critique, Propose, Test, Validate, Promote, and Repeat. The page describes it as ready to use in Claude Code, with nine assistants in separate roles, a record of each version, a spending cap, and a command to roll back. The safety rule is explicit: the system cannot decide to replace the current version with a worse one. In the RSI course, LOOP-R is presented as a conceptual proposal from the project. There is also a LOOP-R course with five tracks and 21 lessons for owners and managers without a technical background.
Why learn it
A cycle without guardrails can replace a good version with a worse one without anyone noticing. A spending cap, version records, and rollback make the cycle auditable and safe to repeat.
Key concepts
Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back
What it is
The “Start here” section has two resources. The RSI Course v6.2 has 18 lessons in six modules, with exercises, reviews, and materials for applying LOOP-R and recording approved knowledge, in Portuguese, English, and Spanish. The RSI Guide maps the topic, including mechanisms, applications, and limits, also in three languages. The page includes an important caveat: the project brings together research and educational content; it is not a ready-to-run autonomous RSI system. The code and materials are in the rsi repository on GitHub.
Why learn it
Starting with the guide gives you the vocabulary and an understanding of the limits; the course turns that into practice with exercises. Knowing the project is educational helps you avoid expecting a tool that improves itself.
Key concepts
RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system
What it is
Along with the guide and course, the resource collection includes three materials with clear caveats. RSI Copiloto is a local assistant with AI, memory, tasks, and routines that compares instructions, asks for human review, and lets you roll back changes; its public demo uses programmed examples, without AI, and real AI is available only in the local version. Dream-RSI is an independent educational guide, not affiliated with Google, about learning from experiment histories; the lab and the eight sessions are future plans, not an implementation of the paper. Alerta IA 2028 is a course with three tracks about the improvement cycle, evidence, and the limits of evaluation. It treats 2028 scenarios as hypotheses, with no guaranteed timeline. The page also lists references such as Self-Refine, Reflexion, Darwin Gödel Machine, and AlphaEvolve, reviewed on September 25, 2026; their results were not reproduced in the project.
Why learn it
Each resource says what is available now and what is still a plan. Reading those caveats is an exercise in the field itself: separating evidence from claims.
Key concepts
RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English