MODULE 3.1 / 7 OF 9

RSI: AI that helps improve AI

Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back.

Area page: eventos.inema.pro/rsi/en/ · in the menu: “AI improvement cycles”

0% 0 of 0
01 / RSIWhat RSI Is02 / RSIThree Distinct Levels03 / RSIFrom Idea to a Verifiable Test04 / RSILOOP-R: The Cycle with Guardrails05 / RSIWhere to Start: Course and Guide06 / RSIThe Resource Collection: Copiloto, Dream-RSI, 2028
The 6 topics in this module. At the end, you’ll fill out the area worksheet.
6 topics
~20 min reading and practice
1 area worksheet
1 ready-to-use prompt

1What RSI Is

What it is

RSI stands for recursive self-improvement: the idea of AI helping improve AI. The area’s central question is practical: how does a system propose changes, test results, and keep what works? The page treats this as a field of study and an open collection of resources, with references, examples, and criteria for evaluating claims. The starting point is a distinction: improving a response, improving an agent, and improving the ability to create new agents are different things. An agent here is an AI system that receives a goal and carries out steps using instructions, memory, and tools.

Why learn it

Claims about AI improving itself come up often and tend to mix these three levels. If you understand the difference, you can ask what exactly improved and what evidence supports the claim.

Key concepts

recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims

In practice

A vendor says its assistant “learns on its own.” The manager asks: does it redo the response, change its own instructions, or produce future improvements? And how was that measured? The answer shows whether the product is at the first, second, or third level.

✓ Do

Write one sentence describing a claim about automatic improvement you have heard, and mark which of the three levels it actually describes.

✗ Avoid

Judging the area by its name: read the main idea on the page before deciding if it’s for you.

2Three Distinct Levels

What it is

The first level is reviewing the response: the model critiques and revises an output. This can improve a task without changing its parameters—the internal values learned during training—or its development process. The second level is improving the system: the agent’s instructions, memory, tools, or code change, and the versions need to be compared on held-out tasks while keeping a way to roll back. The third level is investigating recursion: the improvement helps produce future improvements. This mechanism needs evidence, because more attempts or a higher score do not prove unlimited growth.

Why learn it

Each level calls for a different kind of control. Reviewing a response is inexpensive and reversible; changing the system calls for version tracking and tests; claiming recursion calls for strong evidence. Mixing up the levels can lead you to accept guarantees no one has measured.

Key concepts

review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back

In practice

A team asks the model to review its own summary before delivering it: that is level 1. Then it rewrites the customer service agent’s instruction and compares the two versions on separate cases: level 2. Neither one alone means AI is creating better AI.

Reviewlevel 1Systemlevel 2Recursionlevel 3
The three levels increase in ambition and in the proof they require: reviewing an answer does not change the system, and only level 3 talks about improvements that produce improvements.

3From Idea to a Verifiable Test

What it is

The page suggests starting with a small task: answer based on a policy, create questions from a text, or extract action items from notes. First define the task, the correct reference, and the limits for data, spending, and actions. Then record the initial version and separate development cases from cases held out for validation. Propose one change at a time and compare using the same criteria, including errors, cost, and review effort. Finally, validate before adopting: save results and versions, decide who has the final say, and keep a rollback path.

Why learn it

A familiar evaluation can be exploited, the page warns. Keeping tests independent, checking for invented data, and looking at results outside the sample helps you avoid mistaking a local gain for general autonomy.

Key concepts

small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback

In practice

A school tests an instruction that creates questions from a text. It uses four texts to make adjustments and sets two aside that no one has seen. The new version gets more right on the four, but misses one of the held-out cases; the program lead decides to keep the old version until they understand the error.

Steps to try

  1. Open the area page beside this lesson.
  2. Choose one of your tasks and write four lines: task and reference; initial version and held-out cases; the one change; who decides and how to roll back.
  3. Write one sentence about what changed in your understanding.
Define taskInitial versionOne changeSame criteriaValidate and decide
A verifiable test separates an adjustment from validation: the change is adopted only after it passes reserved cases, with rollback available.

4LOOP-R: The Cycle with Guardrails

What it is

LOOP-R is the project’s framework for organizing improvements: Run, Measure, Critique, Propose, Test, Validate, Promote, and Repeat. The page describes it as ready to use in Claude Code, with nine assistants in separate roles, a record of each version, a spending cap, and a command to roll back. The safety rule is explicit: the system cannot decide to replace the current version with a worse one. In the RSI course, LOOP-R is presented as a conceptual proposal from the project. There is also a LOOP-R course with five tracks and 21 lessons for owners and managers without a technical background.

Why learn it

A cycle without guardrails can replace a good version with a worse one without anyone noticing. A spending cap, version records, and rollback make the cycle auditable and safe to repeat.

Key concepts

Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back

In practice

A manager uses the cycle to improve the instruction that classifies customer requests. Each round produces a recorded version; when the new version does worse in a test, it is not promoted and the current version stays in place.

✓ Do

Open the LOOP-R guide and note which of the eight steps your current prompt-adjustment process usually stops at.

✗ Avoid

Jumping to the tool or course without understanding the problem the area addresses.

ExecuteMeasureCritiqueProposeTestValidatePromoteRepeat
The eight LOOP-R steps repeat the cycle, but Validate and Promote act as gates: a worse version does not replace the current one.

5Where to Start: Course and Guide

What it is

The “Start here” section has two resources. The RSI Course v6.2 has 18 lessons in six modules, with exercises, reviews, and materials for applying LOOP-R and recording approved knowledge, in Portuguese, English, and Spanish. The RSI Guide maps the topic, including mechanisms, applications, and limits, also in three languages. The page includes an important caveat: the project brings together research and educational content; it is not a ready-to-run autonomous RSI system. The code and materials are in the rsi repository on GitHub.

Why learn it

Starting with the guide gives you the vocabulary and an understanding of the limits; the course turns that into practice with exercises. Knowing the project is educational helps you avoid expecting a tool that improves itself.

Key concepts

RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system

In practice

A consultant wants to explain RSI to a leadership team. They use the RSI Guide to map the three levels and recommend the RSI Course v6.2 to the team member who will lead the improvement tests.

6The Resource Collection: Copiloto, Dream-RSI, 2028

What it is

Along with the guide and course, the resource collection includes three materials with clear caveats. RSI Copiloto is a local assistant with AI, memory, tasks, and routines that compares instructions, asks for human review, and lets you roll back changes; its public demo uses programmed examples, without AI, and real AI is available only in the local version. Dream-RSI is an independent educational guide, not affiliated with Google, about learning from experiment histories; the lab and the eight sessions are future plans, not an implementation of the paper. Alerta IA 2028 is a course with three tracks about the improvement cycle, evidence, and the limits of evaluation. It treats 2028 scenarios as hypotheses, with no guaranteed timeline. The page also lists references such as Self-Refine, Reflexion, Darwin Gödel Machine, and AlphaEvolve, reviewed on September 25, 2026; their results were not reproduced in the project.

Why learn it

Each resource says what is available now and what is still a plan. Reading those caveats is an exercise in the field itself: separating evidence from claims.

Key concepts

RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English

In practice

An analyst tries the Copiloto demo and realizes the responses are programmed. Instead of concluding that “the AI got everything right,” they record that they need the local version for a real test.

RSI Copilotdemo without AIDream-RSIindependentAI Alert 2028scenarios = hypothesesLOOP-Ra worse version is not accepted
Each item in the collection includes a caveat: read the subtitle before drawing conclusions about what it delivers.

Criteria for reviewing your worksheet

Use this rubric after the lab. Each row asks for evidence; marking a topic as read doesn’t mean the worksheet is complete.

Criterion Expected evidence If it doesn’t meet the criterion
Main idea You can describe the area in one sentence that reflects the page. Reread the top of the page and topic 1.
Audience You can say who the area is for and who it isn’t for. Return to topic 2 and write an example from your work.
Core elements You can name the area’s core elements. Use the module diagram as a guide.
First step You chose a small, concrete step. Copy the first step recommended on the page itself.
Starting point You know which course, kit, or project to open first. Check topic 6 and the page’s access section.
Source Every statement in the worksheet comes from the page. Replace your assumptions with what the page says.

HANDS-ON / ~10 MIN

Your first improvement cycle on paper

Keep the area page open: https://eventos.inema.pro/rsi/en/. Use an example from your work, without personal or client data.

Prompt: my first improvement test (RSI)

Paste it into Claude, ChatGPT, or Codex. Replace the words inside < and > with details about your situation.

Read https://eventos.inema.pro/rsi/ and the guide at https://inematds.github.io/rsi/guia/.
I want to test an improvement to a task at work without promising more than I can measure.
My task: <describe the task, e.g., extract action items from meeting notes>.
My current instruction: <paste the instruction or prompt you use today>.
1. Say which of the three levels my idea fits: review the response, improve the system, or investigate recursion.
2. Set up the test: correct reference, data/spending/action limits, development cases, and held-out cases.
3. Suggest ONE change to the instruction and a table to compare both versions (errors, cost, review effort).
4. Say who should decide whether to adopt it and how to return to the previous version.
Do not make up results: leave the result fields blank for me to fill in.

Completion criterion

Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back. Keep the worksheet with the main idea, audience, first step, and starting point.

Open the area page ↗

Review what you’ve learned

After ten attempts, the new instruction scored higher on the same set you used to tune it. Does that prove it’s better?

View suggested answer

No. You can overfit to a known evaluation; the change needs to hold up on reserved cases using the same criteria. More attempts or a higher score don’t demonstrate unlimited growth.

If your answer was different, return to the relevant topic and describe the difference in one sentence. This check won’t block your progress.

Module summary

  • recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims
  • review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back
  • small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback
  • Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back
  • RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system
  • RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English

Check the source

Pages read on 28/09/2026. Area content changes; the official page takes precedence over this summary.

Module complete