MODULE 3.1 / 7 OF 9
RSI: AI that helps improve AI
Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back.
Area page: eventos.inema.pro/rsi/en/ · in the menu: “AI improvement cycles”
1What RSI Is
What it is
RSI stands for recursive self-improvement: the idea of AI helping improve AI. The area’s central question is practical: how does a system propose changes, test results, and keep what works? The page treats this as a field of study and an open collection of resources, with references, examples, and criteria for evaluating claims. The starting point is a distinction: improving a response, improving an agent, and improving the ability to create new agents are different things. An agent here is an AI system that receives a goal and carries out steps using instructions, memory, and tools.
Why learn it
Claims about AI improving itself come up often and tend to mix these three levels. If you understand the difference, you can ask what exactly improved and what evidence supports the claim.
Key concepts
recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims
In practice
A vendor says its assistant “learns on its own.” The manager asks: does it redo the response, change its own instructions, or produce future improvements? And how was that measured? The answer shows whether the product is at the first, second, or third level.
✓ Do
Write one sentence describing a claim about automatic improvement you have heard, and mark which of the three levels it actually describes.
✗ Avoid
Judging the area by its name: read the main idea on the page before deciding if it’s for you.
2Three Distinct Levels
What it is
The first level is reviewing the response: the model critiques and revises an output. This can improve a task without changing its parameters—the internal values learned during training—or its development process. The second level is improving the system: the agent’s instructions, memory, tools, or code change, and the versions need to be compared on held-out tasks while keeping a way to roll back. The third level is investigating recursion: the improvement helps produce future improvements. This mechanism needs evidence, because more attempts or a higher score do not prove unlimited growth.
Why learn it
Each level calls for a different kind of control. Reviewing a response is inexpensive and reversible; changing the system calls for version tracking and tests; claiming recursion calls for strong evidence. Mixing up the levels can lead you to accept guarantees no one has measured.
Key concepts
review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back
In practice
A team asks the model to review its own summary before delivering it: that is level 1. Then it rewrites the customer service agent’s instruction and compares the two versions on separate cases: level 2. Neither one alone means AI is creating better AI.
3From Idea to a Verifiable Test
What it is
The page suggests starting with a small task: answer based on a policy, create questions from a text, or extract action items from notes. First define the task, the correct reference, and the limits for data, spending, and actions. Then record the initial version and separate development cases from cases held out for validation. Propose one change at a time and compare using the same criteria, including errors, cost, and review effort. Finally, validate before adopting: save results and versions, decide who has the final say, and keep a rollback path.
Why learn it
A familiar evaluation can be exploited, the page warns. Keeping tests independent, checking for invented data, and looking at results outside the sample helps you avoid mistaking a local gain for general autonomy.
Key concepts
small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback
In practice
A school tests an instruction that creates questions from a text. It uses four texts to make adjustments and sets two aside that no one has seen. The new version gets more right on the four, but misses one of the held-out cases; the program lead decides to keep the old version until they understand the error.
Steps to try
- Open the area page beside this lesson.
- Choose one of your tasks and write four lines: task and reference; initial version and held-out cases; the one change; who decides and how to roll back.
- Write one sentence about what changed in your understanding.
4LOOP-R: The Cycle with Guardrails
What it is
LOOP-R is the project’s framework for organizing improvements: Run, Measure, Critique, Propose, Test, Validate, Promote, and Repeat. The page describes it as ready to use in Claude Code, with nine assistants in separate roles, a record of each version, a spending cap, and a command to roll back. The safety rule is explicit: the system cannot decide to replace the current version with a worse one. In the RSI course, LOOP-R is presented as a conceptual proposal from the project. There is also a LOOP-R course with five tracks and 21 lessons for owners and managers without a technical background.
Why learn it
A cycle without guardrails can replace a good version with a worse one without anyone noticing. A spending cap, version records, and rollback make the cycle auditable and safe to repeat.
Key concepts
Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back
In practice
A manager uses the cycle to improve the instruction that classifies customer requests. Each round produces a recorded version; when the new version does worse in a test, it is not promoted and the current version stays in place.
✓ Do
Open the LOOP-R guide and note which of the eight steps your current prompt-adjustment process usually stops at.
✗ Avoid
Jumping to the tool or course without understanding the problem the area addresses.
5Where to Start: Course and Guide
What it is
The “Start here” section has two resources. The RSI Course v6.2 has 18 lessons in six modules, with exercises, reviews, and materials for applying LOOP-R and recording approved knowledge, in Portuguese, English, and Spanish. The RSI Guide maps the topic, including mechanisms, applications, and limits, also in three languages. The page includes an important caveat: the project brings together research and educational content; it is not a ready-to-run autonomous RSI system. The code and materials are in the rsi repository on GitHub.
Why learn it
Starting with the guide gives you the vocabulary and an understanding of the limits; the course turns that into practice with exercises. Knowing the project is educational helps you avoid expecting a tool that improves itself.
Key concepts
RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system
In practice
A consultant wants to explain RSI to a leadership team. They use the RSI Guide to map the three levels and recommend the RSI Course v6.2 to the team member who will lead the improvement tests.
6The Resource Collection: Copiloto, Dream-RSI, 2028
What it is
Along with the guide and course, the resource collection includes three materials with clear caveats. RSI Copiloto is a local assistant with AI, memory, tasks, and routines that compares instructions, asks for human review, and lets you roll back changes; its public demo uses programmed examples, without AI, and real AI is available only in the local version. Dream-RSI is an independent educational guide, not affiliated with Google, about learning from experiment histories; the lab and the eight sessions are future plans, not an implementation of the paper. Alerta IA 2028 is a course with three tracks about the improvement cycle, evidence, and the limits of evaluation. It treats 2028 scenarios as hypotheses, with no guaranteed timeline. The page also lists references such as Self-Refine, Reflexion, Darwin Gödel Machine, and AlphaEvolve, reviewed on September 25, 2026; their results were not reproduced in the project.
Why learn it
Each resource says what is available now and what is still a plan. Reading those caveats is an exercise in the field itself: separating evidence from claims.
Key concepts
RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English
In practice
An analyst tries the Copiloto demo and realizes the responses are programmed. Instead of concluding that “the AI got everything right,” they record that they need the local version for a real test.
Criteria for reviewing your worksheet
Use this rubric after the lab. Each row asks for evidence; marking a topic as read doesn’t mean the worksheet is complete.
| Criterion | Expected evidence | If it doesn’t meet the criterion |
|---|---|---|
| Main idea | You can describe the area in one sentence that reflects the page. | Reread the top of the page and topic 1. |
| Audience | You can say who the area is for and who it isn’t for. | Return to topic 2 and write an example from your work. |
| Core elements | You can name the area’s core elements. | Use the module diagram as a guide. |
| First step | You chose a small, concrete step. | Copy the first step recommended on the page itself. |
| Starting point | You know which course, kit, or project to open first. | Check topic 6 and the page’s access section. |
| Source | Every statement in the worksheet comes from the page. | Replace your assumptions with what the page says. |
HANDS-ON / ~10 MIN
Your first improvement cycle on paper
Keep the area page open: https://eventos.inema.pro/rsi/en/. Use an example from your work, without personal or client data.
Prompt: my first improvement test (RSI)
Paste it into Claude, ChatGPT, or Codex. Replace the words inside < and > with details about your situation.
Read https://eventos.inema.pro/rsi/ and the guide at https://inematds.github.io/rsi/guia/.
I want to test an improvement to a task at work without promising more than I can measure.
My task: <describe the task, e.g., extract action items from meeting notes>.
My current instruction: <paste the instruction or prompt you use today>.
1. Say which of the three levels my idea fits: review the response, improve the system, or investigate recursion.
2. Set up the test: correct reference, data/spending/action limits, development cases, and held-out cases.
3. Suggest ONE change to the instruction and a table to compare both versions (errors, cost, review effort).
4. Say who should decide whether to adopt it and how to return to the previous version.
Do not make up results: leave the result fields blank for me to fill in.
Completion criterion
Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back. Keep the worksheet with the main idea, audience, first step, and starting point.
Open the area page ↗Review what you’ve learned
After ten attempts, the new instruction scored higher on the same set you used to tune it. Does that prove it’s better?
View suggested answer
No. You can overfit to a known evaluation; the change needs to hold up on reserved cases using the same criteria. More attempts or a higher score don’t demonstrate unlimited growth.
If your answer was different, return to the relevant topic and describe the difference in one sentence. This check won’t block your progress.
Module summary
- recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims
- review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back
- small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback
- Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back
- RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system
- RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English
Check the source
Pages read on 28/09/2026. Area content changes; the official page takes precedence over this summary.