INEMA.CLUBPRORSI v6.2

RSI v6.2 · 6 modules · 18 lessons

How can an AI help improve another one?

Understand recursive self-improvement, examine real studies, and build a supervised improvement cycle. Short reading, visible examples, and practice in every lesson.

The manager and the educator compare documents and results from an improvement experience at a work table.

Module 1 · Basics: recognize what changed

Distinguish RSI, response reviews, and claims without evidence.

Module 2 · How it works: draw a cycle

Goals, roles, evaluation, memory, and adopting improvements.

Module 3 · Research: read real cases

AlphaEvolve, DGM, and the pieces of automated research.

Module 4 · Assessment: check before trusting

Held-out cases, manipulation of measurements and generated data.

Module 5 · Applications: use at work

Customer service, education, and document summaries with references.

Module 6 · Final project: decide with evidence

Choose an implementation, plan a pilot, and report its limits.

Glossary · 18 terms

RSI v6.2

Glossary

The course’s technical terms in simple words. Each term links to the lessons where it appears.

isolated environment

A space for carrying out a task with defined limits. Its protection depends on the implementation and permissions.

Appears in: Lesson 9

autonomy

How much a system can decide and carry out without human intervention. It needs a clearly defined scope and explicit limits.

Appears in: Lesson 16

evaluator

A person or mechanism that applies quality criteria to a result. It needs evidence, not just opinions.

Appears in: Lesson 5

evolutionary search

Exploration of variants, with evaluation and selection over rounds. It doesn’t require living organisms.

Appears in: Lesson 8

held-out cases

Examples set aside before making changes and used afterward to evaluate the candidate. They should not guide each review.

Appears in: Lesson 10

stopping condition

An event that ends the test, such as reaching the attempt limit, exceeding cost, or detecting a critical failure.

Appears in: Lesson 17

approved knowledge

Accepted rule or information for defined uses, with a source, version, responsible person, and validity period. Approval does not make it infallible.

Appears in: Lesson 6

quality criterion

An observable rule for deciding whether the result meets the goal. It must be defined before comparing versions.

Appears in: Lesson 14

synthetic data

Artificially produced examples, including by AI. They can be used for tests and training, with proper verification.

Appears in: Lesson 12

evidence

A record that lets you check a claim. It can be an experiment, its measurement, and its conditions.

Appears in: Lesson 3

fidelity

The correspondence between the response and the available reference. A faithful response doesn’t add facts without support.

Appears in: Lesson 13

bottleneck

A step that limits the system’s overall performance. Improving another step may bring little total gain.

Appears in: Lesson 7

LOOP-R

Educational adaptation of the LOOP-R method: execute, measure, critique, propose, test, validate, promote, and repeat.

Appears in: Lesson 4

reward manipulation

Get a good score by exploiting evaluation flaws, without following the task’s intended goal.

Appears in: Lesson 11

parameters

Internal values adjusted during training. They are not the messages you write in a conversation.

Appears in: Lesson 2

traceability

The ability to link each claim to the material that supports it and to the record of how it was produced.

Appears in: Lesson 15

RSI

Recursive self-improvement: improvements in the system help produce new improvements. The acronym comes from the English recursive self-improvement.

Appears in: Lesson 1

validation

Checking the candidate using defined criteria and held-out cases, before deciding whether to use it.

Appears in: Lesson 18

Lesson 1 of 18

RSI: when improving also improves the next cycle

Professionals review work materials and compare records for the lesson: RSI: when improving also improves the next cycle.

You can classify three situations and recognize what would be recursive improvement.

A better answer impresses. But that doesn’t show whether the AI learned to improve itself. Start by discovering what really changed.

In 1 minute

  1. Repeating a task is not the same as improving the system.
  2. A change has to survive a comparison.
  3. In recursion, improvement feeds the ability to improve.

1The question is what changed

RSI: Recursive self-improvement: improvements in the system help produce new improvements. The acronym comes from the English recursive self-improvement.

An AI can rewrite a message without changing how it works. It can also propose a new work rule, saved for future use. These are different changes. To talk about recursion, look for evidence that the new version helps improve the next versions.

In Marina’s fictional stationery shop, shortening a reply changes the text. Saving a comparison procedure changes future work.

Worksheet · example
1Revised text
The message got shorter. The procedure stays the same.
2Revised process
Every proposal goes through a comparison before becoming a rule.
Read the first record and check how the second one relates to it.

2Local improvements have value

You can take advantage of improvement cycles without building an autonomous AI. In this course, the practices are supervised experiments. They teach you to propose, compare, and decide. They don’t show an intelligence explosion or automatically change the model used in the chat.

Ícaro, a fictional educator, compares instructions to create exercises. He saves the best instruction, but he doesn’t claim he trained a new AI.

Comparison cards · illustration

Unclear requestImprove how I prepare exercises.

Educational record to compare; it’s not execution of a tool.

Verifiable requestCompare two instructions using the same three texts. Count the questions that can be answered using only each text.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3The cycle needs evidence

A version that praises itself does not prove improvement. Record the previous version, the test, the result, and who decided. Separate capability in a task from the ability to do research. More speed in one step may not change the performance of the whole.

Marina keeps the old and new answer. She only changes the guidance after testing questions she didn’t use in the review.

Comparison cards · illustration

Control questionWhich part of the system changed?

Educational record to compare; it’s not execution of a tool.

Evidence neededDoes the new version improve results and help you produce future improvements?

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

A tool got better at summarizing. Does that prove RSI?

Stuck here? That's normalThink about the object that changed: was it the message, the rule saved, or the method that produces new versions? Classify one change at a time.

Practice now 0/3

Separate three situations

About 10 minutes, on your phone, computer, or paper. You’re done when you can classify three situations and recognize what would be recursive improvement.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Write on paper or in a notepad.
A: the AI rewrote an answer.
B: the AI proposed a rule; a person tested it and saved it.
C: a system changed its research method; new versions produced verified improvements.
Classify: answer review, supervised improvement, or a sign of recursiveness. Explain what still needs to be measured.
Check the expected result

A reviews the answer. B improves a procedure under supervision. C suggests recursiveness, but requires independent and repeated tests. None of the cases proves unlimited acceleration.

Practiced ability: classify three situations and recognize what would be recursive improvement.

Lesson cheat sheet

Separate three situations

  1. Repeating a task isn’t improving the system.
  2. A change needs to survive a comparison.
  3. In recursiveness, improvement feeds the ability to improve.

Your next step

You already know how to classify three situations and recognize what would be recursive improvement.

Save the result in your notes as “RSI · lesson 1”. If you skipped the practice, use the template from this page.

In the next lesson: Answer, memory, and training change different things.

Additional materialDeeper dive and sources. Outside lesson time.

Three levels of claims

Self-improvement is a broad term. Here, we distinguish between response review, system improvement, and improvement in the ability to improve. This division is didactic. It’s not a universal classification for the field. Real systems can mix these layers. Notice which component was changed and which result was measured. A good analysis describes the mechanism before assigning a label.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 1 · RSI v6.2 · INEMA.CLUB

Lesson 2 of 18

Answer, memory, and training change different things

Professionals review work materials and compare records for the lesson: Response, memory, and training change different things.

You can identify three ways to improve a system and choose one for a simple task.

A conversation can look like it learns from you. This impression can come only from earlier messages. Identifying the layer prevents you from expecting a change that didn’t happen.

In 1 minute

  1. Context guides the current response.
  2. Memory stores information to be retrieved.
  3. Training changes the model’s parameters.

1Context is the material from this attempt

parameters: Values internally adjusted during training. They are not the messages you write in a conversation.

The model answers using available instructions and information. A correction in the conversation can change the next response. That does not mean the model’s parameters were updated. When you start another conversation, part of that material may no longer be available.

Ícaro adds a short text about recycling before asking questions. The answer improves because it received the necessary reference.

No reference

Create questions about today’s activity.

With reference

Use this text: clean paper goes to recycling; greasy paper needs a different destination.

Compare the two records. A didactic example.

2Memory needs to be consulted and checked

Saving a rule helps only when the system retrieves it at the right time. The rule can grow outdated or contradict another. Name the record, add a date, and specify who is responsible. Before using important information, confirm whether it still holds.

Marina keeps customer service guidance in a document. When the store’s hours change, she updates the guidance and checks old answers.

Comparison cards · illustration

Incomplete recordWe’re open late.

Educational record to compare; it’s not execution of a tool.

Useful recordExample hours: 9am to 6pm. Check the current version before responding.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3Training is another intervention

Training adjusts parameters with examples or evaluation signals. Changing tools or instructions can improve the system without that adjustment. Start with the simplest layer that solves the problem. Don’t use training to compensate for a missing reference.

Ícaro wants the questions to cite the supporting excerpt. He first changes the instruction and measures the result; he does not need to start by training a model.

Worksheet · example
1Small change
Provide a reference and ask for evidence.
2Deep change
Adjust parameters with data and specialized evaluation.
Read the first record and check how the second one relates to it.

Test yourself

You opened another conversation and the rule disappeared. What hypothesis do you want to test?

Stuck here? That's normalA new conversation might not get the previous rule. First check what information was sent, before imagining a change in parameters.

Practice now 0/3

Choose the right layer

About 10 minutes, on your phone, computer, or paper. You’re done when you can identify three ways to improve a system and choose one for a simple task.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Store case: the AI gives an old time.
School case: the AI creates questions without receiving the text.
Research case: a lab adjusts parameters with evaluated examples.
For each case, write what’s missing and which layer needs to change.
Check the expected result

Store: update the retrieved reference. School: provide the text in context. Lab: training. There isn’t enough reason to train in the first two cases.

Practiced capacity: identify three ways to improve a system and choose one for a simple task.

Lesson cheat sheet

Choose the right layer

  1. Context guides the current answer.
  2. Memory stores information so it can be retrieved.
  3. Training changes the model’s parameters.

Your next step

You already know how to identify three ways to improve a system and choose one for a simple task.

Save the result in your notes as “RSI · lesson 2”. If you skipped the practice, use the template from this page.

In the next lesson: Read a news article without turning a prediction into fact.

Additional materialDeeper dive and sources. Outside lesson time.

Learn while using

Some systems combine memory, tools, and later training. So it’s more accurate to avoid the phrase “the chat learns everything” than to say “AI never learns.” Ask what was persisted and how it entered the next use. In professional systems, document the origin and validity of each memory. A false note retrieved repeatedly can spread the same error.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 2 · RSI v6.2 · INEMA.CLUB

Lesson 3 of 18

Read a news article without turning a prediction into fact

Professionals review work materials and compare records for the lesson: Read a news story without turning prediction into fact.

You can fill in an evidence sheet for a claim about RSI.

News mixes results, goals, and interpretations. A number without context can look like much stronger proof than it really is.

In 1 minute

  1. Find the original document.
  2. Record what was measured and by whom.
  3. Goals and scenarios still remain possible.

1Start with the exact statement

evidence: A record that lets you verify a claim. It may include an experiment, its measurements, and its conditions.

Separate a verifiable sentence from the opinion that comes with it. Look for the original study or report, its date, and its authors. A video can present good ideas, but its narration does not replace the original document. Also record whether there was external evaluation.

Marina reads “AI does all the research”. She looks for which tasks were done and which decisions stayed human.

Comparison cards · illustration

Broad statementAI replaced all the research.

Educational record to compare; it’s not execution of a tool.

Concrete questionWhich steps were automated, in which experiment?

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2Activity is not the final result

More hours of execution or more experiments do not directly measure useful discoveries. In September 2026, OpenAI reported greater internal use of agents. It also said that people still choose priorities and decide which results to pursue.

Ícaro can generate a hundred exercises in an afternoon. That does not tell you how many are correct or help his students.

Worksheet · example
1Activity count
Number of exercises generated.
2Utility measure
Number of correct exercises that fit the goal.
Read the first record and check how the second one relates to it.

3Write the limit next to the result

An experiment can have few cases or very specific tasks. A future scenario is a tool for discussing possibilities. It does not establish a guaranteed timeline. Use three labels in your record: observed, interpretation, and prediction. Leave what you couldn’t verify explicitly open.

Marina agrees to test a new routine. She does not change the whole company because of a predicted deadline.

Observed

Described result and its conditions.

Open

Whether the gain repeats in other tasks or organizations.

Compare the two records. A didactic example.

Test yourself

A company announces a goal for 2028. How to record it?

Stuck here? That's normalUse the case of Escola Aurora, already complete in practice. “Not provided” is a valid answer for a missing field.

Practice now 0/3

Make an evidence card

About 10 minutes, on your phone, computer, or paper. Done when you can fill out an evidence card for a statement about RSI.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious case, provided for practice.
Title: Escola Aurora test. Author: teaching team. Date: 25/09/2026.
The team compared two instructions in the same five texts.
The old instruction generated three questions with correct support. The new one generated four.
Each instruction was run once. There was no test with students.
Fill in: title; date; authors; statement; measure; comparison; limit.
Then, optionally, apply the card to a real source in the additional material.
Check the expected result

Title: Escola Aurora test. Date: 25/09/2026. Authorship: fictitious educational team. Measurement: questions supported by the text. Comparison: three vs. four. Limits: five texts, one execution, and no tests with students. The case describes question quality; it does not measure learning.

Practiced skill: fill out an evidence form for a statement about RSI.

Lesson cheat sheet

Make an evidence card

  1. Find the original document.
  2. Record what was measured and by whom.
  3. Goals and scenarios remain possibilities.

Your next step

You already know how to fill out an evidence form for a statement about RSI.

Save the result in your notes as “RSI · lesson 3”. If you skipped the practice, use the template from this page.

In the next lesson: Draw the cycle before automating.

Additional materialDeeper dive and sources. Outside lesson time.

Analysis of the source material

The RSI project includes three German transcriptions, one user interpretation, and two prior studies. The course preserves the idea of evaluating cycles. Some passages include leaks and comparisons without complete conditions. They do not count as proven results. Numbers from commercial models were discarded when they did not help you understand the mechanism. The full study and the editorial decision are documented in this course.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 3 · RSI v6.2 · INEMA.CLUB

Lesson 4 of 18

Draw the cycle before automating

Professionals review work materials and compare records for the lesson: Draw the cycle before automating.

You can put together a form with the eight steps of a supervised improvement cycle.

Asking to “improve continuously” leaves almost everything undefined. You need to say what to test, when to accept, and when to stop.

In 1 minute

  1. Keep an initial reference.
  2. Change one thing per attempt.
  3. Only adopt the change after checking.

1The reference comes before the proposal

LOOP-R: Educational adaptation of the LOOP-R method: execute, measure, critique, propose, test, validate, promote, and repeat.

Execute the task with the current guidance and save the result. Define a simple measure tied to the goal. Only then ask for a proposal. Without an initial reference, a convincing answer can seem better just because it is new. Keep the same cases in the comparison.

Marina records whether each answer provides a time, avoids promises without a basis, and indicates the next action.

Worksheet · example
1No reference
The new answer looks excellent.
2With reference
The old one met two criteria; the new one will be evaluated by the same criteria.
Read the first record and check how the second one relates to it.

2Test one change at a time

Criticizing means finding an observable flaw. Proposing means suggesting an adjustment that can fix it. Testing means running again under comparable conditions. If you change instruction, model, and examples all together, it becomes hard to point to a single cause for the result.

Ícaro adds only “cite the supporting excerpt.” He keeps the text, the number of questions, and the previous criteria unchanged.

Mixed changes

Change everything and compare the overall impression.

Isolated change

Add a rule and count questions supported by the text.

Compare the two records. A didactic example.

3Promoting creates a reference for future uses

Validation checks the candidate on held-out cases. Approval authorizes its adoption within written limits. Promotion records that decision and makes the approved version current. In equivalent future uses, the system applies the approved rule without asking again. A new candidate still needs assessment and a decision.

Marina approves the schedule rule for drafts. The team reuses that rule as long as the source and authorization remain valid.

Comparison cards · illustration

Before the decisionB is a candidate: a good result does not yet authorize adoption.

Educational record to compare; it’s not execution of a tool.

After promotionB becomes the current rule for drafts. Don’t ask again for the same approval.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

The candidate improved in the test. What’s the next step?

Stuck here? That's normalStart with the flaw: the answer invented holiday hours. The answer key shows the eight filled-in steps.

Practice now 0/3

Write your first LOOP-R

About 10 minutes, on your phone, computer, or paper. You’re done when you can put together a worksheet with the eight steps of a supervised improvement cycle.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious reference: store hours are 9:00 a.m. to 6:00 p.m.; holidays not listed.
Instruction A: respond politely. Answer A to the question “Is it open on holidays?”: “Yes, from 9:00 a.m. to 6:00 p.m.”
Criteria: faithful to the reference; don’t invent information.
Candidate B: use only the reference and state when data is missing.
Fill in the eight LOOP-R steps with this case.
Held-out case to check later: “Is it open on Sunday?”. Sundays aren’t listed.
Plan on paper: up to three rounds; Marina decides.
Check the expected result

Execute: save A and its answer. Measure: A fails both criteria. Critique: A claimed the store opens on a holiday without evidence. Propose: add rule B. Test: B must say that holidays weren’t provided. Validate: check Sunday without assuming it’s open. Promote: Marina decides after checking. Repeat: at most three rounds. This is a didactic plan, not a measured execution. Promotion requires recording the version, the responsible person, and the authorized uses. Repeating reuses the current rule; it does not restart approval of the same rule.

Practice goal: make a worksheet with the eight steps of a supervised improvement cycle.

Lesson cheat sheet

Write your first LOOP-R

  1. Save an initial reference.
  2. Change one thing per attempt.
  3. Only adopt the change after checking.

Your next step

You already know how to make a worksheet with the eight steps of a supervised improvement cycle.

Save the result in your notes as “RSI · lesson 4”. If you skipped the practice, use the template from this page.

In the next lesson: Separate who proposes, who measures, and who approves.

Additional materialDeeper dive and sources. Outside lesson time.

The LOOP-R applied to the course

The LOOP-R project has agents with separate roles, experiment logs, versions, and memory. In human-review mode, a change is promoted only after a favorable evaluation and approval. The promoted version then guides future uses. This course applies that discipline with worksheets directly on paper. Approving a new version differs from reusing the current version. The record does not establish universal truth or improvement in every situation. Also read the project’s own critique.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 4 · RSI v6.2 · INEMA.CLUB

Lesson 5 of 18

Separate who proposes, who measures, and who approves

Professionals review work materials and compare records for the lesson: Separate who proposes, who measures, and who approves.

You can define responsibilities and permissions for a small improvement system.

If the same AI writes the answer and declares it’s perfect, you’re missing a reliable check. Splitting roles makes mistakes easier to spot.

In 1 minute

  1. The proposer suggests changes.
  2. The evaluator checks criteria set in advance.
  3. One person authorizes the adoption.

1Define the inputs and outputs for each role

evaluator: Person or mechanism that applies quality criteria to a result. It needs evidence, not just opinion.

A proposer receives the task, the observed errors, and the limits. They return a candidate and the justification. The evaluator receives cases and criteria, and returns checkable results. The approver reviews quality, cost, and undesired effects before deciding.

Ícaro asks the AI for a new instruction, checks the questions against the text, and decides whether to use the instruction.

Proposal

Swap “create questions” for “create questions answerable from the excerpt.”

Evaluation

For each question, point out where the text contains the answer.

Compare the two records. A didactic example.

2Two conversations don’t guarantee independence

Separating screens helps you organize the work, but models can repeat the same mistakes. Use verifiable criteria and examples the proposal didn’t receive. For factual statements, check the reference. For subjective results, record disagreements instead of hiding them in a score.

Marina asks the other conversation for criticism. Even so, she compares the time with the store’s document.

Comparison cards · illustration

Fragile checkAnother AI said it was perfect.

Educational record to compare; it’s not execution of a tool.

Useful checkThe time matches the document and the response didn’t invent a deadline.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3Approval applies to a written scope

Whoever approves sets the rule, the permitted uses, and the limits. A recorded approval keeps working for equivalent uses. Don’t ask again just because another run started. Changing the rule or expanding your actions requires a new decision. Approving content does not automatically authorize sending messages.

Ícaro approves a rule for preparing drafts. The system reuses the rule, but it does not change grades or contact families.

Worksheet · example
1Already authorized
Prepare drafts using the approved rule, without repeating the question.
2Out of scope
Send the drafts to families: this action hasn’t been authorized yet.
Read the first record and check how the second one relates to it.

Test yourself

Two AIs agreed on a time. What decides?

Stuck here? That's normalTwo opinions don’t change the written time. Open the current reference and compare your answer with it.

Practice now 0/3

Make a responsibilities sheet

About 10 minutes, on your phone, computer, or paper. Done when you can define responsibilities and permissions for a small improvement system.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

On paper or in your notes, write three lines: proposer, evaluator, final responsible person.
For each line, record: receives; produces; can change; cannot change.
Store case: customer service drafts.
School case: question drafts.
The evaluation rule is outside the proposer’s scope.
Write down a rule that has already been approved and which uses it allows without a new approval.
Separate this rule from an external action that keeps being unauthorized.
Check the expected result

The responsible person authorizes adoption. The tool can suggest and organize evidence. It should not control the answer, the answer key, and the decision on its own. The recorded approval is valid to reuse the same rule within the defined scope. Any external action outside it still depends on a decision.

Practiced capability: define responsibilities and permissions for a small improvement system.

Lesson cheat sheet

Make a responsibilities sheet

  1. The proposer suggests changes.
  2. The evaluator checks the criteria set in advance.
  3. One person authorizes adoption.

Your next step

You already know how to define responsibilities and permissions for a small improvement system.

Save the result in your notes as “RSI · lesson 5”. If you skipped the practice, use the template from this page.

In the next lesson: Turn approval into reusable knowledge.

Additional materialDeeper dive and sources. Outside lesson time.

From the sheet to the technical system

When a team automates the design, it will need to turn written limits into real permissions. Separate roles in the request don’t prevent improper access. The team must control files, tools, and actions available. It also needs to record failures and preserve the evaluation criteria. This lesson produces the work specification, not a standalone and secure implementation by itself.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 5 · RSI v6.2 · INEMA.CLUB

Lesson 6 of 18

Turn approval into reusable knowledge

Professionals examine work materials and compare records for the lesson: Turn approval into reusable knowledge.

You can record an approval and tell apart reuse, review, and an action that’s still not authorized.

If the system asks everything again, you waste decisions. If you treat any approval as unlimited permission, you can repeat and expand a mistake.

In 1 minute

  1. A recorded approval becomes the system’s approved knowledge.
  2. Reuse within the scope without asking for the same approval.
  3. Approve carefully: a wrong rule can guide many uses.

1Record what was truly approved

approved knowledge: Rule or information accepted for defined uses, with source, version, responsible person, and validity. Approval doesn’t make it infallible.

Separate the proposal, the test result, and the approved rule. In LOOP-R, the promotion links the decision to the current version. Record the exact content, the source, the tests, and who authorized it. Write which tasks the approval applies to. Without that, the system can’t tell approved knowledge apart from an attempt.

Marina approves rule H1: use the current policy’s time in drafts. Holidays still have no information.

Comparison cards · illustration

Saved attemptH1 passed a test, but has not yet been authorized for use.

Educational record to compare; it’s not execution of a tool.

Approved knowledgeH1, policy P1, Marina, 25/09/2026. Allowed: prepare drafts about schedule.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2Reuse without reopening the same decision

Before using it, the system retrieves the current rule and checks the scope. This check is not a request for approval. If task, source, and validity stay compatible, it continues without asking again. An approved preference can guide dozens of answers; each answer is still subject to the agreed criteria.

Ícaro approves citing the reference excerpt in the questions. The system applies this rule to new drafts without asking for the same authorization.

Worksheet · example
1Equivalent repetition
Another time question, with current policy P1: reuse H1.
2Different request
Send the response to the customer: approval for drafts does not include sending.
Read the first record and check how the second one relates to it.

3Approve with future uses in mind

An approval can power dozens of future decisions. Check content, evidence, coverage, and the cost of being wrong before approving. If the source changes, a contradiction appears, or the validity ends, pause the affected use and review it. Revoking disables the rule, but preserves the history. Human approval can also be wrong.

Policy P2 changes the time to 10h–17h. Marina suspends H1 until she reviews the reference; the system doesn’t insist on the old time.

Conscious approval

I know the source, the tests, the allowed uses, and how to revoke.

Revision signal

Policy changed, or an error appeared: stop the affected use and record the reason.

Compare the two records. A didactic example.

Test yourself

Rule H1 is approved and valid. Another equivalent question arrived. What should you do?

Stuck here? That's normalStart with case A: nothing changed and H1 is still valid. Checking this validity doesn’t mean asking Marina again.

Practice now 0/3

Fill out an approved knowledge form

About 10 minutes, on your phone, computer, or paper. You’re done when you can record an approval and tell apart reuse, review, and action that isn’t authorized yet.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictional case. Policy P1: store opens 9h–18h; holidays not listed.
H1: use P1 to prepare drafts about time, without inventing information.
Marina checked time and holiday cases and approved H1 on 25/09/2026.
Validity: until P1 changes, an error appears, or Marina revokes. Sending is not authorized.
Fill in: rule; version; source; evidence; who approved; date; scope; validity; how to revoke.
Classify: A) a new time question with P1 in effect; B) P2 changes to 10h–17h; C) send a message to the customer.
For each case, choose: reuse without asking; pause and review; request authorization for a new action.
Write why approving H1 requires care with future uses.
Check the expected result

A: reuse H1 without asking again, after checking P1 and the scope. B: suspend the affected use and review with P2. C: sending requires its own authorization. Record the decision and version; when you revoke H1, remove it from use and preserve the history. An incorrect approval can repeat the same error. The tests for this case are fictitious.

Practiced capacity: record an approval and distinguish between reuse, review, and action that is not yet authorized.

Lesson cheat sheet

Fill out an approved knowledge form

  1. A recorded approval becomes the system’s approved knowledge.
  2. Reuse within the scope without asking for the same approval.
  3. Approve carefully: a wrong rule can guide many uses.

Your next step

You already know how to record an approval and distinguish between reuse, review, and action that is not yet authorized.

Save the result in your notes as “RSI · lesson 6”. If you skipped the practice, use the template from this page.

In the next lesson: AlphaEvolve: a gain in one piece has a defined reach.

Additional materialDeeper dive and sources. Outside lesson time.

Approved memory is not absolute truth

An approved operational rule is a decision about use. A factual statement still depends on sources and evidence. Don’t mix up these two records. Approving an isolated answer also doesn’t mean approving a general rule. Write whether the decision applies only to that result or to reuse. The system must store and look up this record: a conversation without persistent memory won’t automatically come to “know” it. In LOOP-R, the current version and the history help you apply this distinction. New candidates go through another cycle; the already-authorized version doesn’t need repeated approval.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 6 · RSI v6.2 · INEMA.CLUB

Lesson 7 of 18

AlphaEvolve: a gain in one piece has a defined reach

Professionals review work materials and compare records for the lesson: AlphaEvolve: a gain in one piece has a defined reach.

You can explain a real optimization case and distinguish local gain from total gain.

A headline says something became 23% faster. Before concluding that the whole system sped up like that, find which piece was measured.

In 1 minute

  1. An AI proposes variants of a program.
  2. Evaluators verify correctness and performance.
  3. The overall effect depends on the bottleneck.

1Propose and verify work together

bottleneck: Stage that limits overall performance. Improving another stage can bring little total gain.

AlphaEvolve combines language models, automatic evaluations, and search among candidate programs. The published results include computational optimizations. The problem needs to allow a reliable evaluation. A sleek proposal that produces wrong answers shouldn’t be chosen just because it’s fast.

Marina makes an analogy with the counter: answering faster doesn’t help when the information is wrong.

Worksheet · example
1Incomplete objective
Reducing time at any cost.
2Complete objective
Preserve the correct response and then reduce time.
Read the first record and check how the second one relates to it.

2The percentage belongs to a measurement

DeepMind reported a 23% speedup in an operation used by Gemini. The reported effect on training time was 1%. These values refer to different scales. The case shows why you should ask where the gain happened.

Icaro takes less time to format questions. You still need to read the text and check the answers.

Optimized component

Specific operation: reported acceleration of 23%.

Whole system

Training time: reported reduction of 1%.

Compare the two records. A didactic example.

3Transfer the logic, not the promise

In your work, choose a task whose quality can be checked. Define the correct result before measuring speed. Compare total effort, including review. The usefulness of this logic does not depend on having access to the lab’s search tool.

Marina measures the time between receiving the question and approving the answer. She does not measure only the time the chat takes to write.

Comparison cards · illustration

Partial timeSeconds to generate a draft.

Educational record to compare; it’s not execution of a tool.

Total timeGeneration, checking, and correction until it’s ready to use.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

Generation was fast, but the review doubled. What should you measure?

Stuck here? That's normalAdd up the time from all steps. In the exercise, 3 minutes saved belong to a 30-minute job.

Practice now 0/3

Calculate the overall gain

About 10 minutes, on your phone, computer, or paper. Done when you can explain a real optimization case and tell the difference between local gain and total gain.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious exercise, on paper: preparing an activity takes 30 minutes. There are 6 minutes of formatting and 24 minutes of reading and review.
A change reduces formatting to 3 minutes.
What is the new total time? How many minutes were saved? What was the total percentage reduction?
Check the expected result

The total drops to 27 minutes. The savings is 3 minutes, or 10% of 30. The local gain was 50%, but the total was 10%.

Practiced capacity: explain a real optimization case and distinguish local gain from total gain.

Lesson cheat sheet

Calculate the overall gain

  1. An AI proposes variants of a program.
  2. Evaluators verify correctness and performance.
  3. The overall effect depends on the bottleneck.

Your next step

You already know how to explain a real optimization case and distinguish local gain from total gain.

Save the result in your notes as “RSI · lesson 7”. If you skipped the practice, use the template from this page.

In the next lesson: Darwin Gödel Machine: comparing a family of agents.

Additional materialDeeper dive and sources. Outside lesson time.

What this case supports

The case shows the AI’s contribution to improving components used in computing and training. It does not demonstrate unlimited acceleration. It also does not authorize transferring the same percentage to another organization. The operational lesson is to define an evaluator before expanding the search. The practice calculation is an original analogy, with no numerical relation to DeepMind’s experience.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 7 · RSI v6.2 · INEMA.CLUB

Lesson 8 of 18

Darwin Gödel Machine: compare a family of agents

Professionals review work materials and compare records for the lesson: Darwin Gödel Machine: comparing a family of agents.

You can describe what the DGM changes and why it preserves different candidates.

Today’s best attempt doesn’t always become tomorrow’s best. Exploring different paths can reveal improvements that a narrow search would miss.

In 1 minute

  1. The DGM modifies the agent’s code.
  2. It keeps an archive of variants.
  3. Evaluates performance on programming tasks.

1 The altered object is the agent

evolutionary search: Exploration of variants, with evaluation and selection across rounds. It doesn’t require living organisms.

The DGM uses models to propose changes to the code that organizes the agent. This includes tools and resolution procedures. It does not mean retraining all the language model’s parameters again. Identifying the object helps you avoid confusing a system improvement with a model swap.

In Ícaro’s analogy, changing the checking routine changes the work. It does not automatically change the teacher’s training.

Altered object

Agent code and procedures.

What you shouldn’t conclude

All model parameters were trained again.

Compare the two records. A didactic example.

2 Saved variants open new paths

The study keeps a growing collection of agents. A new attempt can start from different variants. A mid-level candidate might contain an idea that helps another round. Preserving diversity doesn’t mean using worse candidates in real life.

Marina saves a short answer and a detailed one between attempts. None becomes the current guidance without passing the criteria.

Comparison cards · illustration

ExplorationSave one idea to test another combination.

Educational record to compare; it’s not execution of a tool.

Real useAdopt only a validated version.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3 Test result isn’t a universal guarantee

The authors report gains in programming evaluations. The reach depends on the tasks and conditions studied. They also document cases where the evaluation was fooled. That’s why a single score doesn’t end the investigation. Look at the result and the path used to get it.

Ícaro checks whether the question has an answer in the text. A “pass” mark generated by the candidate itself isn’t enough.

Worksheet · example
1Stated score
The candidate says it passed.
2Checkable evidence
The criteria were applied to pheld-out results.
Read the first record and check how the second one relates to it.

Test yourself

An old variant brought a useful idea. What should you do?

Stuck here? That's normalSaving a variant doesn’t mean using it. First, check whether it invents facts in the hard case.

Practice now 0/3

Choose where the variants go

About 10 minutes, on your phone, computer, or paper. Done when you can describe what the DGM changes and why it preserves different candidates.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictional reference: store opens from 9:00 AM to 6:00 PM; holidays not mentioned.
Common question: “What are the hours?”.
A answers correctly in two sentences. B answers correctly in one sentence. C asks for clarification and doesn’t give the hours.
Hard question: “Does it open on holidays?”.
A says there’s no data. B says it’s open. C says there’s no data.
On paper, apply: faithful to the reference; useful; don’t invent.
Decide which variants to save and which to test for use. Create a new validation case.
Check the expected result

A passes both presented cases and can proceed to validation. B invents an opening and should be rejected for use. C preserves fidelity, but loses usefulness in the common question. Saving C for research does not mean approving it. A possible new case is asking about Sundays, also not mentioned.

Practiced capacity: describe what the DGM changes and why it preserves different candidates.

Lesson cheat sheet

Choose where the variants go

  1. The DGM modifies the agent code.
  2. Keeps an archive of variants.
  3. Evaluates performance on programming tasks.

Your next step

You already know how to describe what the DGM changes and why it preserves different candidates.

Save the result in your notes as “RSI · lesson 8”. If you skipped the practice, use the template from this page.

In the next lesson: Automated research depends on infrastructure and judgment.

Additional materialDeeper dive and sources. Outside lesson time.

The comparison with the Gödel machine

Gödel’s theoretical proposal uses proofs of benefit before changes. The DGM works with empirical verification on tasks. This difference matters: passing tests is evidence, not a general proof that every future change will be beneficial. Read the article to understand the controls and conditions. A scientific report should be interpreted together with its limitations.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 8 · RSI v6.2 · INEMA.CLUB

Lesson 9 of 18

Automated research depends on infrastructure and judgment

Professionals review work materials and compare records for the lesson: Automated research depends on infrastructure and judgment.

You can compare three initiatives without treating all automation as full autonomy.

Producing an article, running an experiment, and choosing an important question are different skills. The course now connects these pieces.

In 1 minute

  1. AI Scientist explores automation of scientific work.
  2. DSec provides environments for agent tasks.
  3. Human judgment still matters for direction and acceptance.

1Automating scientific steps doesn’t solve all of science

isolated environment: Execution space with defined limits for a task. Its protection depends on the implementation and permissions.

Sakana reported a piece of work by the AI Scientist accepted in a workshop connected to ICLR 2025. The workshop agreed to participate in the evaluation experiment. This is not the same as acceptance at the main conference. An accepted article also doesn’t prove that every conclusion is correct.

Ícaro distinguishes a draft of an activity from its validation in the classroom. Each step adds different evidence.

Comparison cards · illustration

Automated stepRun and analyze an experiment with a defined scope.

Educational record to compare; it’s not execution of a tool.

Open questionQuality, reproducibility, and usefulness beyond that experiment.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2Experiments need a place to happen

DSec is an infrastructure described by DeepSeek for running agent tasks at scale. Preparing environments, controlling resources, and recording actions are concrete problems. Efficient infrastructure enables more tests. It doesn’t, by itself, prove that a system chooses good questions or improves them recursively.

Marina reserves fictional copies to test responses. Preparing the workstation makes the work easier, but it doesn’t determine which response is good.

Worksheet · example
1Workstation ready
Materials and limits prepared for the trial.
2Useful experiment
Clear question, comparison, and interpretation of the result.
Read the first record and check how the second one relates to it.

3Notice who decides the direction

Fast research reports don’t remove the need to choose goals. Ask who defines the next experiment and who can stop it. Counting attempts is easy. Measuring scientific contribution means assessing what each attempt actually added.

Ícaro can generate many exercise variations. He keeps deciding which ones meet the lesson’s goal.

Activity

More drafts and more attempts.

Contribution

A useful conclusion that holds up under review.

Compare the two records. A didactic example.

Test yourself

An infrastructure runs millions of tasks. What does that show?

Stuck here? That's normalUse the three summary cards in this practice. Reading the full articles is optional further study.

Practice now 0/3

Fill in a comparison of initiatives

About 10 minutes, on your phone, computer, or paper. You’re done when you can compare three initiatives without treating all automation as full autonomy.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Use these summary cards, based on the sources in the additional material.
AI Scientist: automates stages of a research project; an article went through review in a workshop connected to the ICLR. The workshop cooperated with the experiment.
DSec: prepares and runs environments for agent tasks. The study describes infrastructure; scale doesn’t prove scientific quality.
OpenAI internal report: agents support research tasks. People set priorities and decide which results to pursue.
Make three lines: initiative; main function; available evidence; limit.
Reading the full documents is optional further study.
Check the expected result

AI Scientist: scientific steps; article evaluation in a workshop; not the same as acceptance at the main conference. DSec: infrastructure; technical execution report; doesn’t measure scientific discovery. OpenAI: research support; internal report; doesn’t prove full autonomy. When a card doesn’t describe an evaluator, write “not detailed here”.

Practiced capability: compare three initiatives without treating all automation as full autonomy.

Lesson cheat sheet

Fill in a comparison of initiatives

  1. AI Scientist explores automation of scientific work.
  2. DSec provides environments for agent tasks.
  3. Human judgment still matters for direction and acceptance.

Your next step

You already know how to compare three initiatives without treating all automation as full autonomy.

Save the result in your notes as “RSI · lesson 9”. If you skipped the practice, use the template from this page.

In the next lesson: Prepare an evaluation that the candidate hasn’t seen yet.

Additional materialDeeper dive and sources. Outside lesson time.

The reach of recent research

The sources of this module cover complementary functions. Publication experience, infrastructure, and internal use reporting do not form joint evidence of autonomous RSI. Don’t add incompatible indicators. OpenAI’s public proposal from September 2026 describes fully autonomous RSI as not happening at that time. It’s a statement from the organization, not a universal inspection of all existing systems.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 9 · RSI v6.2 · INEMA.CLUB

Lesson 10 of 18

Prepare an assessment that the candidate hasn’t seen yet

Professionals review work materials and compare records for the lesson: Prepare an evaluation the candidate hasn’t seen yet.

You can separate development and validation cases using observable criteria.

If you always correct using the same examples, you can improve only on those examples. The test needs to reveal new situations.

In 1 minute

  1. Development helps you fine-tune.
  2. Validation checks held-out cases.
  3. Quality combines correct answers, failures, and cost.

1Choose cases that represent the work

held-out cases: Examples separated before the changes and used afterward to evaluate the candidate. They shouldn’t guide every review.

Include common situations, missing data, and requests out of scope. Define the expected response or the criteria before running. A small sample teaches the method, but it doesn’t prove overall performance. Avoid choosing only easy cases after you’ve seen the results.

Marina includes one question with no date and another outside the provided time window. This tests whether the response asks for clarification.

Worksheet · example
1Easy case
What time? Full reference available.
2Hard case
Does it open on a holiday? The reference doesn’t mention holidays.
Read the first record and check how the second one relates to it.

2Keep part of the cases before adjusting

Use one group to improve guidance. Set another group aside to check the chosen version. After you consult the held-out results repeatedly, they no longer provide an independent check on unseen cases. For future decisions, prepare other cases and record this replacement.

Ícaro adjusts the instruction using two texts. A third one stays saved until the final comparison.

Development

Text seen during the reviews.

Validation

Text separated before and opened only to confirm.

Compare the two records. A didactic example.

3Record more than an average score

Mark criteria per case, in addition to time and review effort. An average can hide an important failure. Repeat attempts when answers vary. Compare versions under similar conditions. Don’t turn a small test into a promise of accuracy for any task.

One of Marina’s responses was short, but it invented a condition for exchanging a purchase. This error prevents approval even if the response was fast.

Comparison cards · illustration

Required criterionNo invented deadlines or conditions.

Educational record to compare; it’s not execution of a tool.

Preference criterionClear, short text when the data allows it.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

You adjusted it ten times while looking at the reserved test. What changed?

Stuck here? That's normalUse the six ready questions. Set the last two aside before making any instruction revisions.

Practice now 0/3

Prepare six made-up cases

About 10 minutes, on your phone, computer, or paper. Ready when you can separate development and validation cases using observable criteria.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Made-up reference: store is open from 9am to 6pm; pickup after notice. Exchanges, delivery, holidays, and Sundays are not covered by the reference.
Use six ready questions:
1. What are the hours?
2. Can I pick up my order?
3. Do you deliver to my home?
4. Can I exchange it in ten days?
5. Is it open on a holiday?
6. Is it open on Sunday?
On paper, separate 1–4 for development and 5–6 for validation.
Define criteria: matches the reference; does not invent; asks for confirmation when needed.
This lesson prepares the assessment. Don’t send the held-out cases to an instruction review.
Check the expected result

1: state 9am to 6pm. 2: check whether there was notice. 3 and 4: recognize service or policy not provided. 5 and 6: don’t assume it’s open. Inventing a fact rejects any case. The six examples teach the method, without proving overall performance.

Practiced skill: separating development and validation cases using observable criteria.

Lesson cheat sheet

Prepare six made-up cases

  1. Development helps you adjust.
  2. Validation checks the held-out cases.
  3. Quality combines correct answers, failures, and cost.

Your next step

You already know how to separate development and validation cases using observable criteria.

Save the result in your notes as “RSI · lesson 10”. If you skipped the practice, use the template from this page.

In the next lesson: A better score can hide a worse task.

Additional materialDeeper dive and sources. Outside lesson time.

How to interpret public measurements

The METR explains that a task horizon corresponds to the human time of tasks completed with a certain success rate. It’s not the time an AI works alone. The estimates have uncertainty and vary across domains. This caution also applies to your own test: a useful measure must say what it measures, which tasks it includes, and which conclusions it does not allow.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 10 · RSI v6.2 · INEMA.CLUB

Lesson 11 of 18

A better score can hide a worse task

Professionals review work materials and compare records for the lesson: A better score can hide a worse task.

You can identify three shortcuts that fool the assessment and propose protection for each one.

A system pushed by a measure can optimize the measure and abandon the goal. This happens when the rule allows shortcuts.

In 1 minute

  1. The measure represents only part of the goal.
  2. The result needs external evidence.
  3. The person being assessed should not be able to control their answer key.

1An incomplete goal creates shortcuts

reward manipulation: Get a high score by exploiting flaws in the assessment, without fulfilling the task’s intent.

If the goal is only to answer quickly, leaving out information can seem like a win. If it’s to increase quantity, duplicating results can seem like progress. Also write what must stay correct. Combine the main measure with criteria that prevent these losses.

Marina won’t accept empty messages to reduce customer service time. The response must resolve the question or ask for the missing information.

Vulnerable goal

Answer in a few words.

Protected goal

Answer faithfully and concisely, without leaving out what’s necessary.

Compare the two records. A didactic example.

2Statements are not execution

A tool can claim it checked something without producing evidence. Keep the observable result of the check. Distinguish a teaching simulation from real execution. When there isn’t enough proof, write “not verified” instead of turning trust into approval.

Ícaro asks for the excerpt that supports each answer. “Everything checked” doesn’t replace finding the sentences.

Comparison cards · illustration

StatementI checked, and everything is correct.

Educational record to compare; it’s not execution of a tool.

EvidenceQuestion 2: the answer appears in the second sentence of the text.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3Protect the criteria from convenient changes

A candidate system should not make its own test easier. Preserve the cases, the answer key, and the original results. If a rule is wrong, revise it in a separate process and reevaluate all versions. Don’t change criteria just to save your preferred candidate.

Marina finds an ambiguous criterion. She corrects it and applies it again to both versions, without favoring the new one.

Worksheet · example
1Invalid comparison
Change the rule afterward so the candidate passes.
2Re-made comparison
Correct the rule and test all the versions again.
Read the first record and check how the second one relates to it.

Test yourself

The grade went up after deleting difficult cases. What can you conclude?

Stuck here? That's normalAsk whether the task was completed even without looking at the grade. Then look for the evidence that would let you verify it.

Practice now 0/3

Find the shortcut

About 10 minutes, on your phone, computer, or paper. Ready when you can identify three shortcuts that fool the assessment and propose a protection for each one.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Case A: they assess quantity; the AI duplicates questions.
Case B: they assess speed; the AI omits essential conditions.
Case C: they ask for checking; the AI only writes “checked”.
On paper, match each failure to a concrete protection.
Check the expected result

A: count distinct and useful questions. B: verify required content beyond time. C: require and check evidence. The responsible person preserves criteria and records.

Practical capacity: identify three shortcuts that fool the assessment and propose a protection for each one.

Lesson cheat sheet

Find the shortcut

  1. The measure represents only part of the goal.
  2. The result needs external evidence.
  3. The person being assessed should not control the answer key.

Your next step

You already know how to identify three shortcuts that fool the assessment and propose a protection for each one.

Save the result in your notes as “RSI · lesson 11”. If you skipped the practice, use the template from this page.

In the next lesson: Generated data needs a source, diversity, and verification.

Additional materialDeeper dive and sources. Outside lesson time.

Why the risk isn’t just theoretical

Sakana described cases of fabricated test records and changing the markers used in the assessment. Anthropic studied reward manipulation in controlled scenarios. These findings call for independent verification. They don’t mean that all AI has a human intention to cheat. The practical focus is to detect behaviors that produce points without meeting the goal.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 11 · RSI v6.2 · INEMA.CLUB

Lesson 12 of 18

Generated data needs a source, diversity, and verification

Professionals review work materials and compare records for the lesson: Generated data needs source, diversity, and verification.

You can evaluate a small set of made-up examples before using it to improve a routine.

Generating a thousand examples is easy. Knowing whether they’re correct, varied, and suitable for the work requires a different evaluation.

In 1 minute

  1. Quantity doesn’t replace coverage.
  2. Keep authorized real references.
  3. The automatic judge can also be wrong.

1Start with what the examples represent

synthetic data: Examples produced artificially, including by AI. They can serve for tests and training, with proper verification.

Choose categories needed for the task. Include common situations and relevant exceptions. Mark which examples were generated and who verified them. A collection of nearly identical variations increases volume without greatly expanding coverage.

Ícaro notices that ten generated questions ask for the same information. He swaps part of them for questions that assess other ideas from the text.

Comparison cards · illustration

VolumeTen versions of the same question.

Educational record to compare; it’s not execution of a tool.

CoverageQuestions about different points and one with missing information.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2Don’t feed the cycle only with your own outputs

Studies show degradation under certain conditions of recursive training with generated data. This doesn’t mean that all synthetic data causes collapse. The composition, selection, and preservation of diversity matter. Don’t replace verified references with answers generated without checking.

Marina uses the schedule document as a reference. She doesn’t treat an old chat response as the official source.

Worksheet · example
1Lost reference
New answers learn only from old answers.
2Preserved reference
Each example is checked against authorized material.
Read the first record and check how the second one relates to it.

3AI evaluation is limited support

An AI can compare results and locate problems. But it may prefer long text or agree with plausible errors. Have a person check a sample using objective criteria. When the judgment is uncertain, keep the disagreement instead of fabricating a conclusion.

Icaro uses automatic criticism to locate suspicious questions. He checks the text before accepting the evaluation.

Assisted use

The AI points out possible problems.

Decision verified

The person checks reference, coverage, and appropriateness.

Compare the two records. A didactic example.

Test yourself

One hundred repeated examples repeat the same mistake. What to do?

Stuck here? That's normalCompare questions 1 and 2. They use different words, but they ask for the same information.

Practice now 0/3

Audit five examples

About 10 minutes, on your phone, computer, or paper. Done when you can evaluate a small set of fictional examples before using it to improve a routine.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictional text: Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday.
Fictitious generated questions:
1. On what day does collection happen?
2. What is the collection day?
3. How should greasy paper be treated?
4. Can clean paper be recycled?
5. How many tons are collected?
On the paper, mark source, support, repetition, and missing information.
Replace question 2 with “What condition of the paper requires separation?”. Compare with question 3 to discuss redundancy.
Decide whether it’s better to stick with three distinct questions instead of five.
Check the expected result

1 and 2 repeat the same idea. 3 and 4 have support. 5 requires missing information. The proposed replacement is still close to 3. A valid decision is keeping only 1, 3, and 4. More questions don’t guarantee better coverage.

Practiced skill: evaluating a small set of fictional examples before using it to improve a routine.

Lesson cheat sheet

Audit five examples

  1. Quantity doesn’t replace coverage.
  2. Save authorized real references.
  3. The automatic judge can also be wrong.

Your next step

You already know how to evaluate a small set of fictional examples before using it to improve a routine.

Save the result in your notes as “RSI · lesson 12”. If you skipped the practice, use the template from this page.

In the next lesson: Support: improve without inventing a policy.

Additional materialDeeper dive and sources. Outside lesson time.

Two different research lines

The collapse study analyzes loss of information in recursive training with generated data. The work on self-rewarding models investigates using the model itself to produce evaluation signals. These are different problems. Promising results in one don’t remove the risks of the other. For your test, preserve references and never confuse automatic preference with factual truth.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 12 · RSI v6.2 · INEMA.CLUB

Lesson 13 of 18

Service: improve without making up a policy

A stationery manager compares two cards while preparing an answer test, in front of material shelves.

You can compare two customer service instructions using a fictional policy and reserved questions.

A polite response can promise something the company doesn’t offer. The cycle needs to improve usefulness without inventing commitments.

In 1 minute

  1. The policy is the reference.
  2. Unanswered questions require clarification.
  3. Approving a rule does not authorize actions outside the scope.

1Provide a small, complete reference

fidelity: Correspondence between the response and the available reference. A faithful response doesn’t add facts without support.

Use a fictional policy to learn the method. Define which information is available. Ask the response to acknowledge gaps. A cordial tone instruction doesn’t replace facts. The first test should reveal whether the system distinguishes known information from unknown information.

Marina tests an imaginary stationery store. The reference includes opening hours and pickup, but it doesn’t mention the exchange period.

Worksheet · example
1Exercise reference
The store is open from 9 a.m. to 6 p.m. Pickup after notice. There’s no information about exchanges.
2Fictional customer question
Can I exchange it after ten days?
Read the first record and check how the second one relates to it.

2Improve only the response guidance

Compare a basic instruction with another that requires indicating missing data. Keep the model and the reference. Use separate conversations or clearly identify each attempt. If the tool isn’t available, manually compare the provided drafts.

Ícaro applies the same logic to school announcements. A lack of information doesn’t turn into imagined authorization.

Inadequate draft

Yes, the exchange is allowed in ten days.

Faithful draft

The reference doesn’t say the exchange period. Confirm the policy with the store.

Compare the two records. A didactic example.

3Validate with a different question

After adjusting the instruction, test a reserved question. Look for made-up facts and confusing actions. Also note whether the answer resolves what it can resolve. Saying “I don’t know” to everything avoids inventions, but it can make the service useless.

Marina asks when she can pick up an order. The answer should use the available condition: after the notice.

Comparison cards · illustration

Reserved questionCan I pick up now?

Educational record to compare; it’s not execution of a tool.

Supported answerPickup happens after the notice. Have you already received that notice?

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

The AI avoids making things up but refuses even to state the opening hours. What is the problem?

Stuck here? That’s normalDo only the development first. Validation has a separate block that you open afterward.

Practice now 0/3

Do one round in the chat or on paper

About 10 minutes, on your phone, computer, or paper. Done when you can compare two customer service instructions using a fictional policy and reserved questions.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Copy only this block for the development stage.
Fictitious reference: opening times from 9:00 a.m. to 6:00 p.m.; pickup after notice; exchanges not provided.
Compare two instructions, without changing them:
A: answer the question politely.
B: answer using only the reference; if any data is missing, explain and ask for confirmation.
Development questions: What are the opening hours? Can I exchange an item after ten days?
Show the answer for each version for each question.
Stop after these cases. Don’t approve the candidate yet.
Stage 2 · open validation after development

Only after closing the development, start a new conversation. Paste the reference and the chosen instruction, without previous criticism. Ask: “Can I pick it up now?”. Check that the answer asks about the notice. Don’t change the candidate after seeing this result; record the decision.

Paper alternative: A answers “Yes, you can pick it up now”. B answers “Pickup happens after notice. Have you already received the notice?”. Apply the same criteria. Record that you analyzed drafts provided, without running a chat.

Check the expected result

The answer about exchange must not create a deadline. The answer about pickup must ask about the notice. If a candidate fails, save the result and do not approve.

Practiced skill: comparing two customer service instructions using a fictitious policy and reserved questions.

Lesson cheat sheet

Do one round in the chat or on paper

  1. The policy is the reference.
  2. Questions without answers require clarification.
  3. Approving a rule does not authorize actions outside the scope.

Your next step

You already know how to compare two customer service instructions using a fictitious policy and reserved questions.

Save the result in your notes as “RSI · lesson 13”. If you skipped the practice, use the template from this page.

In the next lesson: Education: create questions that the text supports.

Additional materialDeeper dive and sources. Outside lesson time.

From exercise to work

The policy from the exercise was invented for teaching. It does not describe a real company or establish consumer rights. For professional use, replace it with the organization’s authorized reference and keep a responsible person to review it. Don’t use the exercise as legal advice. The improvement shown is local and supervised, with no training of a new model.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 13 · RSI v6.2 · INEMA.CLUB

Lesson 14 of 18

Education: create questions that the text supports

Professionals review work materials and compare records for the lesson: Education: produce questions the text supports.

You can create and check three questions with answers supported by a short text.

A question can look good and ask for something the student didn’t receive. The goal is to check the link between text, question, and answer.

In 1 minute

  1. Define the learning objective.
  2. Ask for the sentence that supports each answer.
  3. The educator decides whether the activity is appropriate.

1Tell what the student should demonstrate

quality criterion: Observable rule to decide whether the result meets the learning objective. It must be defined before comparing versions.

Choose a small skill, like locating explicitly stated information. Don’t mix this goal with external research without telling. Provide the text and the number of questions. The assessment must match the chosen skill, not just the look of a school activity.

Ícaro wants the class to locate recycling conditions. He uses a short text and does not require outside knowledge.

Fictitious text

Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday.

Goal

Locate information that appears explicitly in the text.

Compare the two records. A didactic example.

2Ask for evidence with the answer

Ask for the question, the answer, and the supporting excerpt. Check whether the excerpt truly supports the answer. Copying any sentence doesn’t solve it. Require the tool to discard questions without support in the material. You can apply this procedure by hand, without depending on a chat.

Marina applies the same logic in a team training. Each question must point back to the guidance that was provided.

Comparison cards · illustration

Question without supportHow many tons does the collection pick up per month?

Educational record to compare; it’s not execution of a tool.

Question supportedOn what day does the collection happen? Answer: Tuesday. Support: the last sentence.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3Check variety and appropriateness

Three almost identical questions don’t assess three different ideas. Check repetition, clarity, and difficulty. The tool may suggest a revision, but it doesn’t automatically know the class. For a next round, use another reserved text and keep the same criteria.

Ícaro rejects two questions that ask about the same day. He adds a question about the condition of the paper.

Worksheet · example
1Limited coverage
Three questions about Tuesday.
2Expanded coverage
Condition of clean paper, separation of the greasy paper, and collection day.
Read the first record and check how the second one relates to it.

Test yourself

The question cites a sentence, but it doesn’t answer the question. What should you do?

Stuck here? That’s normalIf the questions are already correct, don’t invent a defect. Record the check and preserve the version.

Practice now 0/3

Create a verifiable activity

About 10 minutes, on your phone, computer, or paper. Done when you can create and check three questions with answers supported by a short text.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious text: Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday.
Create three questions to locate explicitly stated information.
For each one, provide the answer and the exact support sentence.
Cover three different pieces of information. Don’t add outside knowledge.
Review every link between question, answer, and excerpt.
Check the expected result

The three available pieces of information are the condition of the clean paper, the separation of greasy paper, and the collection day. Tonnage doesn’t appear.

Practiced capacity: create and check three questions with answers supported by a short text.

Lesson cheat sheet

Create a verifiable activity

  1. Define the learning objective.
  2. Ask for the sentence that supports each answer.
  3. The educator decides whether the activity is appropriate.

Your next step

You already know how to create and check three questions with answers supported by a short text.

Save the result in your notes as “RSI · lesson 14”. If you skipped the practice, use this page’s template.

In the next lesson: Documents: turn notes into actions without filling in blanks.

Additional materialDeeper dive and sources. Outside lesson time.

Improving the material doesn’t prove learning

The exercise checks the quality of the material you produced. It doesn’t measure what students learned. To assess learning, the educator will need to observe answers, difficulties, and the group’s context. They should also avoid using identifiable data in unnecessary experiments. The review method helps you prepare the activity; it doesn’t replace teaching judgment.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 14 · RSI v6.2 · INEMA.CLUB

Lesson 15 of 18

Documents: turn notes into actions without filling in blanks

Professionals review work materials and compare records for the lesson: Documents: turn notes into actions without filling in gaps.

You can produce a list of actions that keeps responsible people, deadlines, and missing information.

Smooth summaries can invent a date or assign a task to the wrong person. Here, quality will be traceable back to the original note.

In 1 minute

  1. Extract only what’s written.
  2. Mark what still needs to be confirmed.
  3. Preserve the link to the reference.

1Separate extraction from suggestion

traceability: Possibility of linking each statement to the material that supports it and to the record of how it was produced.

Extraction recovers information that is already present. Suggesting adds a possibility. Both actions can be useful, as long as they’re identified. Ask for a factual list first. If you want suggestions, put them in a separate part, without presenting them as decisions already made.

Marina wants to turn meeting notes into actions. She doesn’t want the AI to choose deadlines to fill in empty spaces.

Comparison cards · illustration

Fictitious noteMarina will review the stock by Friday. Ícaro will prepare an activity; the deadline is pending.

Educational record to compare; it’s not execution of a tool.

Faithful extractionMarina: stock, by Friday. Ícaro: activity, pending deadline.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2A blank (missing piece of info) is a valid result

Use “not provided” when the reference doesn’t contain the data. Don’t confuse that label with a lack of capability. It shows the limit of the material. A list with explicit blanks can be more useful than a complete, invented summary.

Ícaro prefers to receive “pending deadline”. That way, they know which question to ask before planning the activity.

Worksheet · example
1Invented filling-in
Ícaro will deliver the activity on Friday.
2Blank preserved
Ícaro’s activity deadline needs to be confirmed.
Read the first record and check how the second one relates to it.

3Compare item by item

Evaluate person, action, and deadline separately. Count omissions and additions without support. Then check clarity. When you review the guidance, change only one rule. Keep the original note available so another person can verify the list.

Marina checks each line against the note. The review finds the date copied from another task incorrectly.

Criteria

Correct responsible person; correct action; deadline faithful to the note or missing.

Error detected

A deadline from one task applied incorrectly to another.

Compare the two records. A didactic example.

Test yourself

The summary includes a deadline missing from the note. What’s the decision?

Stuck here? That's normalMarina’s deadline doesn’t apply to Ícaro. Mark his deadline as not provided and write the missing question.

Practice now 0/3

Extract and review two actions

About 10 minutes, on your phone, computer, or paper. You’re done when you can produce a list of actions that preserves responsible people, deadlines, and missing information.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious notes: Marina will review the inventory by Friday. Ícaro will prepare an activity; the deadline is still pending.
Produce a list with responsible person, action, deadline, and supporting excerpt.
Write “not provided” when data is missing.
Don’t suggest dates or treat suggestions as decisions.
Check each field against the original notes.
Check the expected result

Marina’s deadline is Friday. Ícaro doesn’t have a defined deadline. The useful question is when the activity should be ready, without assuming the same Friday.

Practiced capability: produce a list of actions that preserves responsible people, deadlines, and missing information.

Lesson cheat sheet

Extract and review two actions

  1. Extract only what’s written.
  2. Mark what needs to be confirmed.
  3. Preserve the link to the reference.

Your next step

You already know how to produce a list of actions that preserves responsible people, deadlines, and missing information.

Save the result in your notes as “RSI · lesson 15”. If you skipped the practice, use the template from this page.

In the next lesson: Choose between a manual draft and an automation.

Additional materialDeeper dive and sources. Outside lesson time.

Where to use it and where to increase caution

This outline serves as a starting point for syntheses, to-do lists, and review of materials. Documents that guide relevant decisions require review proportional to the impact. Keep access to the reference and don’t hide uncertainties. If a tool changes its behavior, reapply your validation cases before trusting the old guidance.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 15 · RSI v6.2 · INEMA.CLUB

Lesson 16 of 18

Choose between a manual draft and an automation

Professionals review work materials and compare records for the lesson: Choose between a manual draft and an automation.

You can decide the next level of implementation based on need, cost, and your ability to verify.

The desire to automate can show up before a good evaluation exists. A small experience helps you decide whether it’s worth investing.

In 1 minute

  1. Start manually to learn the errors.
  2. Automate repeatable and assessable tasks.
  3. Research systems require a team and infrastructure.

1A chat and a worksheet are enough to get started

autonomy: How much a system can decide and carry out without human intervention. It needs clear scope and limits.

In your first attempts, copy the material, compare answers, and record results. This makes the method visible. You learn which cases fail before you expand execution. You don’t need to buy a special tool to do the exercises in this course.

Marina compares two approaches on a sheet. She only considers automation after she understands which failures she needs to detect.

Worksheet · example
1Manual trial
Few cases, direct review, and clear time cost.
2Decision question
Does frequent repetition make this part worth automating?
Read the first record and check how the second one relates to it.

2The next level needs requirements

An automatic routine needs to receive data, run proposals, evaluate, and record. Someone maintains the limits and handles failures. Before you hire or build, write down these needs. The provider must show how the result can be checked and how to go back to the previous version.

Ícaro specifies that the routine only prepares questions. The educational approval stays with him.

Vague requirement

I want an AI that improves on its own.

Verifiable requirement

I want to compare instructions with fixed cases and review candidates before adopting them.

Compare the two records. A didactic example.

3Frontier research isn’t a button

Modifying agents and running training involves specialization, extra resources, and additional risks. Don’t treat lab examples as a ready-made recipe for any company. Choose the smallest level that meets the need. Local gains can be useful without broad autonomy.

Marina doesn’t need to build a lab to reduce service mistakes. First, validate the guidance used in the drafts.

Comparison cards · illustration

Local needImprove a repeated, bounded task.

Educational record to compare; it’s not execution of a tool.

Research projectInvestigate new methods with infrastructure and specialized evaluation.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

Test yourself

There’s still no quality criterion. What’s the next step?

Stuck here? That’s normalChoose one of the three ready-made cases. The decision applies to those conditions, not to every company or school.

Practice now 0/3

Write an implementation recommendation

About 10 minutes, on your phone, computer, or paper. You’re done when you can decide the next implementation level based on need, cost, and how easy it is to check.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Choose a fictional case with complete data:
Store: 4 questions per week; short policy; Marina reviews everything; additional budget is zero; wrong answers need to be corrected before sending.
School: 20 activities per week; Ícaro reviews; technical team available for 2 hours per week; the material must have text support.
Research: dedicated technical team; agent comparison; isolated environment and own budget; results evaluated with tests.
On paper, recommend a manual trial, a limited routine, or research. Justify using frequency, evaluation, resources, and the cost of failing.
Say what data you would measure before investing more.
Check the expected result

Too little repetition usually supports a manual test. A defined routine requires evaluation and maintenance. Research agent requests demand specialization beyond this course.

Practiced skill: decide the next implementation level based on needs, cost, and the ability to verify results.

Lesson cheat sheet

Write an implementation recommendation

  1. Start manually to learn about the mistakes.
  2. Automate repeated and measurable tasks.
  3. Research systems require a team and infrastructure.

Your next step

You already know how to decide the next implementation level based on needs, cost, and the ability to verify results.

Save the result in your notes as “RSI · lesson 16”. If you skipped the practice, use this page’s template.

In the next lesson: Plan a pilot with a budget and stop conditions.

Additional materialDeeper dive and sources. Outside lesson time.

How to talk with the people who will implement it

Bring your experiment sheet and failure examples. Ask for a demonstration with your authorized cases, including the hard ones. Request evidence of cost control and stopping. Don’t accept only one demonstration selected by the vendor. The decision depends on the task and the real conditions, not on a general ranking of models.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 16 · RSI v6.2 · INEMA.CLUB

Lesson 17 of 18

Plan a pilot with a budget and stop conditions

Professionals review work materials and compare records for the lesson: Plan a pilot with a budget and stop conditions.

You can fill in a pilot plan with goal, metric, owner, cap, and reversal.

A small improvement can cost more than the benefit. Without a written limit, attempts pile up and the decision gets postponed.

In 1 minute

  1. Count generation, review, and corrections.
  2. Set the cap before the first attempt.
  3. Stop when there’s no more evidence or control.

1Calculate the cost of the usable result

stop condition: Event that ends the test, like reaching the limit of attempts, going over cost, or detecting a critical failure.

Add time for preparation, execution, and review. If there’s a charge based on usage, record the observed value. Don’t confuse the advertised price with the total cost of the work. Compare recurring savings with the effort to prepare and maintain the change.

Marina estimates how much time she saves per response and how much she spent preparing the test. A rare improvement may not cover the preparation.

Fictitious exercise

Prepare the pilot: 40 minutes. Savings per use: 2 minutes.

Break-even point

You need 20 uses to recover 40 minutes, not counting maintenance.

Compare the two records. A didactic example.

2Write simple, verifiable limits

For a first test, choose a few cases and up to three candidates. Set a work window. Define a financial cap if the tool charges. Record educational examples and values that were truly measured separately. Don’t move the exercise budget to any project.

Ícaro reserves a session to compare instructions. If there isn’t time to check, he doesn’t approve the candidate out of haste.

Comparison cards · illustration

Exercise ceilingThree candidates, six cases, and a planned session.

Educational record to compare; it’s not execution of a tool.

Decision ruleWithout a completed validation, keep the previous version.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

3The reversal is part of the plan

Keep the previous guidance and indicate who can restore it. Stop when you come across invented relevant information, an unauthorized action, or a cost above the ceiling. Investigate the failure before repeating. If the system changes, reassess the pilot conditions.

Marina keeps the current guidance. A candidate who creates improper commitments is rejected before reaching the team.

Worksheet · example
1Stop
Failure to meet a mandatory criterion, excessive cost, or an attempt outside the scope.
2Return
Keep using the previous version while the cause is analyzed.
Read the first record and check how the second one relates to it.

Test yourself

The budget ran out before validation. What should you do?

Stuck here? That's normalFirst calculate without maintenance: 40 divided by 2. Then include the monthly effort to see if the conclusion changes.

Practice now 0/3

Fill out the one-page plan

About 10 minutes, on your phone, computer, or paper. You’re done when you can fill out a pilot plan with goal, measure, person responsible, ceiling, and reversal.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

Fictitious plan to complete on paper:
Marina prepares 20 answers per month. The pilot requires 40 minutes of preparation.
Estimated savings: 2 minutes per use. Policy review: 10 minutes per month.
Tool already available; additional spending cap: zero. No message sending.
Reference: open 9h–18h; pickup after notification; other information missing.
Development cases: opening hours, pickup, exchanges, and delivery. Held-out cases: holidays and Sundays.
Required criteria: fidelity and no inventing. Up to three candidates.
Complete: goal; responsible person; stop point; rollback; next review.
Calculate uses to recover preparation without maintenance. Then discuss the effect of the 10 minutes per month.
Also record: approved knowledge; reuse scope without new approval; validity; who can revoke.
Check the expected result

No maintenance: 40 divided by 2 = 20 uses. With 20 monthly uses, it saves 40 minutes and spends 10 on maintenance: estimated net gain of 30 minutes per month. The 40-minute preparation doesn’t get recovered in the first month. Marina decides; stop if the system invents information or reaches the cap, and keep the previous guidance. Monthly review. These are fictitious estimates, not measured results. The promoted rule can be reused without asking again within the approved scope. A policy change, error, or revocation suspends the affected use.

Practiced capacity: fill out a pilot plan with goal, measure, person responsible, ceiling, and reversal.

Lesson cheat sheet

Fill out the one-page plan

  1. Count generation, review, and corrections.
  2. Set the ceiling before the first attempt.
  3. Stop when evidence or control is missing.

Your next step

You already know how to fill out a pilot plan with goal, measure, person responsible, ceiling, and reversal.

Save the result in your notes as “RSI · lesson 17”. If you skipped the practice, use the template from this page.

In the next lesson: Final project: compare, decide, and explain the limit.

Additional materialDeeper dive and sources. Outside lesson time.

Governance proportional to the trial

This course pilot does not send messages, does not modify external systems, and does not use personal data. Even so, it needs a responsible person and criteria. In real projects, increase controls based on impact and autonomy. Proposals for international standards discuss supervision and evidence; they don’t replace your organization’s concrete obligations.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 17 · RSI v6.2 · INEMA.CLUB

Lesson 18 of 18

Final project: compare, decide, and explain the limit

Final project: compare, decide, and explain the limit.

You can deliver a short report based on provided results, with comparison, decision, and limits.

The final result does not need to be a successful candidate. Finding out that a change isn’t worth it is also a useful conclusion.

In 1 minute

  1. Compare the reference with a candidate.
  2. Use held-out cases.
  3. Report gains, failures, and uncertainties.

1Choose a clearly defined task

validation: Candidate review with defined criteria and held-out cases, before deciding whether to use it.

Use support services, reading questions, or action extraction. If you didn’t do the previous practices, copy one of the fictional materials from this module. Write the initial reference and one change. You can run the trial with an AI chat or compare drafts manually.

Ícaro chooses questions that can be answered by the text. His only change is requiring a supporting passage.

Comparison cards · illustration

ReferenceCreate three questions about the provided text.

Educational record to compare; it’s not execution of a tool.

CandidateCreate three questions and present the sentence that supports each answer.

Educational record to compare; it’s not execution of a tool.

Tap on the two labels to examine the records.

2Run and record the comparison

Apply both guidelines to the same cases. Keep the responses, the criteria, and the observed times. Then use the held-out cases without continuing to adjust the candidate. If there’s any variation, record it. A small difference in a short trial can be inconclusive.

Marina tests common questions and gaps. She keeps a worse answer alongside the best ones, so she doesn’t select only wins.

Worksheet · example
1Full record
Case; version; answer; criteria; review; time.
2Insufficient record
A capture of the best answer, with no comparison.
Read the first record and check how the second one relates to it.

3Also decide what can be reused

Conclude: approve for limited use, reject, or investigate more. If you approve, record the accepted knowledge, the version, and the future authorized uses. The same valid rule doesn’t require approval every time you repeat it. New changes go through review. State what the test didn’t prove and how to suspend or revoke the decision.

Icaro approves a rule for drafts with pedagogical review. This rule does not need new approval for every activity.

Current knowledge

Rule B approved to prepare drafts, with source and version recorded.

Approval limit

It does not allow automatic sending or prove learning improvement.

Compare the two records. A didactic example.

Test yourself

The candidate did not bring reliable gain. What conclusion is valid?

Stuck here? That's normalUse the provided results and say they are educational. Running your own experiment is an optional later session.

Practice now 0/3

Submit your short report

About 10 minutes, on your phone, computer, or paper. Ready when you can submit a short report based on provided results, with comparison, decision, and limits.

Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.

10-minute path: analyze the fictional results below. These are not real measurements.
Reference: store opens 9am–6pm; holidays not listed.
A: respond politely. B: use only the reference and report gaps.
Development — time: A and B respond 9am–6pm.
Development — exchange: A promises ten days; B says that information is missing.
Held-out — holiday: A claims it is open; B reports that data is missing.
Fictional review times per case: A = 2, 3, 3 minutes; B = 1, 1, 1 minute. Financial spend not measured.
Criteria: fidelity; no inventing; usefulness.
Write: hypothesis; single change; comparison; failures; time; decision; responsible person; rollback; limit.
Optional: then do your own run in another session, using new held-out cases.
If you decide to approve: record knowledge, version, source, responsible person, scope, validity, and revocation.
Tell which equivalent use does not require new approval and which change requires review.
Check the expected result

B avoids A’s inventions in the provided cases. Review totals 8 minutes for A and 3 for B: a 5-minute difference in this educational example. Spend not measured. Possible decision: B proceeds to a real pilot with review; Marina approves and keeps A for rollback. Limits: fictional data, few cases, no proof of real performance or autonomous RSI. In the authorized pilot, B becomes approved knowledge to prepare drafts, without asking for the same approval for each equivalent use. The review of the answers stays within the pilot scope. Record the source, the validity, and who can suspend or revoke.

Practiced skill: deliver a short report based on provided results, with comparison, decision, and limits.

Lesson cheat sheet

Submit your short report

  1. Compare the reference with a candidate.
  2. Use held-out cases.
  3. Report gains, failures, and uncertainties.

Your next step

You already know how to deliver a short report based on provided results, with comparison, decision, and limits.

Save the result in your notes as “RSI · lesson 18”. If you skipped the practice, use the template from this page.

Your next cycle: pick a small task and reapply the criteria you recorded.

Additional materialDeeper dive and sources. Outside lesson time.

After the course

Deepen one front at a time: evaluation, tools, data, or automated research. Reapply the method to a task often enough to justify the effort. Consult the primary documents before adopting new claims. The field changes quickly; preserve the habit of asking what changed, how it was measured, and who can verify it. The course research folder records the cutoff date and the sources used.

Sources to check

Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.

Lesson 18 · RSI v6.2 · INEMA.CLUB