RSI v6.2 · 6 modules · 18 lessons
Understand recursive self-improvement, examine real studies, and build a supervised improvement cycle. Short reading, visible examples, and practice in every lesson.

Distinguish RSI, response reviews, and claims without evidence.
Goals, roles, evaluation, memory, and adopting improvements.
AlphaEvolve, DGM, and the pieces of automated research.
Held-out cases, manipulation of measurements and generated data.
Customer service, education, and document summaries with references.
Choose an implementation, plan a pilot, and report its limits.
RSI v6.2
The course’s technical terms in simple words. Each term links to the lessons where it appears.
A space for carrying out a task with defined limits. Its protection depends on the implementation and permissions.
Appears in: Lesson 9
How much a system can decide and carry out without human intervention. It needs a clearly defined scope and explicit limits.
Appears in: Lesson 16
A person or mechanism that applies quality criteria to a result. It needs evidence, not just opinions.
Appears in: Lesson 5
Exploration of variants, with evaluation and selection over rounds. It doesn’t require living organisms.
Appears in: Lesson 8
Examples set aside before making changes and used afterward to evaluate the candidate. They should not guide each review.
Appears in: Lesson 10
An event that ends the test, such as reaching the attempt limit, exceeding cost, or detecting a critical failure.
Appears in: Lesson 17
Accepted rule or information for defined uses, with a source, version, responsible person, and validity period. Approval does not make it infallible.
Appears in: Lesson 6
An observable rule for deciding whether the result meets the goal. It must be defined before comparing versions.
Appears in: Lesson 14
Artificially produced examples, including by AI. They can be used for tests and training, with proper verification.
Appears in: Lesson 12
A record that lets you check a claim. It can be an experiment, its measurement, and its conditions.
Appears in: Lesson 3
The correspondence between the response and the available reference. A faithful response doesn’t add facts without support.
Appears in: Lesson 13
A step that limits the system’s overall performance. Improving another step may bring little total gain.
Appears in: Lesson 7
Educational adaptation of the LOOP-R method: execute, measure, critique, propose, test, validate, promote, and repeat.
Appears in: Lesson 4
Get a good score by exploiting evaluation flaws, without following the task’s intended goal.
Appears in: Lesson 11
Internal values adjusted during training. They are not the messages you write in a conversation.
Appears in: Lesson 2
The ability to link each claim to the material that supports it and to the record of how it was produced.
Appears in: Lesson 15
Recursive self-improvement: improvements in the system help produce new improvements. The acronym comes from the English recursive self-improvement.
Appears in: Lesson 1
Checking the candidate using defined criteria and held-out cases, before deciding whether to use it.
Appears in: Lesson 18
Lesson 1 of 18

You can classify three situations and recognize what would be recursive improvement.
A better answer impresses. But that doesn’t show whether the AI learned to improve itself. Start by discovering what really changed.
In 1 minute
RSI: Recursive self-improvement: improvements in the system help produce new improvements. The acronym comes from the English recursive self-improvement.
An AI can rewrite a message without changing how it works. It can also propose a new work rule, saved for future use. These are different changes. To talk about recursion, look for evidence that the new version helps improve the next versions.
In Marina’s fictional stationery shop, shortening a reply changes the text. Saving a comparison procedure changes future work.
You can take advantage of improvement cycles without building an autonomous AI. In this course, the practices are supervised experiments. They teach you to propose, compare, and decide. They don’t show an intelligence explosion or automatically change the model used in the chat.
Ícaro, a fictional educator, compares instructions to create exercises. He saves the best instruction, but he doesn’t claim he trained a new AI.
Unclear requestImprove how I prepare exercises.
Educational record to compare; it’s not execution of a tool.
Verifiable requestCompare two instructions using the same three texts. Count the questions that can be answered using only each text.
Educational record to compare; it’s not execution of a tool.
A version that praises itself does not prove improvement. Record the previous version, the test, the result, and who decided. Separate capability in a task from the ability to do research. More speed in one step may not change the performance of the whole.
Marina keeps the old and new answer. She only changes the guidance after testing questions she didn’t use in the review.
Control questionWhich part of the system changed?
Educational record to compare; it’s not execution of a tool.
Evidence neededDoes the new version improve results and help you produce future improvements?
Educational record to compare; it’s not execution of a tool.
Test yourself
A tool got better at summarizing. Does that prove RSI?
Stuck here? That's normalThink about the object that changed: was it the message, the rule saved, or the method that produces new versions? Classify one change at a time.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can classify three situations and recognize what would be recursive improvement.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Write on paper or in a notepad. A: the AI rewrote an answer. B: the AI proposed a rule; a person tested it and saved it. C: a system changed its research method; new versions produced verified improvements. Classify: answer review, supervised improvement, or a sign of recursiveness. Explain what still needs to be measured.
Practiced ability: classify three situations and recognize what would be recursive improvement.
Lesson cheat sheet
Self-improvement is a broad term. Here, we distinguish between response review, system improvement, and improvement in the ability to improve. This division is didactic. It’s not a universal classification for the field. Real systems can mix these layers. Notice which component was changed and which result was measured. A good analysis describes the mechanism before assigning a label.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 1 · RSI v6.2 · INEMA.CLUB
Lesson 2 of 18

You can identify three ways to improve a system and choose one for a simple task.
A conversation can look like it learns from you. This impression can come only from earlier messages. Identifying the layer prevents you from expecting a change that didn’t happen.
In 1 minute
parameters: Values internally adjusted during training. They are not the messages you write in a conversation.
The model answers using available instructions and information. A correction in the conversation can change the next response. That does not mean the model’s parameters were updated. When you start another conversation, part of that material may no longer be available.
Ícaro adds a short text about recycling before asking questions. The answer improves because it received the necessary reference.
Create questions about today’s activity.
Use this text: clean paper goes to recycling; greasy paper needs a different destination.
Saving a rule helps only when the system retrieves it at the right time. The rule can grow outdated or contradict another. Name the record, add a date, and specify who is responsible. Before using important information, confirm whether it still holds.
Marina keeps customer service guidance in a document. When the store’s hours change, she updates the guidance and checks old answers.
Incomplete recordWe’re open late.
Educational record to compare; it’s not execution of a tool.
Useful recordExample hours: 9am to 6pm. Check the current version before responding.
Educational record to compare; it’s not execution of a tool.
Training adjusts parameters with examples or evaluation signals. Changing tools or instructions can improve the system without that adjustment. Start with the simplest layer that solves the problem. Don’t use training to compensate for a missing reference.
Ícaro wants the questions to cite the supporting excerpt. He first changes the instruction and measures the result; he does not need to start by training a model.
Test yourself
You opened another conversation and the rule disappeared. What hypothesis do you want to test?
Stuck here? That's normalA new conversation might not get the previous rule. First check what information was sent, before imagining a change in parameters.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can identify three ways to improve a system and choose one for a simple task.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Store case: the AI gives an old time. School case: the AI creates questions without receiving the text. Research case: a lab adjusts parameters with evaluated examples. For each case, write what’s missing and which layer needs to change.
Practiced capacity: identify three ways to improve a system and choose one for a simple task.
Lesson cheat sheet
Some systems combine memory, tools, and later training. So it’s more accurate to avoid the phrase “the chat learns everything” than to say “AI never learns.” Ask what was persisted and how it entered the next use. In professional systems, document the origin and validity of each memory. A false note retrieved repeatedly can spread the same error.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 2 · RSI v6.2 · INEMA.CLUB
Lesson 3 of 18

You can fill in an evidence sheet for a claim about RSI.
News mixes results, goals, and interpretations. A number without context can look like much stronger proof than it really is.
In 1 minute
evidence: A record that lets you verify a claim. It may include an experiment, its measurements, and its conditions.
Separate a verifiable sentence from the opinion that comes with it. Look for the original study or report, its date, and its authors. A video can present good ideas, but its narration does not replace the original document. Also record whether there was external evaluation.
Marina reads “AI does all the research”. She looks for which tasks were done and which decisions stayed human.
Broad statementAI replaced all the research.
Educational record to compare; it’s not execution of a tool.
Concrete questionWhich steps were automated, in which experiment?
Educational record to compare; it’s not execution of a tool.
More hours of execution or more experiments do not directly measure useful discoveries. In September 2026, OpenAI reported greater internal use of agents. It also said that people still choose priorities and decide which results to pursue.
Ícaro can generate a hundred exercises in an afternoon. That does not tell you how many are correct or help his students.
An experiment can have few cases or very specific tasks. A future scenario is a tool for discussing possibilities. It does not establish a guaranteed timeline. Use three labels in your record: observed, interpretation, and prediction. Leave what you couldn’t verify explicitly open.
Marina agrees to test a new routine. She does not change the whole company because of a predicted deadline.
Described result and its conditions.
Whether the gain repeats in other tasks or organizations.
Test yourself
A company announces a goal for 2028. How to record it?
Stuck here? That's normalUse the case of Escola Aurora, already complete in practice. “Not provided” is a valid answer for a missing field.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can fill out an evidence card for a statement about RSI.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious case, provided for practice. Title: Escola Aurora test. Author: teaching team. Date: 25/09/2026. The team compared two instructions in the same five texts. The old instruction generated three questions with correct support. The new one generated four. Each instruction was run once. There was no test with students. Fill in: title; date; authors; statement; measure; comparison; limit. Then, optionally, apply the card to a real source in the additional material.
Practiced skill: fill out an evidence form for a statement about RSI.
Lesson cheat sheet
The RSI project includes three German transcriptions, one user interpretation, and two prior studies. The course preserves the idea of evaluating cycles. Some passages include leaks and comparisons without complete conditions. They do not count as proven results. Numbers from commercial models were discarded when they did not help you understand the mechanism. The full study and the editorial decision are documented in this course.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 3 · RSI v6.2 · INEMA.CLUB
Lesson 4 of 18

You can put together a form with the eight steps of a supervised improvement cycle.
Asking to “improve continuously” leaves almost everything undefined. You need to say what to test, when to accept, and when to stop.
In 1 minute
LOOP-R: Educational adaptation of the LOOP-R method: execute, measure, critique, propose, test, validate, promote, and repeat.
Execute the task with the current guidance and save the result. Define a simple measure tied to the goal. Only then ask for a proposal. Without an initial reference, a convincing answer can seem better just because it is new. Keep the same cases in the comparison.
Marina records whether each answer provides a time, avoids promises without a basis, and indicates the next action.
Criticizing means finding an observable flaw. Proposing means suggesting an adjustment that can fix it. Testing means running again under comparable conditions. If you change instruction, model, and examples all together, it becomes hard to point to a single cause for the result.
Ícaro adds only “cite the supporting excerpt.” He keeps the text, the number of questions, and the previous criteria unchanged.
Change everything and compare the overall impression.
Add a rule and count questions supported by the text.
Validation checks the candidate on held-out cases. Approval authorizes its adoption within written limits. Promotion records that decision and makes the approved version current. In equivalent future uses, the system applies the approved rule without asking again. A new candidate still needs assessment and a decision.
Marina approves the schedule rule for drafts. The team reuses that rule as long as the source and authorization remain valid.
Before the decisionB is a candidate: a good result does not yet authorize adoption.
Educational record to compare; it’s not execution of a tool.
After promotionB becomes the current rule for drafts. Don’t ask again for the same approval.
Educational record to compare; it’s not execution of a tool.
Test yourself
The candidate improved in the test. What’s the next step?
Stuck here? That's normalStart with the flaw: the answer invented holiday hours. The answer key shows the eight filled-in steps.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can put together a worksheet with the eight steps of a supervised improvement cycle.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious reference: store hours are 9:00 a.m. to 6:00 p.m.; holidays not listed. Instruction A: respond politely. Answer A to the question “Is it open on holidays?”: “Yes, from 9:00 a.m. to 6:00 p.m.” Criteria: faithful to the reference; don’t invent information. Candidate B: use only the reference and state when data is missing. Fill in the eight LOOP-R steps with this case. Held-out case to check later: “Is it open on Sunday?”. Sundays aren’t listed. Plan on paper: up to three rounds; Marina decides.
Practice goal: make a worksheet with the eight steps of a supervised improvement cycle.
Lesson cheat sheet
The LOOP-R project has agents with separate roles, experiment logs, versions, and memory. In human-review mode, a change is promoted only after a favorable evaluation and approval. The promoted version then guides future uses. This course applies that discipline with worksheets directly on paper. Approving a new version differs from reusing the current version. The record does not establish universal truth or improvement in every situation. Also read the project’s own critique.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 4 · RSI v6.2 · INEMA.CLUB
Lesson 5 of 18

You can define responsibilities and permissions for a small improvement system.
If the same AI writes the answer and declares it’s perfect, you’re missing a reliable check. Splitting roles makes mistakes easier to spot.
In 1 minute
evaluator: Person or mechanism that applies quality criteria to a result. It needs evidence, not just opinion.
A proposer receives the task, the observed errors, and the limits. They return a candidate and the justification. The evaluator receives cases and criteria, and returns checkable results. The approver reviews quality, cost, and undesired effects before deciding.
Ícaro asks the AI for a new instruction, checks the questions against the text, and decides whether to use the instruction.
Swap “create questions” for “create questions answerable from the excerpt.”
For each question, point out where the text contains the answer.
Separating screens helps you organize the work, but models can repeat the same mistakes. Use verifiable criteria and examples the proposal didn’t receive. For factual statements, check the reference. For subjective results, record disagreements instead of hiding them in a score.
Marina asks the other conversation for criticism. Even so, she compares the time with the store’s document.
Fragile checkAnother AI said it was perfect.
Educational record to compare; it’s not execution of a tool.
Useful checkThe time matches the document and the response didn’t invent a deadline.
Educational record to compare; it’s not execution of a tool.
Whoever approves sets the rule, the permitted uses, and the limits. A recorded approval keeps working for equivalent uses. Don’t ask again just because another run started. Changing the rule or expanding your actions requires a new decision. Approving content does not automatically authorize sending messages.
Ícaro approves a rule for preparing drafts. The system reuses the rule, but it does not change grades or contact families.
Test yourself
Two AIs agreed on a time. What decides?
Stuck here? That's normalTwo opinions don’t change the written time. Open the current reference and compare your answer with it.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can define responsibilities and permissions for a small improvement system.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
On paper or in your notes, write three lines: proposer, evaluator, final responsible person. For each line, record: receives; produces; can change; cannot change. Store case: customer service drafts. School case: question drafts. The evaluation rule is outside the proposer’s scope. Write down a rule that has already been approved and which uses it allows without a new approval. Separate this rule from an external action that keeps being unauthorized.
Practiced capability: define responsibilities and permissions for a small improvement system.
Lesson cheat sheet
When a team automates the design, it will need to turn written limits into real permissions. Separate roles in the request don’t prevent improper access. The team must control files, tools, and actions available. It also needs to record failures and preserve the evaluation criteria. This lesson produces the work specification, not a standalone and secure implementation by itself.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 5 · RSI v6.2 · INEMA.CLUB
Lesson 6 of 18

You can record an approval and tell apart reuse, review, and an action that’s still not authorized.
If the system asks everything again, you waste decisions. If you treat any approval as unlimited permission, you can repeat and expand a mistake.
In 1 minute
approved knowledge: Rule or information accepted for defined uses, with source, version, responsible person, and validity. Approval doesn’t make it infallible.
Separate the proposal, the test result, and the approved rule. In LOOP-R, the promotion links the decision to the current version. Record the exact content, the source, the tests, and who authorized it. Write which tasks the approval applies to. Without that, the system can’t tell approved knowledge apart from an attempt.
Marina approves rule H1: use the current policy’s time in drafts. Holidays still have no information.
Saved attemptH1 passed a test, but has not yet been authorized for use.
Educational record to compare; it’s not execution of a tool.
Approved knowledgeH1, policy P1, Marina, 25/09/2026. Allowed: prepare drafts about schedule.
Educational record to compare; it’s not execution of a tool.
Before using it, the system retrieves the current rule and checks the scope. This check is not a request for approval. If task, source, and validity stay compatible, it continues without asking again. An approved preference can guide dozens of answers; each answer is still subject to the agreed criteria.
Ícaro approves citing the reference excerpt in the questions. The system applies this rule to new drafts without asking for the same authorization.
An approval can power dozens of future decisions. Check content, evidence, coverage, and the cost of being wrong before approving. If the source changes, a contradiction appears, or the validity ends, pause the affected use and review it. Revoking disables the rule, but preserves the history. Human approval can also be wrong.
Policy P2 changes the time to 10h–17h. Marina suspends H1 until she reviews the reference; the system doesn’t insist on the old time.
I know the source, the tests, the allowed uses, and how to revoke.
Policy changed, or an error appeared: stop the affected use and record the reason.
Test yourself
Rule H1 is approved and valid. Another equivalent question arrived. What should you do?
Stuck here? That's normalStart with case A: nothing changed and H1 is still valid. Checking this validity doesn’t mean asking Marina again.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can record an approval and tell apart reuse, review, and action that isn’t authorized yet.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictional case. Policy P1: store opens 9h–18h; holidays not listed. H1: use P1 to prepare drafts about time, without inventing information. Marina checked time and holiday cases and approved H1 on 25/09/2026. Validity: until P1 changes, an error appears, or Marina revokes. Sending is not authorized. Fill in: rule; version; source; evidence; who approved; date; scope; validity; how to revoke. Classify: A) a new time question with P1 in effect; B) P2 changes to 10h–17h; C) send a message to the customer. For each case, choose: reuse without asking; pause and review; request authorization for a new action. Write why approving H1 requires care with future uses.
Practiced capacity: record an approval and distinguish between reuse, review, and action that is not yet authorized.
Lesson cheat sheet
An approved operational rule is a decision about use. A factual statement still depends on sources and evidence. Don’t mix up these two records. Approving an isolated answer also doesn’t mean approving a general rule. Write whether the decision applies only to that result or to reuse. The system must store and look up this record: a conversation without persistent memory won’t automatically come to “know” it. In LOOP-R, the current version and the history help you apply this distinction. New candidates go through another cycle; the already-authorized version doesn’t need repeated approval.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 6 · RSI v6.2 · INEMA.CLUB
Lesson 7 of 18

You can explain a real optimization case and distinguish local gain from total gain.
A headline says something became 23% faster. Before concluding that the whole system sped up like that, find which piece was measured.
In 1 minute
bottleneck: Stage that limits overall performance. Improving another stage can bring little total gain.
AlphaEvolve combines language models, automatic evaluations, and search among candidate programs. The published results include computational optimizations. The problem needs to allow a reliable evaluation. A sleek proposal that produces wrong answers shouldn’t be chosen just because it’s fast.
Marina makes an analogy with the counter: answering faster doesn’t help when the information is wrong.
DeepMind reported a 23% speedup in an operation used by Gemini. The reported effect on training time was 1%. These values refer to different scales. The case shows why you should ask where the gain happened.
Icaro takes less time to format questions. You still need to read the text and check the answers.
Specific operation: reported acceleration of 23%.
Training time: reported reduction of 1%.
In your work, choose a task whose quality can be checked. Define the correct result before measuring speed. Compare total effort, including review. The usefulness of this logic does not depend on having access to the lab’s search tool.
Marina measures the time between receiving the question and approving the answer. She does not measure only the time the chat takes to write.
Partial timeSeconds to generate a draft.
Educational record to compare; it’s not execution of a tool.
Total timeGeneration, checking, and correction until it’s ready to use.
Educational record to compare; it’s not execution of a tool.
Test yourself
Generation was fast, but the review doubled. What should you measure?
Stuck here? That's normalAdd up the time from all steps. In the exercise, 3 minutes saved belong to a 30-minute job.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can explain a real optimization case and tell the difference between local gain and total gain.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious exercise, on paper: preparing an activity takes 30 minutes. There are 6 minutes of formatting and 24 minutes of reading and review. A change reduces formatting to 3 minutes. What is the new total time? How many minutes were saved? What was the total percentage reduction?
Practiced capacity: explain a real optimization case and distinguish local gain from total gain.
Lesson cheat sheet
The case shows the AI’s contribution to improving components used in computing and training. It does not demonstrate unlimited acceleration. It also does not authorize transferring the same percentage to another organization. The operational lesson is to define an evaluator before expanding the search. The practice calculation is an original analogy, with no numerical relation to DeepMind’s experience.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 7 · RSI v6.2 · INEMA.CLUB
Lesson 8 of 18

You can describe what the DGM changes and why it preserves different candidates.
Today’s best attempt doesn’t always become tomorrow’s best. Exploring different paths can reveal improvements that a narrow search would miss.
In 1 minute
evolutionary search: Exploration of variants, with evaluation and selection across rounds. It doesn’t require living organisms.
The DGM uses models to propose changes to the code that organizes the agent. This includes tools and resolution procedures. It does not mean retraining all the language model’s parameters again. Identifying the object helps you avoid confusing a system improvement with a model swap.
In Ícaro’s analogy, changing the checking routine changes the work. It does not automatically change the teacher’s training.
Agent code and procedures.
All model parameters were trained again.
The study keeps a growing collection of agents. A new attempt can start from different variants. A mid-level candidate might contain an idea that helps another round. Preserving diversity doesn’t mean using worse candidates in real life.
Marina saves a short answer and a detailed one between attempts. None becomes the current guidance without passing the criteria.
ExplorationSave one idea to test another combination.
Educational record to compare; it’s not execution of a tool.
Real useAdopt only a validated version.
Educational record to compare; it’s not execution of a tool.
The authors report gains in programming evaluations. The reach depends on the tasks and conditions studied. They also document cases where the evaluation was fooled. That’s why a single score doesn’t end the investigation. Look at the result and the path used to get it.
Ícaro checks whether the question has an answer in the text. A “pass” mark generated by the candidate itself isn’t enough.
Test yourself
An old variant brought a useful idea. What should you do?
Stuck here? That's normalSaving a variant doesn’t mean using it. First, check whether it invents facts in the hard case.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can describe what the DGM changes and why it preserves different candidates.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictional reference: store opens from 9:00 AM to 6:00 PM; holidays not mentioned. Common question: “What are the hours?”. A answers correctly in two sentences. B answers correctly in one sentence. C asks for clarification and doesn’t give the hours. Hard question: “Does it open on holidays?”. A says there’s no data. B says it’s open. C says there’s no data. On paper, apply: faithful to the reference; useful; don’t invent. Decide which variants to save and which to test for use. Create a new validation case.
Practiced capacity: describe what the DGM changes and why it preserves different candidates.
Lesson cheat sheet
Gödel’s theoretical proposal uses proofs of benefit before changes. The DGM works with empirical verification on tasks. This difference matters: passing tests is evidence, not a general proof that every future change will be beneficial. Read the article to understand the controls and conditions. A scientific report should be interpreted together with its limitations.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 8 · RSI v6.2 · INEMA.CLUB
Lesson 9 of 18

You can compare three initiatives without treating all automation as full autonomy.
Producing an article, running an experiment, and choosing an important question are different skills. The course now connects these pieces.
In 1 minute
isolated environment: Execution space with defined limits for a task. Its protection depends on the implementation and permissions.
Sakana reported a piece of work by the AI Scientist accepted in a workshop connected to ICLR 2025. The workshop agreed to participate in the evaluation experiment. This is not the same as acceptance at the main conference. An accepted article also doesn’t prove that every conclusion is correct.
Ícaro distinguishes a draft of an activity from its validation in the classroom. Each step adds different evidence.
Automated stepRun and analyze an experiment with a defined scope.
Educational record to compare; it’s not execution of a tool.
Open questionQuality, reproducibility, and usefulness beyond that experiment.
Educational record to compare; it’s not execution of a tool.
DSec is an infrastructure described by DeepSeek for running agent tasks at scale. Preparing environments, controlling resources, and recording actions are concrete problems. Efficient infrastructure enables more tests. It doesn’t, by itself, prove that a system chooses good questions or improves them recursively.
Marina reserves fictional copies to test responses. Preparing the workstation makes the work easier, but it doesn’t determine which response is good.
Fast research reports don’t remove the need to choose goals. Ask who defines the next experiment and who can stop it. Counting attempts is easy. Measuring scientific contribution means assessing what each attempt actually added.
Ícaro can generate many exercise variations. He keeps deciding which ones meet the lesson’s goal.
More drafts and more attempts.
A useful conclusion that holds up under review.
Test yourself
An infrastructure runs millions of tasks. What does that show?
Stuck here? That's normalUse the three summary cards in this practice. Reading the full articles is optional further study.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can compare three initiatives without treating all automation as full autonomy.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Use these summary cards, based on the sources in the additional material. AI Scientist: automates stages of a research project; an article went through review in a workshop connected to the ICLR. The workshop cooperated with the experiment. DSec: prepares and runs environments for agent tasks. The study describes infrastructure; scale doesn’t prove scientific quality. OpenAI internal report: agents support research tasks. People set priorities and decide which results to pursue. Make three lines: initiative; main function; available evidence; limit. Reading the full documents is optional further study.
Practiced capability: compare three initiatives without treating all automation as full autonomy.
Lesson cheat sheet
The sources of this module cover complementary functions. Publication experience, infrastructure, and internal use reporting do not form joint evidence of autonomous RSI. Don’t add incompatible indicators. OpenAI’s public proposal from September 2026 describes fully autonomous RSI as not happening at that time. It’s a statement from the organization, not a universal inspection of all existing systems.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 9 · RSI v6.2 · INEMA.CLUB
Lesson 10 of 18

You can separate development and validation cases using observable criteria.
If you always correct using the same examples, you can improve only on those examples. The test needs to reveal new situations.
In 1 minute
held-out cases: Examples separated before the changes and used afterward to evaluate the candidate. They shouldn’t guide every review.
Include common situations, missing data, and requests out of scope. Define the expected response or the criteria before running. A small sample teaches the method, but it doesn’t prove overall performance. Avoid choosing only easy cases after you’ve seen the results.
Marina includes one question with no date and another outside the provided time window. This tests whether the response asks for clarification.
Use one group to improve guidance. Set another group aside to check the chosen version. After you consult the held-out results repeatedly, they no longer provide an independent check on unseen cases. For future decisions, prepare other cases and record this replacement.
Ícaro adjusts the instruction using two texts. A third one stays saved until the final comparison.
Text seen during the reviews.
Text separated before and opened only to confirm.
Mark criteria per case, in addition to time and review effort. An average can hide an important failure. Repeat attempts when answers vary. Compare versions under similar conditions. Don’t turn a small test into a promise of accuracy for any task.
One of Marina’s responses was short, but it invented a condition for exchanging a purchase. This error prevents approval even if the response was fast.
Required criterionNo invented deadlines or conditions.
Educational record to compare; it’s not execution of a tool.
Preference criterionClear, short text when the data allows it.
Educational record to compare; it’s not execution of a tool.
Test yourself
You adjusted it ten times while looking at the reserved test. What changed?
Stuck here? That's normalUse the six ready questions. Set the last two aside before making any instruction revisions.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Ready when you can separate development and validation cases using observable criteria.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Made-up reference: store is open from 9am to 6pm; pickup after notice. Exchanges, delivery, holidays, and Sundays are not covered by the reference. Use six ready questions: 1. What are the hours? 2. Can I pick up my order? 3. Do you deliver to my home? 4. Can I exchange it in ten days? 5. Is it open on a holiday? 6. Is it open on Sunday? On paper, separate 1–4 for development and 5–6 for validation. Define criteria: matches the reference; does not invent; asks for confirmation when needed. This lesson prepares the assessment. Don’t send the held-out cases to an instruction review.
Practiced skill: separating development and validation cases using observable criteria.
Lesson cheat sheet
The METR explains that a task horizon corresponds to the human time of tasks completed with a certain success rate. It’s not the time an AI works alone. The estimates have uncertainty and vary across domains. This caution also applies to your own test: a useful measure must say what it measures, which tasks it includes, and which conclusions it does not allow.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 10 · RSI v6.2 · INEMA.CLUB
Lesson 11 of 18

You can identify three shortcuts that fool the assessment and propose protection for each one.
A system pushed by a measure can optimize the measure and abandon the goal. This happens when the rule allows shortcuts.
In 1 minute
reward manipulation: Get a high score by exploiting flaws in the assessment, without fulfilling the task’s intent.
If the goal is only to answer quickly, leaving out information can seem like a win. If it’s to increase quantity, duplicating results can seem like progress. Also write what must stay correct. Combine the main measure with criteria that prevent these losses.
Marina won’t accept empty messages to reduce customer service time. The response must resolve the question or ask for the missing information.
Answer in a few words.
Answer faithfully and concisely, without leaving out what’s necessary.
A tool can claim it checked something without producing evidence. Keep the observable result of the check. Distinguish a teaching simulation from real execution. When there isn’t enough proof, write “not verified” instead of turning trust into approval.
Ícaro asks for the excerpt that supports each answer. “Everything checked” doesn’t replace finding the sentences.
StatementI checked, and everything is correct.
Educational record to compare; it’s not execution of a tool.
EvidenceQuestion 2: the answer appears in the second sentence of the text.
Educational record to compare; it’s not execution of a tool.
A candidate system should not make its own test easier. Preserve the cases, the answer key, and the original results. If a rule is wrong, revise it in a separate process and reevaluate all versions. Don’t change criteria just to save your preferred candidate.
Marina finds an ambiguous criterion. She corrects it and applies it again to both versions, without favoring the new one.
Test yourself
The grade went up after deleting difficult cases. What can you conclude?
Stuck here? That's normalAsk whether the task was completed even without looking at the grade. Then look for the evidence that would let you verify it.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Ready when you can identify three shortcuts that fool the assessment and propose a protection for each one.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Case A: they assess quantity; the AI duplicates questions. Case B: they assess speed; the AI omits essential conditions. Case C: they ask for checking; the AI only writes “checked”. On paper, match each failure to a concrete protection.
Practical capacity: identify three shortcuts that fool the assessment and propose a protection for each one.
Lesson cheat sheet
Sakana described cases of fabricated test records and changing the markers used in the assessment. Anthropic studied reward manipulation in controlled scenarios. These findings call for independent verification. They don’t mean that all AI has a human intention to cheat. The practical focus is to detect behaviors that produce points without meeting the goal.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 11 · RSI v6.2 · INEMA.CLUB
Lesson 12 of 18

You can evaluate a small set of made-up examples before using it to improve a routine.
Generating a thousand examples is easy. Knowing whether they’re correct, varied, and suitable for the work requires a different evaluation.
In 1 minute
synthetic data: Examples produced artificially, including by AI. They can serve for tests and training, with proper verification.
Choose categories needed for the task. Include common situations and relevant exceptions. Mark which examples were generated and who verified them. A collection of nearly identical variations increases volume without greatly expanding coverage.
Ícaro notices that ten generated questions ask for the same information. He swaps part of them for questions that assess other ideas from the text.
VolumeTen versions of the same question.
Educational record to compare; it’s not execution of a tool.
CoverageQuestions about different points and one with missing information.
Educational record to compare; it’s not execution of a tool.
Studies show degradation under certain conditions of recursive training with generated data. This doesn’t mean that all synthetic data causes collapse. The composition, selection, and preservation of diversity matter. Don’t replace verified references with answers generated without checking.
Marina uses the schedule document as a reference. She doesn’t treat an old chat response as the official source.
An AI can compare results and locate problems. But it may prefer long text or agree with plausible errors. Have a person check a sample using objective criteria. When the judgment is uncertain, keep the disagreement instead of fabricating a conclusion.
Icaro uses automatic criticism to locate suspicious questions. He checks the text before accepting the evaluation.
The AI points out possible problems.
The person checks reference, coverage, and appropriateness.
Test yourself
One hundred repeated examples repeat the same mistake. What to do?
Stuck here? That's normalCompare questions 1 and 2. They use different words, but they ask for the same information.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can evaluate a small set of fictional examples before using it to improve a routine.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictional text: Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday. Fictitious generated questions: 1. On what day does collection happen? 2. What is the collection day? 3. How should greasy paper be treated? 4. Can clean paper be recycled? 5. How many tons are collected? On the paper, mark source, support, repetition, and missing information. Replace question 2 with “What condition of the paper requires separation?”. Compare with question 3 to discuss redundancy. Decide whether it’s better to stick with three distinct questions instead of five.
Practiced skill: evaluating a small set of fictional examples before using it to improve a routine.
Lesson cheat sheet
The collapse study analyzes loss of information in recursive training with generated data. The work on self-rewarding models investigates using the model itself to produce evaluation signals. These are different problems. Promising results in one don’t remove the risks of the other. For your test, preserve references and never confuse automatic preference with factual truth.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 12 · RSI v6.2 · INEMA.CLUB
Lesson 13 of 18

You can compare two customer service instructions using a fictional policy and reserved questions.
A polite response can promise something the company doesn’t offer. The cycle needs to improve usefulness without inventing commitments.
In 1 minute
fidelity: Correspondence between the response and the available reference. A faithful response doesn’t add facts without support.
Use a fictional policy to learn the method. Define which information is available. Ask the response to acknowledge gaps. A cordial tone instruction doesn’t replace facts. The first test should reveal whether the system distinguishes known information from unknown information.
Marina tests an imaginary stationery store. The reference includes opening hours and pickup, but it doesn’t mention the exchange period.
Compare a basic instruction with another that requires indicating missing data. Keep the model and the reference. Use separate conversations or clearly identify each attempt. If the tool isn’t available, manually compare the provided drafts.
Ícaro applies the same logic to school announcements. A lack of information doesn’t turn into imagined authorization.
Yes, the exchange is allowed in ten days.
The reference doesn’t say the exchange period. Confirm the policy with the store.
After adjusting the instruction, test a reserved question. Look for made-up facts and confusing actions. Also note whether the answer resolves what it can resolve. Saying “I don’t know” to everything avoids inventions, but it can make the service useless.
Marina asks when she can pick up an order. The answer should use the available condition: after the notice.
Reserved questionCan I pick up now?
Educational record to compare; it’s not execution of a tool.
Supported answerPickup happens after the notice. Have you already received that notice?
Educational record to compare; it’s not execution of a tool.
Test yourself
The AI avoids making things up but refuses even to state the opening hours. What is the problem?
Stuck here? That’s normalDo only the development first. Validation has a separate block that you open afterward.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can compare two customer service instructions using a fictional policy and reserved questions.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Copy only this block for the development stage. Fictitious reference: opening times from 9:00 a.m. to 6:00 p.m.; pickup after notice; exchanges not provided. Compare two instructions, without changing them: A: answer the question politely. B: answer using only the reference; if any data is missing, explain and ask for confirmation. Development questions: What are the opening hours? Can I exchange an item after ten days? Show the answer for each version for each question. Stop after these cases. Don’t approve the candidate yet.
Only after closing the development, start a new conversation. Paste the reference and the chosen instruction, without previous criticism. Ask: “Can I pick it up now?”. Check that the answer asks about the notice. Don’t change the candidate after seeing this result; record the decision.
Paper alternative: A answers “Yes, you can pick it up now”. B answers “Pickup happens after notice. Have you already received the notice?”. Apply the same criteria. Record that you analyzed drafts provided, without running a chat.
Practiced skill: comparing two customer service instructions using a fictitious policy and reserved questions.
Lesson cheat sheet
The policy from the exercise was invented for teaching. It does not describe a real company or establish consumer rights. For professional use, replace it with the organization’s authorized reference and keep a responsible person to review it. Don’t use the exercise as legal advice. The improvement shown is local and supervised, with no training of a new model.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 13 · RSI v6.2 · INEMA.CLUB
Lesson 14 of 18

You can create and check three questions with answers supported by a short text.
A question can look good and ask for something the student didn’t receive. The goal is to check the link between text, question, and answer.
In 1 minute
quality criterion: Observable rule to decide whether the result meets the learning objective. It must be defined before comparing versions.
Choose a small skill, like locating explicitly stated information. Don’t mix this goal with external research without telling. Provide the text and the number of questions. The assessment must match the chosen skill, not just the look of a school activity.
Ícaro wants the class to locate recycling conditions. He uses a short text and does not require outside knowledge.
Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday.
Locate information that appears explicitly in the text.
Ask for the question, the answer, and the supporting excerpt. Check whether the excerpt truly supports the answer. Copying any sentence doesn’t solve it. Require the tool to discard questions without support in the material. You can apply this procedure by hand, without depending on a chat.
Marina applies the same logic in a team training. Each question must point back to the guidance that was provided.
Question without supportHow many tons does the collection pick up per month?
Educational record to compare; it’s not execution of a tool.
Question supportedOn what day does the collection happen? Answer: Tuesday. Support: the last sentence.
Educational record to compare; it’s not execution of a tool.
Three almost identical questions don’t assess three different ideas. Check repetition, clarity, and difficulty. The tool may suggest a revision, but it doesn’t automatically know the class. For a next round, use another reserved text and keep the same criteria.
Ícaro rejects two questions that ask about the same day. He adds a question about the condition of the paper.
Test yourself
The question cites a sentence, but it doesn’t answer the question. What should you do?
Stuck here? That’s normalIf the questions are already correct, don’t invent a defect. Record the check and preserve the version.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Done when you can create and check three questions with answers supported by a short text.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious text: Clean paper can be recycled. Greasy paper must be separated. Collection happens on Tuesday. Create three questions to locate explicitly stated information. For each one, provide the answer and the exact support sentence. Cover three different pieces of information. Don’t add outside knowledge. Review every link between question, answer, and excerpt.
Practiced capacity: create and check three questions with answers supported by a short text.
Lesson cheat sheet
The exercise checks the quality of the material you produced. It doesn’t measure what students learned. To assess learning, the educator will need to observe answers, difficulties, and the group’s context. They should also avoid using identifiable data in unnecessary experiments. The review method helps you prepare the activity; it doesn’t replace teaching judgment.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 14 · RSI v6.2 · INEMA.CLUB
Lesson 15 of 18

You can produce a list of actions that keeps responsible people, deadlines, and missing information.
Smooth summaries can invent a date or assign a task to the wrong person. Here, quality will be traceable back to the original note.
In 1 minute
traceability: Possibility of linking each statement to the material that supports it and to the record of how it was produced.
Extraction recovers information that is already present. Suggesting adds a possibility. Both actions can be useful, as long as they’re identified. Ask for a factual list first. If you want suggestions, put them in a separate part, without presenting them as decisions already made.
Marina wants to turn meeting notes into actions. She doesn’t want the AI to choose deadlines to fill in empty spaces.
Fictitious noteMarina will review the stock by Friday. Ícaro will prepare an activity; the deadline is pending.
Educational record to compare; it’s not execution of a tool.
Faithful extractionMarina: stock, by Friday. Ícaro: activity, pending deadline.
Educational record to compare; it’s not execution of a tool.
Use “not provided” when the reference doesn’t contain the data. Don’t confuse that label with a lack of capability. It shows the limit of the material. A list with explicit blanks can be more useful than a complete, invented summary.
Ícaro prefers to receive “pending deadline”. That way, they know which question to ask before planning the activity.
Evaluate person, action, and deadline separately. Count omissions and additions without support. Then check clarity. When you review the guidance, change only one rule. Keep the original note available so another person can verify the list.
Marina checks each line against the note. The review finds the date copied from another task incorrectly.
Correct responsible person; correct action; deadline faithful to the note or missing.
A deadline from one task applied incorrectly to another.
Test yourself
The summary includes a deadline missing from the note. What’s the decision?
Stuck here? That's normalMarina’s deadline doesn’t apply to Ícaro. Mark his deadline as not provided and write the missing question.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can produce a list of actions that preserves responsible people, deadlines, and missing information.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious notes: Marina will review the inventory by Friday. Ícaro will prepare an activity; the deadline is still pending. Produce a list with responsible person, action, deadline, and supporting excerpt. Write “not provided” when data is missing. Don’t suggest dates or treat suggestions as decisions. Check each field against the original notes.
Practiced capability: produce a list of actions that preserves responsible people, deadlines, and missing information.
Lesson cheat sheet
This outline serves as a starting point for syntheses, to-do lists, and review of materials. Documents that guide relevant decisions require review proportional to the impact. Keep access to the reference and don’t hide uncertainties. If a tool changes its behavior, reapply your validation cases before trusting the old guidance.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 15 · RSI v6.2 · INEMA.CLUB
Lesson 16 of 18

You can decide the next level of implementation based on need, cost, and your ability to verify.
The desire to automate can show up before a good evaluation exists. A small experience helps you decide whether it’s worth investing.
In 1 minute
autonomy: How much a system can decide and carry out without human intervention. It needs clear scope and limits.
In your first attempts, copy the material, compare answers, and record results. This makes the method visible. You learn which cases fail before you expand execution. You don’t need to buy a special tool to do the exercises in this course.
Marina compares two approaches on a sheet. She only considers automation after she understands which failures she needs to detect.
An automatic routine needs to receive data, run proposals, evaluate, and record. Someone maintains the limits and handles failures. Before you hire or build, write down these needs. The provider must show how the result can be checked and how to go back to the previous version.
Ícaro specifies that the routine only prepares questions. The educational approval stays with him.
I want an AI that improves on its own.
I want to compare instructions with fixed cases and review candidates before adopting them.
Modifying agents and running training involves specialization, extra resources, and additional risks. Don’t treat lab examples as a ready-made recipe for any company. Choose the smallest level that meets the need. Local gains can be useful without broad autonomy.
Marina doesn’t need to build a lab to reduce service mistakes. First, validate the guidance used in the drafts.
Local needImprove a repeated, bounded task.
Educational record to compare; it’s not execution of a tool.
Research projectInvestigate new methods with infrastructure and specialized evaluation.
Educational record to compare; it’s not execution of a tool.
Test yourself
There’s still no quality criterion. What’s the next step?
Stuck here? That’s normalChoose one of the three ready-made cases. The decision applies to those conditions, not to every company or school.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can decide the next implementation level based on need, cost, and how easy it is to check.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Choose a fictional case with complete data: Store: 4 questions per week; short policy; Marina reviews everything; additional budget is zero; wrong answers need to be corrected before sending. School: 20 activities per week; Ícaro reviews; technical team available for 2 hours per week; the material must have text support. Research: dedicated technical team; agent comparison; isolated environment and own budget; results evaluated with tests. On paper, recommend a manual trial, a limited routine, or research. Justify using frequency, evaluation, resources, and the cost of failing. Say what data you would measure before investing more.
Practiced skill: decide the next implementation level based on needs, cost, and the ability to verify results.
Lesson cheat sheet
Bring your experiment sheet and failure examples. Ask for a demonstration with your authorized cases, including the hard ones. Request evidence of cost control and stopping. Don’t accept only one demonstration selected by the vendor. The decision depends on the task and the real conditions, not on a general ranking of models.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 16 · RSI v6.2 · INEMA.CLUB
Lesson 17 of 18

You can fill in a pilot plan with goal, metric, owner, cap, and reversal.
A small improvement can cost more than the benefit. Without a written limit, attempts pile up and the decision gets postponed.
In 1 minute
stop condition: Event that ends the test, like reaching the limit of attempts, going over cost, or detecting a critical failure.
Add time for preparation, execution, and review. If there’s a charge based on usage, record the observed value. Don’t confuse the advertised price with the total cost of the work. Compare recurring savings with the effort to prepare and maintain the change.
Marina estimates how much time she saves per response and how much she spent preparing the test. A rare improvement may not cover the preparation.
Prepare the pilot: 40 minutes. Savings per use: 2 minutes.
You need 20 uses to recover 40 minutes, not counting maintenance.
For a first test, choose a few cases and up to three candidates. Set a work window. Define a financial cap if the tool charges. Record educational examples and values that were truly measured separately. Don’t move the exercise budget to any project.
Ícaro reserves a session to compare instructions. If there isn’t time to check, he doesn’t approve the candidate out of haste.
Exercise ceilingThree candidates, six cases, and a planned session.
Educational record to compare; it’s not execution of a tool.
Decision ruleWithout a completed validation, keep the previous version.
Educational record to compare; it’s not execution of a tool.
Keep the previous guidance and indicate who can restore it. Stop when you come across invented relevant information, an unauthorized action, or a cost above the ceiling. Investigate the failure before repeating. If the system changes, reassess the pilot conditions.
Marina keeps the current guidance. A candidate who creates improper commitments is rejected before reaching the team.
Test yourself
The budget ran out before validation. What should you do?
Stuck here? That's normalFirst calculate without maintenance: 40 divided by 2. Then include the monthly effort to see if the conclusion changes.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. You’re done when you can fill out a pilot plan with goal, measure, person responsible, ceiling, and reversal.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
Fictitious plan to complete on paper: Marina prepares 20 answers per month. The pilot requires 40 minutes of preparation. Estimated savings: 2 minutes per use. Policy review: 10 minutes per month. Tool already available; additional spending cap: zero. No message sending. Reference: open 9h–18h; pickup after notification; other information missing. Development cases: opening hours, pickup, exchanges, and delivery. Held-out cases: holidays and Sundays. Required criteria: fidelity and no inventing. Up to three candidates. Complete: goal; responsible person; stop point; rollback; next review. Calculate uses to recover preparation without maintenance. Then discuss the effect of the 10 minutes per month. Also record: approved knowledge; reuse scope without new approval; validity; who can revoke.
Practiced capacity: fill out a pilot plan with goal, measure, person responsible, ceiling, and reversal.
Lesson cheat sheet
This course pilot does not send messages, does not modify external systems, and does not use personal data. Even so, it needs a responsible person and criteria. In real projects, increase controls based on impact and autonomy. Proposals for international standards discuss supervision and evidence; they don’t replace your organization’s concrete obligations.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 17 · RSI v6.2 · INEMA.CLUB
Lesson 18 of 18

You can deliver a short report based on provided results, with comparison, decision, and limits.
The final result does not need to be a successful candidate. Finding out that a change isn’t worth it is also a useful conclusion.
In 1 minute
validation: Candidate review with defined criteria and held-out cases, before deciding whether to use it.
Use support services, reading questions, or action extraction. If you didn’t do the previous practices, copy one of the fictional materials from this module. Write the initial reference and one change. You can run the trial with an AI chat or compare drafts manually.
Ícaro chooses questions that can be answered by the text. His only change is requiring a supporting passage.
ReferenceCreate three questions about the provided text.
Educational record to compare; it’s not execution of a tool.
CandidateCreate three questions and present the sentence that supports each answer.
Educational record to compare; it’s not execution of a tool.
Apply both guidelines to the same cases. Keep the responses, the criteria, and the observed times. Then use the held-out cases without continuing to adjust the candidate. If there’s any variation, record it. A small difference in a short trial can be inconclusive.
Marina tests common questions and gaps. She keeps a worse answer alongside the best ones, so she doesn’t select only wins.
Conclude: approve for limited use, reject, or investigate more. If you approve, record the accepted knowledge, the version, and the future authorized uses. The same valid rule doesn’t require approval every time you repeat it. New changes go through review. State what the test didn’t prove and how to suspend or revoke the decision.
Icaro approves a rule for drafts with pedagogical review. This rule does not need new approval for every activity.
Rule B approved to prepare drafts, with source and version recorded.
It does not allow automatic sending or prove learning improvement.
Test yourself
The candidate did not bring reliable gain. What conclusion is valid?
Stuck here? That's normalUse the provided results and say they are educational. Running your own experiment is an optional later session.
Practice now 0/3
About 10 minutes, on your phone, computer, or paper. Ready when you can submit a short report based on provided results, with comparison, decision, and limits.
Use fictional materials. Don’t send personal data or real messages. If something comes out strange, compare with the reference and record the failure.
10-minute path: analyze the fictional results below. These are not real measurements. Reference: store opens 9am–6pm; holidays not listed. A: respond politely. B: use only the reference and report gaps. Development — time: A and B respond 9am–6pm. Development — exchange: A promises ten days; B says that information is missing. Held-out — holiday: A claims it is open; B reports that data is missing. Fictional review times per case: A = 2, 3, 3 minutes; B = 1, 1, 1 minute. Financial spend not measured. Criteria: fidelity; no inventing; usefulness. Write: hypothesis; single change; comparison; failures; time; decision; responsible person; rollback; limit. Optional: then do your own run in another session, using new held-out cases. If you decide to approve: record knowledge, version, source, responsible person, scope, validity, and revocation. Tell which equivalent use does not require new approval and which change requires review.
Practiced skill: deliver a short report based on provided results, with comparison, decision, and limits.
Lesson cheat sheet
Deepen one front at a time: evaluation, tools, data, or automated research. Reapply the method to a task often enough to justify the effort. Consult the primary documents before adopting new claims. The field changes quickly; preserve the habit of asking what changed, how it was measured, and who can verify it. The course research folder records the cutoff date and the sources used.
Consultation: September 25, 2026. Professional examples and exercise numbers are fictional.
Lesson 18 · RSI v6.2 · INEMA.CLUB