LOOP-R · Track 03
Leave knowing who handles each part of the cycle in your business, with answer 3 of the loop worksheet written: what never changes without you.
Lessons
Leave knowing how to run the Critic on one of your proposals and read its output—without letting it make suggestions.
Leave able to identify, in a real case, who “convinced themselves”—and why that never becomes an official version.
Leave with answer 3 on the loop worksheet written: 3 invariants and 2 guardrails for your process.
Leave with the first line of your loop history written: what you tried, what happened, and why.
Track 03 · Lesson 1
By the end of this lesson, you will be able to run the Critic on one of your proposals and distinguish strong evidence from guesswork in its response—without letting it make any suggestions.
Today you ask AI for a proposal and, in the same conversation, ask, “Was it good?” It always thinks so. The person doing the work never evaluates their own work—not on your team and not in the loop. Separating the roles is what brings the evidence to light.
↓ scroll to study
LOOP-R divides the cycle into nine assistants, each with a single role. Think of an employee who does one thing and does it well. They do not decide what the colleague next to them should have done.
The rule applies to everyone. Each assistant reads only what its role requires, writes only that role’s result, and refuses anything outside it. It may sound bureaucratic. It is the opposite: the only way to stop an AI from “convincing itself” that it has improved.
Applied example: a support manager at a small software company has one person answer tickets and another count how many tickets were reopened. If it were the same person, the reopen count would never go up—no one likes counting their own mistakes. In the loop, the AI that writes the reply is not the one that counts the outcome.
The person who does the work does not measure it. The person who measures does not explain it. The person who explains does not propose changes.
O Executor is the only assistant that produces the actual deliverable: the proposal, customer reply, or campaign. It follows the official version’s script exactly and records one row in the log: date, version, and what was produced.
What it never does: change its own script, depart from the designated version, or ignore what must never happen. And—this is the strangest part for anyone used to chat—it does not evaluate its work or suggest improvements. It only executes.
Applied example: a real estate agent receives leads from the portal. The Executor replies to each lead using the agency’s current script (version 1) and records the date, version 1, property, neighborhood, and response time. If the agent says, “reply better,” the Executor refuses: “better” is not an instruction; it is a verdict, and it comes from another assistant.
Test yourself
The Executor wrote a proposal and it seems weak. What is the right next step in the loop?
O Observer turns the rows in the spreadsheet into evidence. It counts, groups, and checks that the rows are complete. If a column is missing or a row is duplicated, it writes “inconsistent data” and stops the cycle there.
What it never does: explain the cause, propose a change, or use information from outside the spreadsheet. Its answer is a table of numbers and counts. No “it seems.” If a number has no count to support it, it is not evidence—it is an impression.
Applied example: the support manager’s company received 10,000 tickets in a year. The Observer reports that 37% concern the same issue, gives the average response time, and counts how many were reopened within 7 days. It does not say, “the product has a serious defect”—that would be an explanation. It only provides the counts for the next assistant to read.
O Critic reads the Observer’s table and the current script version. It says what works, what fails, and where there is waste—always with the number that proves it. It also labels each claim with an evidence strength: strong, moderate, or weak.
Weak claims go in a separate section, “observed, inconclusive.” The Critic also checks the memory: has something that “fails” today already been resolved? And what it never does is propose changes. A diagnosis is not a prescription.
Applied example: the reference example is a beauty clinic sending proposals over WhatsApp, across four cycles using simulated data. The owner’s worksheet stated a current conversion rate of 3%. The Critic reviewed 200 proposals and identified three findings. Observed conversion was 2.5% (5 of 200), below the stated rate. The response rate was 17% (34 of 200). And 83% of proposals got no reply within 48 hours—nearly all the effort went into messages that received no response. Every finding was labeled strong. The Critic did not suggest shortening the message. It only showed where the problem was.
Before—the owner reads and jumps to a conclusion
The clinic owner looks at 200 proposals and feels that “the Tuesday ones get more replies.” She changes the send day. Nothing is recorded or counted.
After—the Critic reads the count
“83% received no reply within 48 hours (166 of 200)—strong. Conversion is 2.5%, below the stated 3%—strong. The mid-sized segment replies at 10%, half the rate of the others—moderate.”
Takeaway: In this example, 3 claims have counts and strength labels, with 0 suggestions—and a discrepancy (3% stated vs. 2.5% observed) that no one had noticed.
Practice now 0/4 complete
Goal: in about 10 minutes, get a diagnosis from your AI chat with evidence strength labeled and no suggestions.
Nothing in your business changes until you approve it: you are only reading a diagnosis in a chat. If the answer contains suggestions or adjectives, paste the request again and add: “you made suggestions—delete them and redo only the diagnosis.”
Before pasting: this is an instruction sheet for a new employee. The first line states their role; the middle says what they receive; the end says what they never do. Replace only the text between the angle brackets.
You are the Critic for the process "<name of your process, e.g., replying to portal leads>". Your role is diagnosis. You do NOT propose a solution, suggest an improvement, or rewrite anything. Proposal/reply sent (current version): <paste here the text your AI or team sent> What happened afterward (only what you have recorded): <e.g., 40 sent this month, 7 replied, 1 sale, 0 complaints> Reply in 3 sections: 1. What the current version does well—with the number that proves it. 2. What fails and where there is waste—with the number that proves it. 3. Label the strength of each claim: STRONG (happened many times), MODERATE (the minimum), or WEAK (few times). Put WEAK claims in a separate section called "Observed, inconclusive." If I did not give a number to support a claim, write "no count—inconclusive." Do not propose anything.
You have just received a diagnosis with evidence strength labeled and no recommendations built in—the first honest review of one of your proposals.
Summary
Track 03 · Lesson 2
By the end of this lesson, you will be able to identify, in a case from your business, who “convinced themselves”—and say which of the three assistants would have prevented it.
The most common change in a small business starts like this: someone has an idea, tests it for two weeks, “feels” that it worked, and turns it into a rule. No sample size, no comparison, no judge. The loop puts the idea, test, and verdict in three different hands—and the third hand can say only one of three things.
↓ scroll to study
O Optimizer reads the Critic’s diagnosis and writes hypotheses. Each hypothesis has a fixed structure. IF one specific change to the script, THEN the metric should move from this value to that value, BECAUSE the evidence is this (strong or moderate). One change per hypothesis. Never two.
Before writing, it checks the drawer of discarded hypotheses: anything that already lost or was vetoed does not come back. Hypotheses based only on weak evidence are not tested; they go into “observed, untested” and wait for a larger sample.
Applied example: the real estate agent received two numbers from the Critic. 70% of portal leads never reply to the first message, and the average response time is 3 hours. The Optimizer writes: “IF the first reply is sent within 15 minutes, THEN the response rate should rise from 30% to 40%.” It adds: “BECAUSE 70% silence is strong evidence and the lead goes cold.” One change only. It did not write “reply quickly, add more photos, and include the price”—that would mix three hypotheses.
O Experimenter takes the approved hypothesis and designs the test: A is the current version, and B is the version with the change. Before sending anything, it calculates the sample size needed for proof —how many runs each side needs so the difference is not just luck.
Then it splits the work evenly, alternating by arrival order—the Executor does not choose who gets A or B. It also sets the stopping rule: the test can stop early only if a guardrail is breached. Never because “it is already winning.”
Applied example: in the reference clinic, the hypothesis that “a message of no more than 80 words” would raise replies from 17% to 22%. The Experimenter calculated that proving such a small lift would require 985 proposals per side. At the clinic’s pace, that meant 40 weeks, above the 16-week ceiling. It did not lower the sample size on its own. It took three options to the owner. The decision was to measure a worthwhile lift (17% to 30%): 166 per side, 7 weeks. The test became feasible because someone did the math before, not after.
Before—the “gut-feel” test
The clinic changes three things in the message at once, sends it for two weeks, checks WhatsApp, and concludes that it “improved.” No one knows which change affected the result—or whether it was luck.
After—the designed test
One change. Two groups, split evenly by arrival order. 166 proposals per group before any verdict. Early stopping only if a guardrail is breached.
Takeaway: In the example, the test went from an impossible 40 weeks to a feasible 7—and the calculation was recorded for anyone to check.
O Evaluator receives the test numbers and gives exactly one of three possible answers. Insufficient sample: one side has not reached the minimum yet—continue. A stays: B did not beat A by a sufficient margin, or B worsened a guardrail—even if the target metric went up. B won: B beat A by a sufficient margin and worsened nothing.
There is no “B looks better.” There is no “B won, but…” The third answer, “insufficient sample,” is the most common—and must be acceptable. An honest loop spends many weeks saying, “I still do not know.”
Applied example: in the clinic’s second cycle, version B used “up to 80 words.” It raised the response rate from 19.3% to 29.7%, with 300 proposals on each side. A real lift, not luck. And the Evaluator answered: A stays. Reason: there were 6 complaints out of 300 for version B, versus 1 out of 300 for version A, and the worksheet set zero tolerance for complaints. The rule was applied as written. (In the Guardian lesson, you will see that the rule was poorly calibrated—and how the clinic corrected it without changing the verdict.)
Test yourself
A support manager’s test is in week 3. B resolves 70% on first contact versus 55% for A—but each side has only 40 tickets, below the 170 minimum. What does the Evaluator answer?
Put the three lessons together. The person who does the work does not measure; the one who measures does not explain; the one who explains does not propose. Now add this: the one who proposes does not test, and the one who tests does not judge. Whenever someone “convinces themselves,” the last three roles were in the same hands. That person had the idea, chose how to test it, and decided it had won.
The separation is not about distrusting people. It accounts for a human—and AI—pattern: whoever has the idea wants it to work and sees the result they expect. The loop does not ask you to be neutral. It asks for a different judge.
Applied example: a support manager changes the standard reply for a recurring issue. The next week, they “feel” that reopenings have fallen and tell the team to adopt it. They acted as Optimizer (the idea), Experimenter (one week, uncounted, with no version A running alongside), and Evaluator (the verdict that they “fell”). In the loop, the idea would become a hypothesis. A and B would run side by side until each had 170 tickets. The verdict would come from someone who did not write the reply.
The idea can be yours. The test and verdict never are.
Practice now 0/3 complete
Goal: in about 8 minutes, read the case, write answers to the 3 questions, and compare them with the answer key.
This is safe: the case is about another company, and nothing changes in your business until you approve it—you are only practicing how to spot the pattern. If your answer does not match the key, reread step 04 and try again. The difference is usually who delivered the verdict.
The case. Marcos manages support at a software company with 6 agents. He noticed that tickets about “I can’t export the report” were reopened most often. He wrote a new reply with a short video and asked everyone to use it starting Monday. On Friday, he checked the dashboard: 3 of the 23 new replies had been reopened—“it used to be much more.” He made the new reply the standard and deleted the old one.
1. All three. He proposed: “wrote a new reply.” He tested it his own way: “asked everyone to use it starting Monday,” with no version A running alongside. He judged: “it used to be much more,” without counting the baseline. The clearest sign is “deleted the old one”: there is no A left to compare or roll back to.
2. One hypothesis with one change—the new reply has new wording e and a video, so there are two changes. Calculate the required sample size in advance: at his pace, it would probably take weeks, not days. Half the tickets should keep using the old reply, alternating by arrival order. And the test can stop only if a guardrail is breached.
3. “Insufficient sample.” Twenty-three replies in one week do not meet the minimum for any reasonable test—and there is no count for A during the same period. Three out of 23 could be luck, a quiet day, or a real effect: no one knows. Marcos turned a guess into a rule and erased the evidence.
You have just separated the idea, test, and verdict in a real case—and identified the exact moment the evidence was erased.
Summary
Track 03 · Lesson 3
By the end of this lesson, you will have written answer 3 on the worksheet: 3 things that never change without you and 2 metrics that must not get worse.
Any loop learns to game its target unless someone states what is forbidden. “More replies” can become an irritating message; “faster support” can become a ticket closed without a solution. The Guardian exists to make AI improve the right metric in the right way. It works only if you write down what it must watch.
↓ scroll to study
O Guardian reads the Optimizer’s hypotheses before any test. For each one, it gives a one-word answer: approved or vetoed. It is the only assistant whose decision is not reviewed within the cycle. Arguments do not persuade it; it does not approve “with reservations” or suggest alternatives. When in doubt, it vetoes.
It sounds strict because it is. The Guardian is not there to exercise common sense; it is there to apply the list you wrote, consistently. Common sense is yours when you write the list. After that, the Guardian only checks.
Applied example: the real estate agent and support manager write different lists, and each Guardian watches different things—but applies the rules in the same way.
In your profession
The Guardian’s first list is the set of invariants: things no hypothesis may touch under any circumstances. They are not targets; they are boundaries: safety, law, promises the company does not make, and people who asked not to be contacted.
A good invariant is short, concrete, and starts with “never.” If it takes two sentences to explain, it is still an opinion—not a boundary. You write it once on the loop worksheet. The Guardian does not invent invariants; it only applies them.
Before looking at the template: what follows is the clinic’s section of the worksheet. Read it like a list of “house rules” posted at reception. You will write your own at the end of this lesson.
Applied example: the real estate agent writes three lines on the worksheet: “Never promise loan approval.” “Never quote condo fees or property tax without checking the property record.” “Never contact anyone who asked not to receive messages.” The hypothesis “reply immediately with an approved loan simulation” never reaches the test—the Guardian vetoes it on review because it violates the first invariant. For comparison, the clinic has four: never discount more than 10%; never promise a clinical result or timeline; never contact anyone who asked not to receive messages; and never name competitors.
The second list is the set of “without getting worse”. Every target in the loop takes the form “increase X without worsening Y and Z.” Without that escort, the loop learns to game the target: it raises X by the cheapest route, which usually damages Y.
The Guardian uses this list in two ways. Before the test, it vetoes hypotheses that “predictably” worsen Y. After the test, the Evaluator rejects B if Y falls beyond the tolerance—even if X is way up, as you saw in the previous lesson.
Applied example: in the clinic’s first cycle, the Optimizer proposed ending the proposal with a single-choice question: “Would you rather start this week or next?” The hypothesis had strong evidence (83% of proposals got no reply) and promised to raise replies from 17% to 25%. The Guardian vetoed it. The one-line reason: the question presumes a purchase and sets a deadline. That creates a predictable risk of complaints and requests to stop receiving messages—the two “without getting worse” measures on the worksheet. It did not negotiate or suggest a softer question. It vetoed the idea, and the Optimizer moved on.
Common mistake
A target without “without getting worse.” “Increase the response rate” by itself is an invitation: the most irritating message in the world gets a response (“stop sending me this” counts as a reply). This happens because the target seems obvious and the escort seems like a detail. Avoid it by never writing a target without at least two “without getting worse” measures beside it—and define what counts as a response before counting.
Each “without getting worse” measure has a number beside it: the tolerance. Zero seems safest. Almost always, it is the wrong choice, because rare events fluctuate by chance: 1 complaint one month, 6 the next, even though nothing changed.
Rule of thumb: tolerance for a rare event must be greater than its normal fluctuation at the test’s sample size. For something that happens 1% of the time, with 300 on each side, normal fluctuation is about 1 percentage point. So tolerance should be 0.01, not zero.
Applied example: this is exactly what the clinic learned in cycle two. B raised replies from 19.3% to 29.7%, but failed because of 6 complaints versus 1, with zero tolerance. Six versus one out of 300 is not a difference; it is noise. The owner did not reopen the verdict (the written rule had been applied, and the record remained). She changed tolerance to 0.01 for future tests. In the next cycle, even with the new tolerance, the Guardian vetoed “putting the price and payment methods right after the greeting.” Introducing the price before any explanation predictably risks opt-outs and complaints. The hypothesis itself acknowledged the risk. Tolerance is not permission; it is only a gauge for chance variation.
The third thing the Guardian watches is the ceiling: AI cost per cycle and the number of simultaneous tests. When the cycle’s cost exceeds the ceiling, it approves nothing else—not even a perfect hypothesis. If one test is already running and the limit is one, the next approved hypothesis waits in line.
The ceiling is answer 4 on the worksheet, which you will write in track 4. For now, just know it exists and that the Guardian applies it as strictly as the invariants. In the clinic example, the ceiling was R$ 50 per weekly cycle and one test at a time. That is why the third approved hypothesis in the first cycle waited for the previous test to finish.
Applied example: the support manager has 6 agents and 800 tickets per month. If the loop ran three tests at once, each would get 130 tickets per side, and none would reach the sample size needed for proof. With a ceiling of “one test at a time,” each test gets all the tickets and finishes in weeks, not months. The ceiling limits waste, not improvement.
How to write answer 3 on the worksheet
Practice now 0/4 complete
Goal: in about 12 minutes, write down 3 invariants and 2 “without getting worse” measures with tolerances for the process you chose in track 2.
This is safe: you are writing a list, not switching anything on—nothing changes in your business until you approve it. If an invariant seems excessive, keep it for now: loosening a rule later is far cheaper than discovering one was missing mid-test.
Example of an acceptable result (the example clinic’s worksheet, translated): Never discount more than 10%. Never promise a clinical result or timeline. Never contact anyone who asked not to receive messages. Without getting worse: average margin (tolerance 0.01) · complaints (tolerance 0.01) · requests to stop receiving messages (tolerance 0.005).
You have just written answer 3 on the loop worksheet—the list the Guardian will apply consistently to every hypothesis. It is ready when you have used it to veto or approve a real idea.
Summary
Track 03 · Lesson 4
By the end of this lesson, you will have the first line in your loop history: what you tried, what happened (with a number), and what you decided.
Your company has already tested dozens of things—and forgotten almost all of them. Six months from now, someone will have the same idea, test it again, and discover the same result. Without memory, the loop does not learn; it rediscovers. Memory is the only part of the system another company using the same AI cannot copy.
↓ scroll to study
A Memory is the assistant that closes the loop. It reads everything the others wrote, records one line in the history, and stores each hypothesis in the right place. It never deletes an old entry, rewrites it, or summarizes it so much that the “why” is lost.
It is read at the start of the next cycle. The Critic checks what has already been resolved; the Optimizer checks what has already been discarded. Without this, every cycle starts from scratch, with the same AI, ideas, and mistakes.
Applied example: at the agent’s real estate agency, someone tested “send the photos before the price” in 2024, and results got worse. No one recorded it. In 2026, a new agent has the same idea, tests it for two weeks, and reaches the same result. Cost: a month of poorly answered leads to relearn what the company already knew. With Memory, the hypothesis never reaches testing—the Optimizer finds it in “discarded,” with the date and number.
Every hypothesis ends up in one of three drawers. Learned: tested, won, and became the official version—with the number. Discarded: tested and lost, or was vetoed beforehand—with the reason, so it does not return in the same form. Observed, untested: the evidence was too weak to become a hypothesis; it waits for a larger sample.
The third drawer is the most forgotten and the most valuable. It holds the “it seems” ideas—stored without turning them into rules or throwing them away. When the sample grows, the Critic reviews them again.
Before looking at the example folder: these are three text lists, one per drawer. Each entry has a date, cycle, and supporting number. You only need to know how to read the three labels.
Applied example: after the clinic’s four cycles, the drawers looked like this. Learned: “open the proposal by mentioning one specific point from the client’s review.” Replies rose from 19.3% to 29.3% with 300 on each side, guardrails intact; the owner approved it, and it became version 2. Discarded: the forced-choice question (vetoed), the price right after the greeting (vetoed), and the 80-word limit, which won on replies but failed the zero-complaint tolerance. Observed, untested: “soften the approach for the small-client segment”—only 3 events, waiting for more volume.
Common mistake
Twelve proposals became a rule. “Tuesday proposals convert better,” based on 12 proposals, belongs in “observed, untested”—never in “learned.” A pattern in a few cases can look too clear to be luck. Avoid this by putting nothing in “learned” without an A/B test at the required sample size. Guesses based on few cases have their own drawer.
Alongside the drawers, Memory keeps the history: one row per event, in order. Date, cycle, hypothesis, verdict with a number, what was monitored, who decided, what happened to the version, and cost. One row. No narrative.
This history is the version history with an undo button for your process. Every official version has a row explaining where it came from and why. If version 2 starts getting worse, you go back to version 1—and that rollback gets its own row.
Applied example: the support manager writes the first line of the history like this: “2026-09-10 · before the loop · standard reply to ‘I can’t export’ rewritten with a video · 23 tickets in one week, 3 reopened, no A side for comparison · decided by me, no test · old reply deleted—no way back.” It is an honest line about a bad test. That is exactly what keeps the next manager from repeating it.
Test yourself
Version 2 was promoted 3 weeks ago, and the reopening guardrail rose beyond its tolerance. What does Memory record?
Two companies use the same AI chat, model, and type of script. After a year, what sets them apart? Not the AI—they have the same one. It is what each knows about its own customers: what it tried, what worked, what failed, for whom, and at what cost. That is the loop’s memory.
LOOP-R calls this a learning advantage. The competitive barrier stops being “who has more data” and becomes “who learns faster from what they already do.” The ninth assistant, the Meta-agent, looks at this history and only reports cost per cycle and how many hypotheses became versions. It does not act. You will see it in track 4.
Applied example: two real estate agencies in the same neighborhood use the same AI chat to reply to leads. The agent’s agency has a 14-row history: three script versions, five discarded hypotheses with numbers, and one recorded rollback. The other has a script that “someone improved” three times without a record. If the competitor copies the agent’s current script, it gets version 3—but not the five ideas already known not to work. It will test them one by one.
How to write the first line of your history
Practice now 0/4 complete
Goal: in about 10 minutes, write one line about a change your company already made to the process on your worksheet. Include what happened, a number, who decided, and the drawer.
This is safe: you are recording the past, not changing anything—nothing in your business changes until you approve it. If you cannot remember the exact number, write what you remember and label it “approximate”; an honest line with an approximate number is better than no line.
Example of an acceptable result (one line from the example clinic’s history, translated): 2026-07-27 · cycle 4 · “open by mentioning one point from the client’s review” · B won: replies 19.3% → 29.3% with 300 per side · monitored: margin OK, complaints 4 vs. 2 (within tolerance), opt-outs 1 vs. 3 (within tolerance) · approved by the owner · became official version 2 · drawer: learned.
You have just written the first line in your loop history. It is the beginning of an asset no other company using the same AI can copy.
Summary