Track 02 · LOOP-R: The Business That Learns on Its Own
Five lessons, five letters. You’ll leave with your loop’s goal written as a number, ten lines of recorded evidence, three testable hypotheses—and a criterion for deciding, without guesswork, whether the new version won.
Lessons
Leave with answer 1 on the loop worksheet written down: “move [number] from X to Y without making Z worse.”
Leave with the current version of your process written in up to 6 lines—the one that will be compared with the next version.
Leave with 10 rows completed in your process log—the answer to question 2 on the loop worksheet.
Leave with 3 hypotheses in the IF / THEN / BECAUSE format, each tied to a number in your spreadsheet.
Leave knowing how to read an A-versus-B result and give one of three answers: B won, A stays, or not enough data.
Track 02 · Lesson 1 · L for Locate
By the end of this lesson, you’ll be able to write your loop’s numeric goal in one sentence—“move [number] from X to Y without worsening Z”—and that sentence is answer 1 on the loop worksheet.
You already ask AI for things every day: proposals, replies to leads, support messages. But when someone asks “did it improve?”, the answer is an impression. Without a chosen process and a number that exists today, there’s nothing to compare—and the loop can’t even start.
↓ scroll to study
LOOP-R doesn’t improve “the company.” It improves one process at a time. Here, a process is a task that happens often, in the same way, and ends with an observable outcome. Three criteria help you choose: it happens at least dozens of times a month; it has a clear beginning and end; and the person doing it can observe the result.
Marcos, a real estate agent, receives about 80 contacts a month from listing sites. Replying to the first contact is a process: it starts when a lead arrives and ends when the lead replies (or doesn’t) within 24 hours. Listing properties, negotiating fees, and coordinating viewings are processes too—but they belong to other loops. If he tries to improve everything at once, he won’t know what caused what.
Think of a renovation done room by room: you don’t paint the whole house on a Saturday. You choose the living room, finish it, and only then move to the kitchen. The loop works the same way—and the first room should be the one that repeats most, because that’s where evidence accumulates fastest.
The loop’s goal has a fixed form: move [a number] from X to Y without worsening Z. Each part has a job. “A number” says what you measure. “From X” is where you are today. “To Y” is a target you’d be glad to reach. And “without worsening Z” is the part almost everyone forgets—and the part that protects the business.
Dr. Renata owns an aesthetics clinic and sends proposals over WhatsApp after consultations. Out of every 100 proposals, 3 become paid packages. Her goal became: “move conversion from 3% to 5% without worsening average margin or increasing complaints”. Why does the second half matter? Because a number by itself becomes a target—and a system that chases a target learns to cheat. If the request were only “close more proposals,” the version offering a 20% discount would win every time.
In support, the same trap looks different. Cláudio manages customer service at a small software company and asked to “reduce service time.” Without a “without worsening” condition, the winning version would close a ticket in two minutes without solving anything. His condition was: without increasing tickets reopened within 7 days.
Common mistake
Writing a goal without “without worsening.” It looks like a clean goal—“increase reply rate”—but it leaves a door wide open: a message that prompts “stop sending me this” also counts as a reply. Before any test, write at least one number that must not fall. Without it, the loop optimizes against you.
“From X” is a number you have, not the number you wish you had. If nobody tracks it today, the loop’s first task is to count it (that’s what lesson 3 does). There’s a second requirement: the number must respond quickly. There are three speeds. Metrics for money (conversion, ticket size, margin) take weeks to show up. Metrics for signal (customer replied, viewing booked) appear in hours. Metrics for artifact (does the proposal meet the checklist?) appear in minutes.
The difference changes everything because of the sample size that proves it. To prove a version converts 5% versus the current 3%, you need about 1,500 proposals per version. Dr. Renata sends 50 a week—it would take more than a year to complete a single test. But to prove the reply rate moved from 20% to 30%, about 300 per version is enough: 12 weeks at her pace. That’s why the clinic’s loop tests reply rate within 48 hours and only monitors conversion.
Marcos made the same switch. He wanted “more properties sold” (2 a month—impossible to test). The number his loop pursues is leads who reply within 24 hours of the first contact: 22 out of every 100 today. Sales are still what he wants; they’re just not what he measures every week.
Test yourself
Cláudio from support wants “fewer customers canceling their subscription” (currently 4% a month). Which number should his loop test first?
A loop worksheet is the only thing you fill in. There are five questions—like the intake form a clinic asks for at a first appointment: you answer once, and every professional who sees you later reads it. The loop’s assistants read the worksheet the same way at every cycle. Answer 1 is exactly the sentence you built in the previous steps.
You’ll see an excerpt from the worksheet below. It looks like a registration form—and that’s exactly what it is: each row is a question you answer once, and the rest of the loop relies on it. You only edit the answers; the labels on the left stay as they are.
QUESTION 1 — goal (example: Dr. Renata’s clinic) process: WhatsApp sales proposal number: conversion — became a paid package within 15 days today: 3% target: 5% fast metric: reply rate within 48h — now 20%, target 30% without worsening: average margin · complaints · requests to stop
Notice what the worksheet doesn’t promise: that conversion will reach 5%. What it guarantees is something else—no version that hurts margin or increases complaints becomes official, everything is recorded, and the cycle runs the same way every week. If the number rises, great; if it doesn’t, you haven’t lost what you had.
How to write your answer 1
Practice now 0/4 done
Leave with a written sentence in the format “move [number] from X to Y without worsening Z,” saved somewhere you can find again—about 10 min.
Nothing in your business changes until you approve it: this sentence is only the loop’s starting point; no test begins because of it. If you realize tomorrow that you chose the wrong process, delete it and write another.
Copy the template below into your notes or the AI chat you already use, and replace what’s between < > with your situation. The filled-in example is Marcos’s.
Process: <reply to the first contact from a listing portal lead>
Current number: <22% of leads reply within 24h>
Target: <30%>
Without worsening:<viewings booked> · <leads who ask to stop receiving messages>
Sentence: move <the 24-hour reply rate> from <22%> to <30%>
without worsening <viewings booked> or <stop requests>.You’ve just written answer 1 on the loop worksheet: the numeric goal. Everything in the next four lessons builds on this sentence.
Recap
Track 02 · Lesson 2 · O for Operate
By the end of this lesson, you’ll be able to write the current version of your process in up to 6 lines—the script followed today, the same way, in every execution—and recognize, in a real case, what varies from one time to the next.
Everyone executes. The problem is that each execution is done “however it works out”: one proposal is longer, another shorter, one includes a discount, another doesn’t. At the end of the month there are 200 proposals and none can be compared with another. Without a written current version, AI has nothing to improve—only something to repeat.
↓ scroll to study
The second letter in LOOP-R is Operate: do the actual work. It sounds obvious, but this is where many people try to skip ahead. You can’t improve a proposal that was never sent or a support exchange that exists only on paper. Every real execution is a chance to learn—and every execution that never happens is a missed chance.
Marcos, the real estate agent, replied to portal leads from memory. On Monday, with time to spare, he wrote three paragraphs and sent a property photo. On Thursday, in traffic, he sent “hi, interested?”. He executed 80 times that month—but they were 80 executions of 80 different things. When he asked AI “how can I improve my reply?”, AI had no way to answer: which one?
It’s like the difference between a recipe and “cooking like Grandma.” Grandma gets it right, but nobody can repeat or improve it, because her hand weighs things differently every day. A written recipe may be wrong—and that’s exactly why it can be corrected.
Before the first cycle, the process gets a current version: a short written script that describes how the work is done today. It doesn’t need to be good. It needs to be the same in every execution, so 200 proposals become 200 comparable cases—not 200 opinions.
Dr. Renata’s clinic version 1 had five blocks: greeting by name and thanking the client for the visit; a summary of what was observed during the assessment, in two or three sentences; an explanation of the recommended package; price and payment options; a friendly close, “available for any questions.” It fits in six lines. That script generated the first 200 logged proposals—and revealed the first finding: 166 out of 200 got no reply within 48 hours.
Before
Every proposal written from memory, however it came out that day. At month’s end: 200 messages, none alike, none comparable.
After
Written version 1 script in 6 lines, followed for every proposal. After 4 weeks: 200 proposals with the same structure, all logged.
Result: 200 comparable cases in 4 weeks —enough for the first finding (166 without a reply). Without the written script, the same month produces zero.
Following the current version to the letter may seem rigid. It’s the opposite: it’s what lets you make changes safely later. If two executions differ in five things at once and one does better, you don’t know which of the five made the difference. If they differ in one thing, you know. The fixed script is what lets the next lesson change one thing at a time.
At Cláudio’s support team, two agents replied to the same type of ticket in opposite ways: one sent step-by-step instructions with images, the other asked for remote access right away. The month’s numbers said 61% of tickets were resolved on first contact. But 61% of what? A mixture. When both agents started following the same version 1 (steps first, remote access only if that failed), the number fell to 58%—and for the first time it meant something.
This doesn’t mean nobody can improvise. It means improvisation doesn’t count. Anyone who goes off script writes “off script” on that execution’s row—and that row stays out of the comparison.
Test yourself
Marcos followed version 1 for 70 leads and improvised on 10. When comparing with version 2, what should he do with the 10?
In LOOP-R, the one who executes is a single-purpose assistant called Executor. It gets the current version and the case data, produces the artifact (proposal, reply to a lead, ticket response), and logs one row. That’s all. It doesn’t judge whether the result is good, suggest changes, or “improve” anything on its own.
For Marcos, the Executor is the assistant that receives version 1 of the script plus the lead’s details (name, property clicked, neighborhood) and returns the first reply ready to paste into WhatsApp. If the script says “one question at the end,” it includes one question at the end—even if two might seem better in that case. The urge to “improve it on the fly” is exactly what it doesn’t do, because an on-the-fly improvement is invisible: nobody recorded it, nobody compared it.
It’s the same separation as on an assembly line: the person who assembles isn’t the inspector. Not because of distrust, but because of method. The person who did the work never sees it the way the person checking it does.
An unrecorded, on-the-spot improvement isn’t an improvement. It’s luck nobody will be able to repeat.
Practice now 0/3 done
Read the case, write the current version in up to 6 lines, and compare it with the answer key—about 10 min.
Nothing in your business changes until you approve it: writing the current version only describes what already happens. If the description is wrong, nobody acts on it—you correct it, and that’s it.
Marcos, a real estate agent for 12 years, describes how he replies to a listing-site lead: “It depends. If the property is good, I send a photo and price right away. If the lead seems curious, I ask what they’re looking for first. I usually greet them by name because the portal sends it. Sometimes I offer a viewing at the end; sometimes I forget. On weekends, I reply more briefly.” He sent 80 replies last month.
Always: greets them by name. Sometimes: sends a photo and price right away or asks first; offers a viewing or forgets; shortens the reply on weekends. Each “sometimes” is a variation that prevents comparison.
One possible version 1: 1) Greet them by name. 2) Mention the property they clicked. 3) Ask what they’re looking for (bedrooms, neighborhood, timeline) before sending a price. 4) Say you have photos and a price range, without sending them yet. 5) Close by saying you’re available for questions. 6) Use the same text every day of the week.
If your version chose “send the photo and price right away” instead of asking first, that’s equally valid. Version 1 doesn’t need to be the best—it needs to be one consistent version. Finding out which one is better is the work of lessons 4 and 5, not this step.
You’ve just written a current version: a script that turns “it depends” into one fixed decision at a time. Every new version will be compared against it.
Recap
Track 02 · Lesson 3 · O for Observe
By the end of this lesson, you’ll be able to fill in 10 rows in your process log—one row per run, with the column that answers your metric—and identify where each row comes from. This spreadsheet is answer 2 on the loop worksheet.
Your company runs this process hundreds of times a month and keeps almost none of the data. The outcome of each proposal is in WhatsApp, in the sender’s memory, or in “I think it closed.” Everything the loop can learn depends on this step—and it’s the step no one thinks matters until it’s missing.
↓ scroll to study
The third letter is Observe. This is where the loop’s intelligence begins—and it’s an unglamorous step: every run becomes a row, with its outcome recorded. No “I think shorter proposals work better.” The question is: how many were short, and how many of those got a reply?
At Dr. Renata’s clinic, the first 200 proposals using version 1 became 200 rows. The “replied within 48 hours” column added up to 34: a 17% response rate. Her worksheet said 20%—the number she “thought” was right. A 3-point difference isn’t huge, but it’s the first time the number left her head and made it onto paper. From then on, everything the loop proposes starts from 17%, not 20%.
Think of how a clinic handles blood pressure: it doesn’t ask “do you feel pressure?”; it measures, records, and compares with the last reading. Evidence is the recorded measurement. Opinion is “I feel.”
A process log is an ordinary spreadsheet—the one you already use will do. What changes is the discipline: one row per run, always with the same columns. It is answer 2 on the loop worksheet: “where the evidence begins.”
Below you’ll see the spreadsheet header. It looks like an accountant’s table, but works like the reception desk’s logbook: each column is a question answered once per row. The first three say what and when; the middle ones say which version; the last ones say what happened. You only fill it in; the column names never change.
id date version variant replied outcome complained notes L0041 2026-09-02 v1 A yes no no portal lead, 2-bedroom apartment L0042 2026-09-02 v1 A no no no L0043 2026-09-03 v1 A yes yes no visit scheduled for Saturday
Marcos uses exactly these columns. “id” is any code he makes up (L0041, L0042…); “version” is the script he followed (v1 for now); “variant” stays “A” until there’s a test; “replied” is his quick metric; “outcome” is his money metric (did the visit close or not); “complained” is the “no harm” metric. “Notes” is free-form—and where he writes “off script” when he improvises.
How to fill in a row (30 seconds each)
The loop assistants are the easy part. The hard part is that the evidence lives in different places: the clinic’s WhatsApp, the portal inbox, the ticketing system, a partner’s spreadsheet, the salesperson’s memory. Where the evidence begins is question 2 on the worksheet precisely because no one can answer it alone. No assistant will fetch data from your WhatsApp; someone—you, the receptionist, or the Executor when they write the artifact—has to bring each outcome to the row.
And every column needs a definition that fits yes or no. In Cláudio’s support team, “resolved” seemed obvious—until agents closed tickets without the customer agreeing. The definition became: “resolved = the customer did not reopen it within 7 days.” At the clinic, “replied” became “replied within 48 hours with something other than a refusal”: does “no, thank you” count as a reply? By this definition, no. Without that sentence, the version that gets more “no, thank you” replies would look like the winner.
Test yourself
Marcos defined the “replied” column as “the lead sent any message back.” What is the risk of this definition?
Ten rows teach you how to fill it in. They don’t support conclusions. The loop only accepts a pattern with at least 30 rows of the same type —and treats it as strong starting at 90. Below that, what looks like a pattern is almost always chance, so the spreadsheet saves the observation in a separate “observed, not tested” list to revisit when there are more rows.
With 12 leads recorded, Marcos noticed that “Tuesday leads reply more.” Three of the four Tuesday leads replied; only two of the other eight did. It looks like a pattern. With 12 rows, it’s nothing: one lead changing days would make the “discovery” disappear. The note stayed on the observations list. Eight weeks later, with 160 rows, Tuesday was no different from any other day.
Common mistake
Twelve proposals became a rule. It happens because the first rows are exciting: for the first time there’s a number, and it seems to say something. How to avoid it: make no claims with fewer than 30 rows of the same type; below that, the observation goes in “observed, not tested” and waits. The loop counts for you; your job is not to jump ahead.
Practice now 0/4 done
Leave with a spreadsheet (the one you already use) with the header below and 10 rows from your last 10 process runs—about 12 minutes.
Nothing changes in your business until you approve it: the spreadsheet only describes what already happened. If you get a row wrong, delete it and try again; if a column won’t fit yes/no, that means its definition needs another sentence—not that you did it wrong.
Copy the header below into the first row of your spreadsheet. Each word becomes a column; the “replied” column gets the name of your quick metric, and “outcome” gets the name of your money metric. Then add one row per run, starting with the most recent.
id date version variant <your quick metric: replied> <your money metric: outcome> <your no-harm metric: complained> notes
You’ve just created the source of evidence for your loop: 10 rows checked at the source, with definitions that fit yes or no. From now on, each new run is just one more row.
Summary
Track 02 · Lesson 4 · P for Propose
By the end of this lesson, you’ll be able to write 3 hypotheses in the IF / THEN / BECAUSE format, each tied to a number in your spreadsheet, and say which one has the strongest evidence.
You already have plenty of ideas to improve the process—the problem is that they’re ideas, not hypotheses. “Make the proposal shorter” can’t be tested or disproved. A hypothesis says what will change, by how much, and why—and it’s the only thing the loop will take into a test.
↓ scroll to study
The fourth letter is Propose. But proposing doesn’t start with creativity; it starts with a cool-headed reading of what the spreadsheet showed. In the loop, that reading has an owner: the Critic, a single-purpose assistant that examines the evidence and says what worked, what failed, and how strongly the evidence supports each claim. It doesn’t suggest a solution. It only diagnoses.
Looking at the clinic’s 200 proposals, the Critic noted: 166 of 200 got no reply within 48 hours (83%); the cost and time per proposal were steady (about five and a half minutes each); and version 1 ended with “let me know if you have any questions”—without asking the client to do anything. Three statements, three numbers. No suggestions. Only then does someone propose an idea.
It’s the mechanic’s order of operations: diagnosis first (“the noise is coming from the front suspension”), then the estimate. Anyone who proposes a fix before diagnosing may replace a good part.
A LOOP-R hypothesis has three required parts. IF: what changes in the script, concretely. THEN: which number changes, and from what to what. BECAUSE: what evidence in the spreadsheet supports the bet. And a fourth practical part: the CHANGE —the exact edit to the current version’s text, so the Executor can follow it.
Here is the loop’s second real hypothesis from the clinic, as written by the Optimizer: IF we add a rule requiring no more than 80 words in total, THEN the response rate should rise from 17% (200 proposals) to at least 22%, BECAUSE 83% of proposals get no reply and the 5-part structure makes for a long WhatsApp message; shortening it reduces the reading cost before the price. CHANGE: under “Rules,” add the line “maximum 80 words in the full message.”
Before
“AI, make the proposal shorter and better.” It doesn’t say how short, what better means, or why. If it succeeds or fails, no one will know.
After
“IF no more than 80 words, THEN response rises from 17% to ≥22%, BECAUSE 166 of 200 don’t reply and the message is long.” It can be tested, and it can lose.
The low version can be disproved by 300 spreadsheet rows. The one above can’t even be wrong—and that’s why it never improves anything.
Not every hypothesis starts out equal. The Critic labels each claim: strong when backed by 90 or more rows of the same type; moderate starting at 30; weak below that. Strength determines the queue: strong hypotheses go to testing first; weak ones stay on the “observed, not tested” list until the spreadsheet grows. You can only test one thing at a time—the evidence strength tells you which one.
Marcos had two ideas. The first—“leads who ask for the price in their first message book fewer visits”—came from 25 leads, so its evidence was weak; it went on the waiting list. The second—“the current reply doesn’t end with a question”—came from the 130 rows recorded so far, 101 with no reply within 24 hours; its evidence was strong. The second went to testing: IF the first reply ends with one clear question (“Would you rather visit this week or next?”), THEN the 24-hour response rate should rise from 22% to at least 28%, BECAUSE 101 of 130 leads don’t reply and the current text doesn’t ask for a response.
Test yourself
Dr. Renata has a strong hypothesis (166 of 200 got no reply) and an idea she loves, based on 9 customers who praised the “more personal” tone. What does the loop do with the second idea?
Writing the hypothesis ends this letter, not the loop. It still has to pass through two gates. The first is the Guardian: an assistant that checks whether a change touches something that must never change on its own or puts a “no harm” metric at risk. If it does, the Guardian vetoes it—and that veto is final. The second gate is next lesson’s test. Only what passes both becomes the official version.
In Cláudio’s support team, 37% of the month’s tickets involved the same issue: the payment screen wouldn’t load. The hypothesis was clear: IF the first reply to this kind of ticket includes the three steps that solve 80% of cases, THEN first-contact resolution should rise from 58% to at least 65%, BECAUSE 112 of 300 tickets are about this issue and the standard reply only asks for more information. The Guardian approved it. At the clinic, though, the cycle’s first hypothesis—ending the proposal with one closing question such as “Would you rather start this week or next?”—was vetoed before any test. The same idea Marcos was allowed to test stopped there: Dr. Renata’s worksheet had zero tolerance for complaints, and a forced choice puts pressure on exactly that metric. The Guardian reads each loop’s worksheet, not a general rule.
A hypothesis that seems great but gets vetoed isn’t a failure. It’s the loop showing that the worksheet was taken seriously.
An idea that can’t be disproved can’t improve anything.
Practice now 0/4 done
Leave with 3 hypotheses in IF / THEN / BECAUSE / CHANGE format, each labeled with its evidence strength and based on your lesson 3 spreadsheet—about 10 minutes.
Nothing changes in your business until you approve it: a hypothesis is just a sentence in a notes app. None changes the script, and none goes to testing without the Guardian and your approval. If the AI suggests something that touches a no-change rule, cross it out and ask for another option.
The text below is a prompt —a ready-to-paste instruction for the AI chat you already use. It looks long because it does the Optimizer’s job: it provides the format, numbers, and rules. Replace only what’s inside < > with your data; leave the rest as is.
You are the Optimizer in an improvement loop. Only propose hypotheses; do not execute or evaluate them. Process: <reply to the first message from a portal lead> Current version (script, 6 lines): <paste the 6 lines of your version 1 here> Evidence (my latest recorded rows): - rows recorded: <10> - quick metric: <replied within 24h> = <2 of 10> - money metric: <visit booked> = <1 of 10> - no-harm metric: <asked us to stop> = <0 of 10> - what I observed: <7 of the 8 who didn’t reply got a text with no question at the end> Rules: 1. Write exactly 3 hypotheses, each with IF / THEN / BECAUSE / CHANGE. 2. THEN must name the quick metric and its current and target values. 3. BECAUSE can only use the evidence above. If there are fewer than 30 rows, label STRENGTH: weak. 4. CHANGE is the exact text to edit in the current version, in one line. 5. No hypothesis may touch: <discounts over 10% · promising a deadline>.
You’ve just written three hypotheses that can be disproved—the difference between an idea and an experiment. When the spreadsheet passes 30 rows, their strength changes and the first one can go to testing.
Summary
Track 02 · Lesson 5 · R for Reinforce
By the end of this lesson, you’ll be able to review an A-versus-B test result and choose one of three answers—B won, A stays, or not enough data—based on the rule written on the worksheet, not on your impression.
This is where most loops die: someone looks at twenty replies using the new version, thinks “it’s better,” and changes the script. Three weeks later, no one knows whether it improved, and the old version is gone. Reinforce is the letter that turns “seems” into “proved”—and keeps a way back.
↓ scroll to study
The fifth letter is Reinforce: test the hypothesis, evaluate the result, and only then incorporate—or discard it. The important words are only then. Between the hypothesis and the official version there is a test, and the test has an outcome that doesn’t depend on anyone’s taste.
Marcos read the first twenty replies written with the new version (the one that ends with a question) and liked them: more direct, friendlier. He wanted to change the script that day. What he had was an impression about twenty messages—not a single outcome row. Those twenty replies could be lovely and still get fewer responses than the old ones. He wouldn’t know.
It’s the difference between a medicine that “seems to have helped” and one tested against another. The first is a story; the second is a result. The loop only accepts the second.
Common mistake
Making a version official because it seems better. It happens because the new version always looks better to the person who proposed it. How to avoid it: the official version only changes with a written test verdict—and the verdict must be one of the three on the worksheet. If you can’t point to the spreadsheet row supporting the change, it doesn’t happen.
A LOOP-R test is simple to describe: half of the runs follow the current version (A), the other half follows the candidate (B), during the same weeks, chosen at random. The spreadsheet’s “variant” column, which was set to A, now alternates. The Experimenteris a single-purpose assistant: it calculates the sample size needed for proof, estimates how many weeks the test will take at your current pace, and writes the stopping rule before the test begins.
Why at the same time, instead of “this month with B, last month with A”? Because months change: at Dr. Renata’s clinic, December has holidays and January has vacations, and the response rate shifts by 8 points without anyone changing anything. If B runs in January and A ran in December, the difference may come from the calendar, not the script. With A and B mixed into the same weeks, the calendar affects both equally.
Test yourself
Marcos used version 1 in August (22% response) and version 2 in September (27%). He says B won. What’s the problem?
When the test ends—or each week while it runs—the Evaluator reads the spreadsheet and gives one of three answers. B won: the difference is bigger than chance can explain, the required sample size has been reached, and no “no harm” metric fell beyond its allowed margin. A stays: B didn’t win, or it won on the main metric but lost on a guardrail. Not enough data: there aren’t enough rows to say anything yet—keep testing. There is no fourth answer. There is no “it’s winning.”
The third answer is the most common, and it has to be acceptable. To prove that the response rate rose from 20% to 30%, you need about 300 runs in each version; at Dr. Renata’s pace of 50 a week, that takes 12 weeks. In Cláudio’s support team, after three weeks of testing, B had resolved 41 tickets on first contact versus 38 for A—66% versus 61%. He wanted to stop: “B is winning.” The Evaluator said there wasn’t enough data because the worksheet required 220 tickets per version, and there were only 62 each. Stopping early because it’s winning is the most common way to make an imaginary difference official.
“Not enough data” is the loop’s most common answer—and the day it stops being acceptable, the loop stops working.
B only wins if it raises the number and none of the “no harm” metrics exceeds the tolerance written on the worksheet. Both conditions must hold. If the number rises 40% and the margin drops by one point with zero tolerance, the answer is A stays—no exceptions, no “but look how much it went up.” The Evaluator applies the rule as written; if the rule is poorly written, you fix it for future tests.
In support, Cláudio’s version B finished testing with 220 runs per variant: first-contact resolution rose from 58% to 66%. But tickets reopened within 7 days rose from 4% to 9%, with a 2-point tolerance. A stays. B resolved tickets “faster” by closing ones that came back—the exact trick the “no harm” metric was designed to catch.
There’s one thing to watch with the tolerance. For frequent metrics (margin, reopenings), zero tolerance works. For rare events —one complaint in every 100, one “stop messaging me” in every 200—zero tolerance can’t separate an effect from chance: in 300 runs, 1 or 4 complaints may be the same noise. The loop’s rule is that the tolerance for a rare event should be at least as large as the expected noise at the planned sample size—about 1 point for 1% events in 300 rows. Writing zero there isn’t rigor; it can reject good hypotheses by chance.
When the verdict is B won, the decision still goes through you: at most once a week, you get a five-line card—what changed, how much it improved, what was monitored, what it cost, and three buttons: approve, reject, or wait for more data. If you don’t respond by the deadline, the default is not to make it official. Once approved, the candidate becomes the current version, and the old one is saved in the version history with rollback.
The rollback button isn’t decoration. During the first cycle after a change, the loop checks whether the test metric or a “no harm” metric fell below the old version’s result. If it did, the old version becomes official again, and the reason is recorded. This is the LOOP-R’s one unambiguous guarantee: the system won’t choose a worse version—and if one slips through, there’s a way back.
Before · clinic version 1
Same opening for every client: greeting and thanks for visiting. 48-hour response rate: 19.3% (300 proposals during the test period).
After · version 2, promoted in cycle 4
The first sentence mentions a specific detail from that client’s assessment. Response rate: 29.3% (300 proposals). Margin, complaints, and requests to stop all stayed within tolerance.
Net result: +10 points in replies, recorded and with a rollback button. Paid-package conversion was 3.7% versus 3.0%—a difference 300 proposals can’t prove. The loop delivered what it promised; what it didn’t promise is still being tested.
Practice now 0/3 done
Read the result below, write which of the three answers applies and what to do with the rule, then compare with the answer key—about 10 minutes.
Nothing changes in your business until you approve it: you’re deciding on another company’s loop, on paper. If you get the verdict wrong, the answer key explains why—and this is exactly the mistake the lesson wants you to make here, not in your own process.
Dr. Renata’s clinic, second cycle. Hypothesis tested: a proposal with no more than 80 words. A-versus-B test over the same 12 weeks, 300 proposals per version (the worksheet required 166). 48-hour response rate: A 19.3% · B 29.7% —a difference far greater than chance can explain. Guardrails on the worksheet: average margin, 1-point tolerance—A 31.3% · B 31.1%; “don’t message me again” requests, 0.5-point tolerance—A 0 · B 1 in 300; complaints, zero tolerance—A 1 in 300 · B 6 in 300.
1. A stays. The required sample size was reached (300 ≥ 166), and B won on the main metric by a wide margin—but complaints exceeded the written tolerance: 6 versus 1, with zero tolerance. The worksheet says B only wins if no “no harm” metric exceeds its tolerance. The complaint row decides it. This was the loop’s actual verdict: A stays, and the 80-word hypothesis was recorded as tested and rejected.
2. Fix the rule, not the verdict. A complaint is a rare event (about 1 in 100). With 300 proposals, 6 versus 1 isn’t a meaningful difference—it’s the size of the noise. Zero tolerance was poorly written: future tests will use 1 point (the expected noise for a 1% event in 300 rows). This test not get re-judged under the new rule: the rule in force when the test ran applies, and changing the yardstick after seeing the result is the most elegant way to “promote it because it seems better.” The hypothesis can be proposed in a future cycle, with the corrected tolerance, and tested again.
3. In one sentence: the loop doesn’t promote a metric that rises while violating the worksheet—and when the worksheet is wrong, you fix it for the next test, never the result of this one. Two things were saved: the rejected hypothesis (so no one tests it again by accident) and what was learned about tolerance (for the loop worksheet and this course).
You’ve just given a rule-based verdict against your impression—with a 10-point difference shouting the opposite. That’s the skill that separates a loop that learns from one that only changes scripts.
Summary