LOOP-R · Track 04
Learn to say how many runs prove an improvement, write a yes-or-no checklist, and decide on a real card—approve, reject, or wait—with rollback at hand.
Lessons
Learn to spot, in one line, how any business target can be inflated without improving anything—and what to monitor to prevent it.
Learn to say how many runs a test needs and how many weeks it will take at your business’s pace—before it starts.
Leave with six yes-or-no questions about the text your AI produces, answered the same way by two different people.
Learn to decide on a five-line card and complete answers 4 and 5 on the loop worksheet: the ceiling and who approves.
Learn to distinguish the loop’s four real guarantees from the four promises it does not make—and recognize a false promise.
Track 04 · Lesson 1
By the end of this lesson, you will be able to take a business target and write, in one line, what AI might do to inflate the number without improving anything—and which two metrics you will monitor to prevent it.
Any target given alone to a learning system becomes something to game. “Reply faster” turns into closing tickets without resolving them; “sell more” turns into discounting until the margin disappears. This is not an AI flaw: it is what any employee might do when judged by a single number. Before measuring for real, you need to know how the measure can be gamed.
↓ scroll to study
There is an old rule in every business: when a measure becomes a target, it stops measuring. Not because people are dishonest, but because they are smart. If you pay by the number of support interactions, those interactions get shorter. A learning system does the same thing, only faster and without guilt.
In LOOP-R, the target metric is the number you asked the loop to increase. If it is the only number it watches, it will find the shortest path to it—and that path rarely means serving people better.
The owner of a beauty clinic asks, “increase the proposal response rate.” Three cycles later, it has risen from 20% to 34%. She opens the conversations and finds out why: the new version sends a second message the next day, and many of the “replies” say, “please stop sending me this.” They counted as responses. The needle went up; the clinic did not.
You do not need to imagine every possible way to cheat. The same patterns repeat. Ask for maximum clicks and the loop learns clickbait. Ask it to minimize handling time and it learns to close tickets quickly without resolving them. Ask for more sales and it learns to discount until the margin disappears. Ask for more replies and it learns to provoke people.
Notice the pattern: in every case, the requested number really does go up. The trick is not in the number; it is in what the number leaves out.
A support manager at a small software company set a target of “average ticket time, from 18 to 10 minutes.” In two weeks it reached 9. But tickets closed in 2 minutes reopened the next day, and satisfaction fell from 4.3 to 3.6. None of this appeared on her dashboard, which showed only time.
Before—a bare target
“Average ticket time: from 18 to 10 minutes.” A ticket closed in 2 minutes and reopened tomorrow counts in favor of the target.
After—a target with guardrails
“From 18 to 10 minutes without worsening 7-day reopenings (currently 9%) or satisfaction (currently 4.3).” The same ticket now counts against it.
Takeaway: the same target and dashboard—but “close without resolving” is no longer a shortcut; it automatically fails.
In LOOP-R, a target never stands alone. Every target metric is paired with at least one number that must not fall—a guardrail. The complete statement is always “maximize X without worsening Y, Z, and W.” This is not decoration; it is the part of the target that prevents gaming.
Typical guardrails vary by area. Sales: average margin, complaints, unsubscribes, average discount. Support: satisfaction, 7-day reopenings, customer effort. Content: time on page, unsubscribes, reports. You do not need to watch them all—only those most likely to be affected by the shortcut.
A real estate agent who replies to portal leads writes: “maximize booked visits per lead without worsening the no-show rate or average discount offered.” If the new version books more visits by promising discounts in conversation, the average discount rises—and the version fails, even with a full calendar.
Two loop assistants use this statement. The Guardian vetoes any hypothesis that proposes changing a guardrail. The Evaluator rejects any version that raises X but worsens a guardrail beyond tolerance—even if X rose a lot. If replies rise 40% and margin falls 1% with zero tolerance, the answer is “A stays.” No exceptions.
Test yourself
Version B raised the response rate by 40%. Average margin fell 1%, and the tolerance for the “margin” guardrail is zero. What does the Evaluator answer?
There is a quieter trick than the three above: it happens in the definition of the number. What exactly counts as a “reply”? If any message back counts, “stop sending me this” is a reply. If “resolved” means any closed ticket, a forcibly closed ticket counts as resolved.
That is why the metric definition belongs on the loop worksheet; it is not a detail. “Reply” must exclude negative replies; “resolved” must exclude tickets reopened within 7 days. The Observer—the assistant that counts—checks that the log matches this definition before adding anything up.
At the clinic, the “replied” column in the log gained a neighbor: “positive reply (yes/no).” At the software company, “closed on” gained “reopened within 7 days (yes/no).” Two extra columns, and both tricks from the start of this lesson disappeared from the count.
A target without guardrails is not a target. It is an invitation.
Practice now 0/4 complete
For each target below, write the most likely trick and two guardrails. About 10 minutes. This is analysis, not execution: running a loop until it cheats would take weeks; the completed case shows in minutes what you need to recognize.
Nothing in your business changes until you approve it: you are writing on paper, not on the loop worksheet. If your answer differs from the key, compare the reasoning—each target can have more than one possible trick and more than one valid guardrail.
Target 1—Renata, beauty clinic owner. “I want to double the number of proposals that get a reply on WhatsApp.”
Target 2—Marcos, real estate agent. “I want more visits booked per lead from the portal.”
Target 3—Paula, support manager. “I want to reduce the average time per ticket.”
Target 1 (clinic). Trick: keep messaging (second and third messages) or provoke, because “stop sending me this” counts as a reply. Guardrails: complaints and unsubscribes (requests to stop receiving messages). Extra: define “reply” as a positive reply in the log.
Target 2 (real estate agent). Trick: book at any cost—promise a discount or say “just come take a look,” filling the calendar with visits that never happen. Guardrails: no-show rate and average discount promised. Acceptable alternative: proposals closed per visit.
Target 3 (support). Trick: close tickets quickly without resolving them. Guardrails: 7-day reopenings and satisfaction. Extra: “resolved” must exclude reopened tickets.
If you chose a different guardrail, ask: “Would the trick I described affect this number?” If yes, it is valid.
You have just spotted the trick before it happens—in three different businesses, using the same question.
Summary
Track 04 · Lesson 2
By the end of this lesson, you will be able to look at a target like “from 20% to 30%” and say how many runs the test needs per version and how many weeks it will take at your business’s pace—before you start.
Most loops die here: someone compares 12 new proposals with 12 old ones, sees “double,” and changes the team’s rule. Three weeks later, the number is back to normal and no one knows why. Luck with a small sample looks like a pattern. There is a calculation that separates the two—and it fits in a five-row table.
↓ scroll to study
Flip a coin twelve times. Getting eight or more heads happens in nearly one out of five tries—with a fair coin, no trick. Now replace “heads” with “customer who booked” and “tails” with “customer who did not reply”: with twelve customers, a new message “beating” the old one is what chance predicts, not proof that it is better.
That is why “I tested it and it worked” is rarely true with a small sample. The person is not wrong about what they saw; the same thing can happen when there is no difference at all.
A real estate agent changes the first message sent to portal leads. Of the next 12 leads, 4 book a visit. With the old message, 2 of the previous 12 did. “It doubled,” he says at Monday’s meeting, and everyone switches to the new message. Three weeks later, the rate is back where it always was—and no one knows whether the new message is better, worse, or the same.
The sample size needed for proof is a standard calculation. The Experimenter—the assistant that designs the test—works it out and shows you before the test starts. You do not need to know the formula; you need to know how to read the table.
| from → to | per version | total | at 50/week | at 200/week |
|---|---|---|---|---|
| 3% → 5% | ~1.500 | ~3.000 | 60 semanas | 15 semanas |
| 3% → 6% | ~700 | ~1.400 | 28 semanas | 7 semanas |
| 20% → 30% | ~300 | ~600 | 12 semanas | 3 semanas |
| 20% → 35% | ~140 | ~280 | 6 semanas | 1,5 semana |
| 50% → 65% | ~170 | ~340 | 7 semanas | 2 semanas |
Notice two things. First: the lower the starting number, the more runs it requires—3% is almost impossible to prove in a small business. Second: the bigger the lift you expect, the fewer runs you need—but the lift must be real, not wishful thinking.
The clinic owner sends 50 proposals a week and wants to raise sales conversion from 3% to 5%. The table calls for 1,500 proposals per version: 60 weeks for one test. More than a year to answer one question. The loop is not broken; that question cannot be answered at that pace.
The clinic’s answer is not to give up measuring; it is to measure a faster number that still relates to revenue. LOOP-R separates metrics into three types. Revenue: conversion, order value, margin—the customer determines these weeks later; proving them takes months. Signal: reply rate, bookings, useful clicks—the customer responds in hours or days; proof takes weeks. Deliverable: does the proposal meet the checklist? Is there an error? Is the tone right? You can answer by reading the text, in minutes.
The loop’s rule: test the fastest metric that still has a plausible link to revenue. Revenue does not disappear; it becomes a guardrail (“without worsening conversion”) until there is enough volume to test it directly.
At the software company, the support manager has all three: revenue is contract renewal (months); the signal is “resolved on first contact” (days); the deliverable is “does the reply meet the checklist?” (minutes). Her loop runs weekly on the deliverable and signal, while monitoring renewals.
Before—measuring revenue
Clinic, 50 proposals/week. Target: sales conversion from 3% to 5%. Required sample: 1,500 per version. 60 semanas until the first reliable result.
After—measuring the signal, monitoring revenue
Same clinic, same pace. Target: reply rate from 20% to 30%, without worsening conversion. Required sample: 300 per version. 12 semanas.
Takeaway: 48 fewer weeks to a reliable first answer—revenue stays in view, just no longer as the target.
The Evaluator, the assistant that judges the test, can give only three answers: B won by a margin, A stays, ou insufficient sample—continue the test. If the number of runs has not reached the required sample size, the answer is the third one. Always. There is no “but it is winning.”
Stopping early because it is winning is the most common way to turn noise into a rule. The previous step’s table says when a test ends—not impatience to see the result. The only allowed early stop is a falling guardrail: the test stops because it is causing harm, not because it is going well.
In the real beauty-clinic case followed by this course, the first cycle nearly started on the wrong foot: the Optimizer expected a small effect, and the Experimenter warned that the test would take 40 weeks. The owner resized it around the target on the worksheet (20% to 30%): 300 proposals per version. The test fit the calendar before the first message was sent.
Test yourself
Week 5 of a 12-week test for the real estate agent: message B has 31% booked visits versus 22% for A, with 140 leads per side. The required sample is 300 per version. What should you do?
After the loop counts, the Critic—the assistant that says what worked and failed—will want to segment by customer type, channel, or day of the week. Every segment obeys the same small-sample rule. “Tuesdays convert better” based on 12 proposals is not a pattern. The Critic can call something a pattern only with at least 30 runs in that segment.
That is why every claim it makes has a strength label. Strong: 90 or more runs—becomes a priority hypothesis. Moderate: 30 to 89—becomes a hypothesis. Weak: fewer than 30—goes into memory as “observed, untested” and is reviewed when the sample grows. Nothing is lost; nothing becomes a rule too early.
At the clinic, the Critic noticed that facial-treatment customers replied more often than body-treatment customers: 14 on one side and 11 on the other. Label: weak. The observation was stored; three months later, with 60 on each side, it became a moderate hypothesis to test. If it had become a rule in the first month, it would have been a rule based on 25 people.
Twelve proposals say nothing. Three hundred start to speak.
Pratique agora 0/4 feito
For each business below, choose a realistic metric, state the number of runs per version (use the table in step 02), and calculate how many weeks it takes at the stated pace. About 10 minutes. This is analysis: a real test takes weeks, but here you do the Experimenter’s calculation in minutes.
Nothing in your business changes until you approve it: you are only doing calculations on paper. If your number differs from the key, check that you read the right table row—the most common mistake is using “total” instead of “per version.”
Case 1—Renata, beauty clinic. 50 proposals a week. Today, 20% reply and 3% buy. She wants to “sell more.”
Case 2—Marcos, real estate agent. 200 portal leads per week. Today, 20% book a visit; he thinks the new message will raise it to 35%.
Case 3—Paula, support. 400 tickets per week. Today, 50% are resolved on first contact; the target is 65%.
Case 1. “Sell more” is a revenue metric (3% → 5%): 1,500 per version, 60 weeks at 50/week. Not feasible. A realistic metric is reply rate, 20% → 30%, with conversion as a guardrail: 300 per version, 600 total, 12 weeks.
Case 2. Booking is a signal, and he has enough volume: 20% → 35% requires 140 per version, 280 total. At 200/week: a week and a half. If the real effect is smaller than he expects (only 30%), the requirement rises to 300 per version and 3 weeks—the Experimenter designs for the worksheet target, not for hope.
Case 3. Resolved on first contact is a signal; 50% → 65% requires 170 per version, 340 total. At 400/week: less than a week. Guardrails: 7-day reopenings and satisfaction, or “resolved” becomes “closed.”
You have just done the calculation that separates luck from a pattern—for three different business paces.
Summary
Track 04 · Lesson 3
By the end of this lesson, you will be able to write a checklist of 6 yes-or-no questions about the text your AI produces—questions that two different people answer the same way, without debate.
When a result depends on a customer replying, a test takes weeks. But much of what you want to know is in the text itself: is it under 80 words? Does it mention something the customer said? Does it promise a timeline? You can check that in minutes. But “give it a score from 1 to 10” does not work—two people, or the same AI two weeks apart, will not give the same score.
↓ scroll to study
Ask two people to score a sales proposal. One gives it 7, the other 8, and both are right—because “7” means nothing that can be checked. Ask an AI for a score and repeat next week: it changes too. Text evaluators, human or not, have known biases: they prefer longer text, prefer the second text they read, and change their minds from week to week.
The answer is not to find a better judge. It is to change the question. “Is this proposal good?” cannot be checked. “Does this proposal contain exactly one closing question?” can: yes or no, and anyone reading the text will answer the same way. A checklist is just that: questions that leave no room for personal taste.
The clinic owner spent months asking her assistant to “improve this proposal” and getting different versions without knowing whether they were better. She switched to “does the proposal mention something specific the client said? Is it at most 80 words? Does it promise a timeline?” Now she and the receptionist give the same answers about the same proposal—for the first time.
What follows looks like a car inspection form, and that is exactly what it is: each row is a question answered by looking only at the text. You do not need to fill it out now—just recognize the format, because your checklist will look the same.
R1 Is it no more than 80 words? yes/no R2 Does the first sentence mention something specific to the customer? yes/no R3 Is there exactly ONE closing question? yes/no R4 Does it avoid promising a timeline? (invariant) yes/no R5 Is any discount no more than 10%? (invariant) yes/no R6 Does it contain no product errors (checked against source)? yes/no R7 Tone: no double exclamation, “must-have,” or ALL CAPS? yes/no
Notice three things. Every question is countable or visible—words, sentences, or something present or absent in the text. Two questions are marked as invariant: they come from answer 3 on the loop worksheet, what never changes on its own. A version that fails one in even a single case loses, regardless of everything else. No question uses “appropriate,” “good,” or “professional.”
For the real estate agent, the checklist for portal replies is similar, with different questions: does it name the property or address the lead viewed? Does it offer two visit times—not one or three? Does it avoid quoting below-list price (invariant)? Does it avoid “once-in-a-lifetime opportunity”?
How to write your checklist
You are not the one who answers the checklist day to day—the Evaluator is. It must be checked against your answers because judges drift. The product asks for one manual measurement task across the whole loop: answer yes or no to each checklist question for 10 to 20 texts. Twenty minutes, once. Those texts and your answers become the calibration set.
At every cycle, before judging anything new, the Evaluator answers those same texts. If it agrees with you on fewer than 80% of answers, it is labeled “uncalibrated judge”: the cycle promotes nothing, and the Meta-agent—the assistant that only observes and reports on the loop—alerts you. If one checklist question has under 80% agreement, the question is the problem and is rewritten or removed.
Two more safeguards require no action from you, but are worth knowing: the Evaluator compares versions in pairs, with shuffled order, three times, and takes the majority (preventing a bias toward the second one); and, when the budget allows, it runs on a different AI model from the one that wrote the text, so it does not approve its own style.
The support manager selected 15 ticket replies—8 she would have sent and 7 she would have returned—and marked yes/no for her 6 checklist questions. In the first calibration, the Evaluator agreed 71% of the time. One question (“is the reply cordial?”) caused almost all disagreement. She replaced it with “does the reply address the customer by name and end by asking whether the issue is resolved?” Agreement: 93%.
A text’s score is simply the number of “yes” answers. To compare the current version with a candidate, the loop does not inspect one proposal; it uses a fixed set of 20 to 50 saved cases and generates both versions for each. Version B “wins on the deliverable” if its average yes score is at least one criterion above A’s and and does not fail any invariant in any case.
This test runs in minutes, without customers, every week. That is why the loop keeps turning while the signal metric is still gathering 300 runs: while the world responds slowly, the text is checked quickly.
The real estate agent saved 30 old leads and their messages as checklist cases. The new reply averaged 5.4 out of 6 versus 4.1 for the old one—but in one of the 30 cases, it quoted a price below the list price. It lost. One invariant violation outweighs a 1.3-point average advantage.
Common mistake
Review 12 proposals, decide the new version is better, and make it the rule. This happens because we notice what confirms our expectations, and because 12 is below even the minimum for a deliverable test (20 to 50 cases). Avoid it with a fixed set of checklist cases saved before the new version exists, and count yeses—never rely on an informal read.
Practice now 0/6 complete
Leave with your checklist: 6 questions, 2 of them invariants, answered by you and another person for the same 6 texts. About 15 minutes. It is ready when the other person can fill all 36 boxes without asking “depends on what?”
Nothing in your business changes until you approve it: the checklist is just paper until it goes on the loop worksheet. Even then, it only checks text—it does not send, change, or delete anything. If the checklist seems wrong, cross it out and rewrite it; a bad version costs nothing.
Example of an acceptable result (support manager): 1. Does it address the customer by name? 2. Does it restate the issue in one sentence before answering? 3. Is it at most 120 words? 4. Does it avoid promising a fix date? (invariant) 5. Does it avoid asking for information already in the ticket? 6. Does it end by asking whether the issue is resolved? (invariant)
You have just written the ruler the loop will use every week—and confirmed that two people read it the same way.
Summary
Track 04 · Lesson 4
By the end of this lesson, you will be able to read a 5-line decision card and answer approve, reject, or wait—and write answers 4 and 5 on the loop worksheet: how much a cycle can cost and who approves.
After the test comes the moment that decides whether the loop is worth anything: someone must say, “yes, make it the official version.” If this decision lands in your inbox constantly, you will stop reading by month three. If it becomes automatic too soon, a worse version slips in unnoticed. The card solves this with one rule: one decision per week, five lines, three answers—and always a rollback button.
↓ scroll to study
The decision card is the only thing in the loop that needs your attention. It arrives at most once a week and has five fixed lines: the hypothesis tested; the result with sample size and margin; guardrail results; cost versus the ceiling; and the three options. Nothing else. If it needed more, no one could fit it into Friday.
What follows looks like a receipt, and it is one: each line is a number you already know how to read from previous lessons. This is the card the beauty clinic owner received in cycle four of the real case followed in this course.
Read from top to bottom as you would read a test result: first what was tested, then whether the main number moved and with how many runs (300 per side was the required sample), then whether any guardrail failed, then the cost. Only then decide. Sales conversion is not shown because it was not the test target—it was monitored and did not fall.
Before the card reaches you, the Evaluator has already given one of three verdicts: B won by a margin, A stays, or insufficient sample. The card translates that into your three options. Approve: the candidate becomes the official version. Reject: the current version stays, and the hypothesis goes into memory as discarded, with the reason. Wait: the test continues for one more period.
What if you do not answer? The safe default applies: nothing is promoted, and the history records “decision pending.” The loop does not sit idle—it keeps running the current version—but it does not decide for you.
The real estate agent received the card during a Friday shift and did not open it. On Monday, the history showed “cycle 0003: pending for 3 days.” The old version kept replying to leads, unchanged. When he opened the card on Tuesday, it was still the same—the loop had made no decision on his behalf.
Common mistake
Approving because the new version “looks better.” This happens because the new text looks nicer, or because you proposed the hypothesis. Avoid it by approving only when line 2 says B won by a margin and the sample reached its minimum. If your preference disagrees with line 2, line 2 wins—the loop stores the hypothesis for a future test, not for forgetting.
In the clinic’s second cycle, the hypothesis was “proposals with at most 80 words.” The primary result was excellent: replies rose from 19.3% to 29.7%, with 300 proposals per side—a ten-point gain, not luck. But the guardrail line showed 6 complaints versus 1, and the tolerance complaint tolerance was zero. Verdict: A stays. The hypothesis was discarded.
The system did exactly what it should: applied the rule as written, did not promote, and recorded the result. The rule was wrong. Complaints are rare, about 1%: with 300 proposals, getting 6 on one side and 1 on the other can easily happen by chance. It was noise, not a real difference. Zero tolerance turned noise into rejection.
The lesson is not “ignore the guardrail when the gain is large”—that would destroy the loop’s only strong guarantee. Tolerance is a calculation, not a preference: for a 1% event in 300 runs, normal fluctuation is about one point, so the right tolerance is 0.01, not zero. The clinic owner adjusted the worksheet for future tests. The discarded hypothesis stayed in memory with its reason—and two cycles later, a different hypothesis truly won.
The card’s fourth line—cost—exists because you set a ceiling on the loop worksheet. Answer 4 is how much a cycle may cost and how much of your time it may require each week. The Guardian stops the cycle when the ceiling is reached. Without one, a loop can cost more than the gain it pursues, and no one notices until the bill arrives.
The Meta-agent—the assistant that only observes and reports on the loop—also monitors a second number from the first cycle: cost per hypothesis that became official. If nothing is promoted in five cycles, it proposes one of three options: run less often, use cheaper assistants for reading steps, or load less memory. It proposes; you decide.
In the product’s sample worksheet, the cost line reads “R$ 41 this cycle (R$ 50 ceiling).” The amount is illustrative, but the format is always spend versus limit. The support manager set her ceiling in two parts: up to R$ 60 per cycle and at most one hour of her attention per week—to read the card and evidence, if she chooses.
How to fill in worksheet answers 4 and 5
Answer 5 on the worksheet is who approves, and it has three levels. L0: the loop only reports; nothing becomes official. L1: the loop proposes; you approve every card. L2: when the verdict is B won, the loop promotes automatically and notifies you—but only after five correct L1 promotions. L3 (the loop redesigning itself) exists, and the last lesson in this track explains why it is still research.
What makes L2 acceptable and L1 safe is the rollback button. In the first cycle after a promotion, if the test metric or any guardrail falls below the previous version beyond tolerance, the previous version returns. At L1, the card alerts you and asks for confirmation. At L2, it rolls back automatically and notifies you. In both cases, the history keeps the reason: nothing disappears or is overwritten.
The real estate agent promoted the new lead reply in cycle 5. In cycle 6, bookings fell from 31% to 19%, below the 22% achieved by the old version. The card had a different line: “drop beyond tolerance; revert to previous version? [yes] [keep].” He reverted. The promoted version remained in the history with the date, numbers, and the label “reverted,” so it would never be tested again as if it were new.
Practice now 0/5 complete
For each card, write approve, reject, or wait, with a one-line reason. Then fill in answers 4 and 5 on your worksheet. About 12 minutes. This is analysis because a real decision changes a business’s official version; here you practice with prepared cards, two of them real, before deciding for your own business.
Nothing in your business changes until you approve it: cards 1 and 2 are from the real clinic case (already decided); card 3 was created for this exercise. Your answers 4 and 5 go on the worksheet only when you take them to the course’s final lesson.
Card 1—reject (A stays). The written rule says zero tolerance and a guardrail worsened; the correct card answer is not to approve. Next, do what the clinic owner did: recognize that 6 versus 1 out of 300 is noise for a rare event, adjust complaint tolerance to 0.01 on the worksheet, and keep the hypothesis in memory as discarded, with the reason, so it can be retested under the right rule. Neither “approve anyway” nor “zero tolerance forever.”
Card 2—approve. B won by a margin, the sample reached its minimum, and no guardrail fell. Conversion did not move significantly—but it was monitored, not targeted, and did not worsen. The version becomes official; rollback is armed for the next cycle. That is what actually happened: it was promoted as v2.
Card 3—wait. 140 is less than half the required sample. A nine-point lead with this sample is “insufficient sample,” with no “but it is winning.” No guardrail fell, so there is no reason to stop: the test continues to 300 per version.
Answers 4 and 5 (acceptable example for the clinic). Ceiling: R$ 50 per cycle, 1 hour of the owner’s time per week. Starting level: L1. Card response deadline: 7 days; no response means nothing is promoted. Move to L2 after 5 L1 approvals with no rollbacks. Your worksheet may have different numbers—that is fine; each line must have a number.
You have just decided on two real cards and one constructed card—and written your loop’s ceiling and approver as numbers.
Summary
Track 04 · Lesson 5
By the end of this lesson, you will be able to state, in one sentence each, the four things the loop guarantees by design and the four it does not—and recognize a false promise when someone sells one.
This entire track has asked for evidence. It would be strange to end by making a promise without any. So this is the honesty lesson: what an improvement loop guarantees because it was built to, and what depends on your volume, your market, and luck. Knowing the difference keeps you from abandoning the loop in month three out of frustration—and from trusting it beyond its limits.
↓ scroll to study
No learning loop guarantees that a number will rise. Not even this one—the beauty clinic owner heard that on the product’s first screen, before filling in the worksheet. LOOP-R guarantees four things, all by design: they hold because the rules from earlier lessons exist, not because someone promised. A worse version never replaces the current one by system decision. Every change is recorded and reversible. The cycle always runs the same way. Cost never exceeds the ceiling.
Beside each guarantee is what it does not cover. Not replacing a version with a worse one does not mean the next will be better. Having a record does not mean it contains anything useful—it may say “no evidence” for ten cycles in a row. Always running does not mean every cycle produces a good hypothesis. Staying under the ceiling does not mean the cost is worthwhile.
The rule this product follows in every text: the word “guarantee” appears only alongside non-regression, a reversible record, or a consistent cycle. If you read “guarantees more sales” in any material about AI loops—including this one—you are reading a promise no one can make.
Before—the promise being sold
“The AI will learn from every proposal and improve your sales.” No number, no condition, no timeline.
After—what is guaranteed
For the clinic: no version with worse replies or margin worse becomes official; every change has a date, metric, and rollback; the cycle runs every Monday; spending stops at R$ 50.
Takeaway: four statements that can be checked in the history, instead of one that cannot be verified anywhere.
In lesson 2, you replaced revenue with a signal as the test target because revenue takes too long. That has a cost worth stating: raising the signal does not prove that revenue rises. It is plausible—people who reply are more likely to buy—but plausible is not proven. Until the volume is there, the link between them is an explicit bet, not a result.
The clinic’s real case shows this plainly. In cycle 0004, the reply rate rose from 19.3% to 29.3%, and the version was rightly promoted. Sales conversion went from 3.0% to 3.7%—a figure that, with 300 proposals on each side, cannot be distinguished from luck. The cycle report says exactly that: “the signal → target link remains unproven.” Ten points more replies, and still nothing to claim about sales.
The real estate agent faces the same link under different names: more booked visits have not yet proved more closed sales. His loop optimizes bookings and monitors closings. Once there are 1,500 leads on each side, closings may become the target. Before then, saying “the loop increased sales” would be the same unsupported promise.
Test yourself
Clinic cycle 0004: replies 19.3% → 29.3% (N = 300 vs. 300, significant); conversion 3.0% → 3.7% (not significant). What can you claim?
The loop guarantees that a cycle runs every week. It does not guarantee that the week brings a good idea. It is possible—and common in a small business—for the Critic to find only weak observations for several cycles, for the Optimizer to have nothing to propose, and for the history to accumulate “insufficient evidence” line after line. That does not mean the loop is broken. It is telling the truth about its volume.
Two things still matter during those weeks. The memory of discarded hypotheses grows and is as valuable as the memory of promoted ones: it prevents you or a future teammate from testing the same bad idea twice. The Meta-agent, which only observes and reports, tracks cost per promoted hypothesis. If nothing is promoted in five cycles, it suggests running less often instead of pretending to make progress.
The support manager went through six such cycles between March and April: too few tickets per type for the Critic to identify a pattern. Following the Meta-agent’s suggestion, she reduced the loop to every two weeks and let the log spreadsheet accumulate data. In June, with 30 tickets in the most common category, the first moderate hypothesis appeared. Her history has six “no evidence” entries—and she considers all six honest.
One part of the LOOP-R idea is not ready for any business: the loop redesigning itself—level L3, where the Meta-agent proposes replacing assistants, changing the cycle order, or rewriting its own rules. That is research. In the product taught in this course, the Meta-agent only reports: cost, hypotheses generated versus promoted, who gets the most Guardian vetoes, and whether the judge is calibrated. No action. Anyone promising a system that “improves itself” is selling something that does not yet exist safely.
And there is a bottleneck no assistant can solve: the log spreadsheet. Without evidence, everything else is theater—and your business’s evidence is in WhatsApp, the reception notebook, or the salesperson’s head. The product connects to one thing: a spreadsheet with columns defined in the worksheet. You fill it in or export data. Each direct connection to another system is its own project; none comes free with the five answers.
For both the clinic and the real estate agent, the conclusion is the same—and it tests any promise: LOOP-R as a process discipline—logging, tests with sufficient sample sizes, and non-regression—has value on its own today. LOOP-R as a “self-learning system” depends on a volume of data most small businesses do not have. An honest product says so on its first screen.
The loop does not promise the number will rise. It promises it will not fall because of a system decision.
Practice now 0/3 complete
Mark each statement T or F. For false ones, write one line naming the real guarantee it stretches. About 8 minutes. This is analysis: you cannot “practice” a guarantee; you learn to recognize it when reading sales proposals.
Nothing changes in your business until you approve it: this is reading and marking. Mistakes help—every false statement you mark true is a promise someone may try to sell you.
1 — F. It stretches non-regression: the loop guarantees the number will not fall by its decision, not that it will rise.
2 — T. This is the core guarantee: the Evaluator promotes only when margin and guardrails remain intact; otherwise, “A stays.”
3 — T. A record with a rollback button. Every promotion and rollback appears in the history with its date, numbers, and reason.
4 — F. It stretches the cycle’s consistency: the loop keeps running, but the default when there is no answer is not to promote and to record “pending.”
5 — T. The Guardian stops the cycle when it reaches the ceiling.
6 — F. It stretches the ceiling: staying within the limit guarantees the spend, not the return. Cost per promoted hypothesis is a separate calculation, reported by the Meta-agent.
7 — F. It stretches consistency: running every week is guaranteed; producing a good hypothesis each week is not. Ten “no evidence” results are honest, and discarded ideas remain in memory.
8 — T. Cycle consistency. It is the cheapest guarantee and the best protection against “we stopped doing it in month 3.”
You have just separated, claim by claim, what the loop can support from what no one can support—the same reading you will apply to the next promise you hear.
Summary