INEMA.CLUBPROLOOP-R PT · EN · ES

LOOP-R · Track 04

Measure what matters

Learn to say how many runs prove an improvement, write a yes-or-no checklist, and decide on a real card—approve, reject, or wait—with rollback at hand.

Track 04 · Lesson 1

The loop that learns
to to game the target

By the end of this lesson, you will be able to take a business target and write, in one line, what AI might do to inflate the number without improving anything—and which two metrics you will monitor to prevent it.

Any target given alone to a learning system becomes something to game. “Reply faster” turns into closing tickets without resolving them; “sell more” turns into discounting until the margin disappears. This is not an AI flaw: it is what any employee might do when judged by a single number. Before measuring for real, you need to know how the measure can be gamed.

↓ scroll to study

01 A target on its own teaches the system to cheat

There is an old rule in every business: when a measure becomes a target, it stops measuring. Not because people are dishonest, but because they are smart. If you pay by the number of support interactions, those interactions get shorter. A learning system does the same thing, only faster and without guilt.

In LOOP-R, the target metric is the number you asked the loop to increase. If it is the only number it watches, it will find the shortest path to it—and that path rarely means serving people better.

The owner of a beauty clinic asks, “increase the proposal response rate.” Three cycles later, it has risen from 20% to 34%. She opens the conversations and finds out why: the new version sends a second message the next day, and many of the “replies” say, “please stop sending me this.” They counted as responses. The needle went up; the clinic did not.

02 Cheating takes three familiar forms

You do not need to imagine every possible way to cheat. The same patterns repeat. Ask for maximum clicks and the loop learns clickbait. Ask it to minimize handling time and it learns to close tickets quickly without resolving them. Ask for more sales and it learns to discount until the margin disappears. Ask for more replies and it learns to provoke people.

Notice the pattern: in every case, the requested number really does go up. The trick is not in the number; it is in what the number leaves out.

A support manager at a small software company set a target of “average ticket time, from 18 to 10 minutes.” In two weeks it reached 9. But tickets closed in 2 minutes reopened the next day, and satisfaction fell from 4.3 to 3.6. None of this appeared on her dashboard, which showed only time.

Before—a bare target

“Average ticket time: from 18 to 10 minutes.” A ticket closed in 2 minutes and reopened tomorrow counts in favor of the target.

After—a target with guardrails

“From 18 to 10 minutes without worsening 7-day reopenings (currently 9%) or satisfaction (currently 4.3).” The same ticket now counts against it.

Takeaway: the same target and dashboard—but “close without resolving” is no longer a shortcut; it automatically fails.

03 The required format: maximize X without worsening Y, Z, W

In LOOP-R, a target never stands alone. Every target metric is paired with at least one number that must not fall—a guardrail. The complete statement is always “maximize X without worsening Y, Z, and W.” This is not decoration; it is the part of the target that prevents gaming.

Typical guardrails vary by area. Sales: average margin, complaints, unsubscribes, average discount. Support: satisfaction, 7-day reopenings, customer effort. Content: time on page, unsubscribes, reports. You do not need to watch them all—only those most likely to be affected by the shortcut.

A real estate agent who replies to portal leads writes: “maximize booked visits per lead without worsening the no-show rate or average discount offered.” If the new version books more visits by promising discounts in conversation, the average discount rises—and the version fails, even with a full calendar.

Two loop assistants use this statement. The Guardian vetoes any hypothesis that proposes changing a guardrail. The Evaluator rejects any version that raises X but worsens a guardrail beyond tolerance—even if X rose a lot. If replies rise 40% and margin falls 1% with zero tolerance, the answer is “A stays.” No exceptions.

Test yourself

Version B raised the response rate by 40%. Average margin fell 1%, and the tolerance for the “margin” guardrail is zero. What does the Evaluator answer?

04 The count itself can be gamed

There is a quieter trick than the three above: it happens in the definition of the number. What exactly counts as a “reply”? If any message back counts, “stop sending me this” is a reply. If “resolved” means any closed ticket, a forcibly closed ticket counts as resolved.

That is why the metric definition belongs on the loop worksheet; it is not a detail. “Reply” must exclude negative replies; “resolved” must exclude tickets reopened within 7 days. The Observer—the assistant that counts—checks that the log matches this definition before adding anything up.

At the clinic, the “replied” column in the log gained a neighbor: “positive reply (yes/no).” At the software company, “closed on” gained “reopened within 7 days (yes/no).” Two extra columns, and both tricks from the start of this lesson disappeared from the count.

A target without guardrails is not a target. It is an invitation.

Practice now 0/4 complete

Find the trick in three targets

For each target below, write the most likely trick and two guardrails. About 10 minutes. This is analysis, not execution: running a loop until it cheats would take weeks; the completed case shows in minutes what you need to recognize.

Nothing in your business changes until you approve it: you are writing on paper, not on the loop worksheet. If your answer differs from the key, compare the reasoning—each target can have more than one possible trick and more than one valid guardrail.

Target 1—Renata, beauty clinic owner. “I want to double the number of proposals that get a reply on WhatsApp.”

Target 2—Marcos, real estate agent. “I want more visits booked per lead from the portal.”

Target 3—Paula, support manager. “I want to reduce the average time per ticket.”

Annotated answer key

Target 1 (clinic). Trick: keep messaging (second and third messages) or provoke, because “stop sending me this” counts as a reply. Guardrails: complaints and unsubscribes (requests to stop receiving messages). Extra: define “reply” as a positive reply in the log.

Target 2 (real estate agent). Trick: book at any cost—promise a discount or say “just come take a look,” filling the calendar with visits that never happen. Guardrails: no-show rate and average discount promised. Acceptable alternative: proposals closed per visit.

Target 3 (support). Trick: close tickets quickly without resolving them. Guardrails: 7-day reopenings and satisfaction. Extra: “resolved” must exclude reopened tickets.

If you chose a different guardrail, ask: “Would the trick I described affect this number?” If yes, it is valid.

You have just spotted the trick before it happens—in three different businesses, using the same question.

Summary

  • When a measure becomes a target, it stops measuring—and a learning system finds the shortcut faster than any employee.
  • The tricks repeat: clickbait, closing without resolving, discounting until the margin disappears, provoking people to “win” a reply.
  • In LOOP-R, a target has a fixed format, “maximize X without worsening Y, Z, and W”; the Guardian vetoes changes that touch a guardrail, and the Evaluator rejects versions that breach one.
  • The number’s definition is also part of the target: “reply” excludes negative replies, and “resolved” excludes reopened tickets.

Your next step

You can now look at a target and spot the shortcut before AI finds it.

Over the next 15 minutes: take the main target for the process you chose in Track 2 and rewrite it as “maximize X without worsening Y and Z.” Send the sentence to yourself on WhatsApp—it goes on the loop worksheet.

In the next lesson, you will learn why 12 proposals prove nothing, how many are needed—and how many weeks that takes at your pace.

Track 04 · Lesson 2

The sample size
that proves it

By the end of this lesson, you will be able to look at a target like “from 20% to 30%” and say how many runs the test needs per version and how many weeks it will take at your business’s pace—before you start.

Most loops die here: someone compares 12 new proposals with 12 old ones, sees “double,” and changes the team’s rule. Three weeks later, the number is back to normal and no one knows why. Luck with a small sample looks like a pattern. There is a calculation that separates the two—and it fits in a five-row table.

↓ scroll to study

01 With twelve, luck disguises itself as a pattern

Flip a coin twelve times. Getting eight or more heads happens in nearly one out of five tries—with a fair coin, no trick. Now replace “heads” with “customer who booked” and “tails” with “customer who did not reply”: with twelve customers, a new message “beating” the old one is what chance predicts, not proof that it is better.

That is why “I tested it and it worked” is rarely true with a small sample. The person is not wrong about what they saw; the same thing can happen when there is no difference at all.

A real estate agent changes the first message sent to portal leads. Of the next 12 leads, 4 book a visit. With the old message, 2 of the previous 12 did. “It doubled,” he says at Monday’s meeting, and everyone switches to the new message. Three weeks later, the rate is back where it always was—and no one knows whether the new message is better, worse, or the same.

02 The sample-size table

The sample size needed for proof is a standard calculation. The Experimenter—the assistant that designs the test—works it out and shows you before the test starts. You do not need to know the formula; you need to know how to read the table.

from → toper versiontotalat 50/weekat 200/week
3% → 5%~1.500~3.00060 semanas15 semanas
3% → 6%~700~1.40028 semanas7 semanas
20% → 30%~300~60012 semanas3 semanas
20% → 35%~140~2806 semanas1,5 semana
50% → 65%~170~3407 semanas2 semanas

Notice two things. First: the lower the starting number, the more runs it requires—3% is almost impossible to prove in a small business. Second: the bigger the lift you expect, the fewer runs you need—but the lift must be real, not wishful thinking.

The clinic owner sends 50 proposals a week and wants to raise sales conversion from 3% to 5%. The table calls for 1,500 proposals per version: 60 weeks for one test. More than a year to answer one question. The loop is not broken; that question cannot be answered at that pace.

03 Three kinds of metrics, three speeds

The clinic’s answer is not to give up measuring; it is to measure a faster number that still relates to revenue. LOOP-R separates metrics into three types. Revenue: conversion, order value, margin—the customer determines these weeks later; proving them takes months. Signal: reply rate, bookings, useful clicks—the customer responds in hours or days; proof takes weeks. Deliverable: does the proposal meet the checklist? Is there an error? Is the tone right? You can answer by reading the text, in minutes.

The loop’s rule: test the fastest metric that still has a plausible link to revenue. Revenue does not disappear; it becomes a guardrail (“without worsening conversion”) until there is enough volume to test it directly.

At the software company, the support manager has all three: revenue is contract renewal (months); the signal is “resolved on first contact” (days); the deliverable is “does the reply meet the checklist?” (minutes). Her loop runs weekly on the deliverable and signal, while monitoring renewals.

Before—measuring revenue

Clinic, 50 proposals/week. Target: sales conversion from 3% to 5%. Required sample: 1,500 per version. 60 semanas until the first reliable result.

After—measuring the signal, monitoring revenue

Same clinic, same pace. Target: reply rate from 20% to 30%, without worsening conversion. Required sample: 300 per version. 12 semanas.

Takeaway: 48 fewer weeks to a reliable first answer—revenue stays in view, just no longer as the target.

04 “Insufficient sample” is the most common answer—and it must be acceptable

The Evaluator, the assistant that judges the test, can give only three answers: B won by a margin, A stays, ou insufficient sample—continue the test. If the number of runs has not reached the required sample size, the answer is the third one. Always. There is no “but it is winning.”

Stopping early because it is winning is the most common way to turn noise into a rule. The previous step’s table says when a test ends—not impatience to see the result. The only allowed early stop is a falling guardrail: the test stops because it is causing harm, not because it is going well.

In the real beauty-clinic case followed by this course, the first cycle nearly started on the wrong foot: the Optimizer expected a small effect, and the Experimenter warned that the test would take 40 weeks. The owner resized it around the target on the worksheet (20% to 30%): 300 proposals per version. The test fit the calendar before the first message was sent.

Test yourself

Week 5 of a 12-week test for the real estate agent: message B has 31% booked visits versus 22% for A, with 140 leads per side. The required sample is 300 per version. What should you do?

05 Segmenting requires thirty per segment

After the loop counts, the Critic—the assistant that says what worked and failed—will want to segment by customer type, channel, or day of the week. Every segment obeys the same small-sample rule. “Tuesdays convert better” based on 12 proposals is not a pattern. The Critic can call something a pattern only with at least 30 runs in that segment.

That is why every claim it makes has a strength label. Strong: 90 or more runs—becomes a priority hypothesis. Moderate: 30 to 89—becomes a hypothesis. Weak: fewer than 30—goes into memory as “observed, untested” and is reviewed when the sample grows. Nothing is lost; nothing becomes a rule too early.

At the clinic, the Critic noticed that facial-treatment customers replied more often than body-treatment customers: 14 on one side and 11 on the other. Label: weak. The observation was stored; three months later, with 60 on each side, it became a moderate hypothesis to test. If it had become a rule in the first month, it would have been a rule based on 25 people.

Twelve proposals say nothing. Three hundred start to speak.

Pratique agora 0/4 feito

Choose the metric you can prove

For each business below, choose a realistic metric, state the number of runs per version (use the table in step 02), and calculate how many weeks it takes at the stated pace. About 10 minutes. This is analysis: a real test takes weeks, but here you do the Experimenter’s calculation in minutes.

Nothing in your business changes until you approve it: you are only doing calculations on paper. If your number differs from the key, check that you read the right table row—the most common mistake is using “total” instead of “per version.”

Case 1—Renata, beauty clinic. 50 proposals a week. Today, 20% reply and 3% buy. She wants to “sell more.”

Case 2—Marcos, real estate agent. 200 portal leads per week. Today, 20% book a visit; he thinks the new message will raise it to 35%.

Case 3—Paula, support. 400 tickets per week. Today, 50% are resolved on first contact; the target is 65%.

Annotated answer key

Case 1. “Sell more” is a revenue metric (3% → 5%): 1,500 per version, 60 weeks at 50/week. Not feasible. A realistic metric is reply rate, 20% → 30%, with conversion as a guardrail: 300 per version, 600 total, 12 weeks.

Case 2. Booking is a signal, and he has enough volume: 20% → 35% requires 140 per version, 280 total. At 200/week: a week and a half. If the real effect is smaller than he expects (only 30%), the requirement rises to 300 per version and 3 weeks—the Experimenter designs for the worksheet target, not for hope.

Case 3. Resolved on first contact is a signal; 50% → 65% requires 170 per version, 340 total. At 400/week: less than a week. Guardrails: 7-day reopenings and satisfaction, or “resolved” becomes “closed.”

You have just done the calculation that separates luck from a pattern—for three different business paces.

Summary

  • With a small sample, chance produces “wins” all the time—eight heads in twelve coin flips happens once in five tries.
  • The required sample size is calculated and shown in advance by the Experimenter: 3% to 5% costs 1,500 per version; 20% to 30% costs 300.
  • When revenue takes too long to test, the loop tests a faster signal linked to it and monitors revenue as a guardrail.
  • Before the finish line, the only answer is “insufficient sample”; the only early stop is a falling guardrail, never an advantage.
  • Segmenting requires 30 per segment; below that, the observation stays stored as “observed, untested.”

Your next step

You can now use the table to say whether a test fits your business calendar.

Over the next 15 minutes: count how many runs your process gets per week (proposals, leads, tickets) and find your target’s row in the table. Write “N per version, X weeks” beside the target you sent yourself on WhatsApp in the previous lesson.

In the next lesson, you will learn to measure what does not depend on a customer reply: the yes/no checklist, which runs in minutes—and is where the loop works every week.

Track 04 · Lesson 3

The checklist:
yes or no

By the end of this lesson, you will be able to write a checklist of 6 yes-or-no questions about the text your AI produces—questions that two different people answer the same way, without debate.

When a result depends on a customer replying, a test takes weeks. But much of what you want to know is in the text itself: is it under 80 words? Does it mention something the customer said? Does it promise a timeline? You can check that in minutes. But “give it a score from 1 to 10” does not work—two people, or the same AI two weeks apart, will not give the same score.

↓ scroll to study

01 A 1-to-10 score measures the evaluator’s mood

Ask two people to score a sales proposal. One gives it 7, the other 8, and both are right—because “7” means nothing that can be checked. Ask an AI for a score and repeat next week: it changes too. Text evaluators, human or not, have known biases: they prefer longer text, prefer the second text they read, and change their minds from week to week.

The answer is not to find a better judge. It is to change the question. “Is this proposal good?” cannot be checked. “Does this proposal contain exactly one closing question?” can: yes or no, and anyone reading the text will answer the same way. A checklist is just that: questions that leave no room for personal taste.

The clinic owner spent months asking her assistant to “improve this proposal” and getting different versions without knowing whether they were better. She switched to “does the proposal mention something specific the client said? Is it at most 80 words? Does it promise a timeline?” Now she and the receptionist give the same answers about the same proposal—for the first time.

02 A sample checklist: seven questions for a proposal

What follows looks like a car inspection form, and that is exactly what it is: each row is a question answered by looking only at the text. You do not need to fill it out now—just recognize the format, because your checklist will look the same.

R1  Is it no more than 80 words?                              yes/no
R2  Does the first sentence mention something specific to the customer? yes/no
R3  Is there exactly ONE closing question?                    yes/no
R4  Does it avoid promising a timeline? (invariant)           yes/no
R5  Is any discount no more than 10%? (invariant)             yes/no
R6  Does it contain no product errors (checked against source)? yes/no
R7  Tone: no double exclamation, “must-have,” or ALL CAPS?     yes/no

Notice three things. Every question is countable or visible—words, sentences, or something present or absent in the text. Two questions are marked as invariant: they come from answer 3 on the loop worksheet, what never changes on its own. A version that fails one in even a single case loses, regardless of everything else. No question uses “appropriate,” “good,” or “professional.”

For the real estate agent, the checklist for portal replies is similar, with different questions: does it name the property or address the lead viewed? Does it offer two visit times—not one or three? Does it avoid quoting below-list price (invariant)? Does it avoid “once-in-a-lifetime opportunity”?

How to write your checklist

  1. Set aside 5 texts you approved and 5 you sent back for revision. Real ones, not made-up examples.
  2. Write what the good ones have that the bad ones do not—as a question. “Does it mention something the customer said?” not “Is it personalized?”
  3. Remove every question that needs an answer of “it depends.” If you hesitated to answer, the judge will too.
  4. Mark the invariants. Anything from answer 3 on the worksheet becomes a question labeled “(invariant).”
  5. Answer the questions yourself for all 10 texts. If a “good” text gets fewer yeses than a “bad” one, the checklist measures the wrong thing.
  6. Ask someone else to answer without talking to you. Where you disagree, rewrite the question.

03 Calibrating the judge: twenty minutes, once

You are not the one who answers the checklist day to day—the Evaluator is. It must be checked against your answers because judges drift. The product asks for one manual measurement task across the whole loop: answer yes or no to each checklist question for 10 to 20 texts. Twenty minutes, once. Those texts and your answers become the calibration set.

At every cycle, before judging anything new, the Evaluator answers those same texts. If it agrees with you on fewer than 80% of answers, it is labeled “uncalibrated judge”: the cycle promotes nothing, and the Meta-agent—the assistant that only observes and reports on the loop—alerts you. If one checklist question has under 80% agreement, the question is the problem and is rewritten or removed.

Two more safeguards require no action from you, but are worth knowing: the Evaluator compares versions in pairs, with shuffled order, three times, and takes the majority (preventing a bias toward the second one); and, when the budget allows, it runs on a different AI model from the one that wrote the text, so it does not approve its own style.

The support manager selected 15 ticket replies—8 she would have sent and 7 she would have returned—and marked yes/no for her 6 checklist questions. In the first calibration, the Evaluator agreed 71% of the time. One question (“is the reply cordial?”) caused almost all disagreement. She replaced it with “does the reply address the customer by name and end by asking whether the issue is resolved?” Agreement: 93%.

04 How the checklist becomes a verdict

A text’s score is simply the number of “yes” answers. To compare the current version with a candidate, the loop does not inspect one proposal; it uses a fixed set of 20 to 50 saved cases and generates both versions for each. Version B “wins on the deliverable” if its average yes score is at least one criterion above A’s and and does not fail any invariant in any case.

This test runs in minutes, without customers, every week. That is why the loop keeps turning while the signal metric is still gathering 300 runs: while the world responds slowly, the text is checked quickly.

The real estate agent saved 30 old leads and their messages as checklist cases. The new reply averaged 5.4 out of 6 versus 4.1 for the old one—but in one of the 30 cases, it quoted a price below the list price. It lost. One invariant violation outweighs a 1.3-point average advantage.

Common mistake

Review 12 proposals, decide the new version is better, and make it the rule. This happens because we notice what confirms our expectations, and because 12 is below even the minimum for a deliverable test (20 to 50 cases). Avoid it with a fixed set of checklist cases saved before the new version exists, and count yeses—never rely on an informal read.

Practice now 0/6 complete

Write 6 yes-or-no questions about your text

Leave with your checklist: 6 questions, 2 of them invariants, answered by you and another person for the same 6 texts. About 15 minutes. It is ready when the other person can fill all 36 boxes without asking “depends on what?”

Nothing in your business changes until you approve it: the checklist is just paper until it goes on the loop worksheet. Even then, it only checks text—it does not send, change, or delete anything. If the checklist seems wrong, cross it out and rewrite it; a bad version costs nothing.

Example of an acceptable result (support manager): 1. Does it address the customer by name? 2. Does it restate the issue in one sentence before answering? 3. Is it at most 120 words? 4. Does it avoid promising a fix date? (invariant) 5. Does it avoid asking for information already in the ticket? 6. Does it end by asking whether the issue is resolved? (invariant)

You have just written the ruler the loop will use every week—and confirmed that two people read it the same way.

Summary

  • A 1-to-10 score measures the evaluator, not the text; a yes/no question answered by looking at the text measures the text.
  • The checklist contains countable or visible questions. Those from answer 3 on the worksheet are invariants: one “no” is enough to fail a version.
  • The Evaluator is calibrated against you once using 10 to 20 texts. Below 80% agreement, nothing is promoted, and the question causing disagreement is rewritten.
  • The verdict comes from counting yeses across 20 to 50 saved cases—a test that runs in minutes while customer outcomes take weeks.

Your next step

You can now turn “is it good?” into six questions anyone on your team can answer the same way.

Over the next 15 minutes: save the 10 texts you used (the 6 from the exercise and 4 older ones) in a folder. They are your first checklist cases—the new AI version will be judged against them, not against your memory.

In the next lesson comes the moment that determines whether the loop is worth anything: the 5-line decision card—and the real case where the clinic gained 10 points and the right answer was not to approve.

Track 04 · Lesson 4

The decision card
and the rollback button

By the end of this lesson, you will be able to read a 5-line decision card and answer approve, reject, or wait—and write answers 4 and 5 on the loop worksheet: how much a cycle can cost and who approves.

After the test comes the moment that decides whether the loop is worth anything: someone must say, “yes, make it the official version.” If this decision lands in your inbox constantly, you will stop reading by month three. If it becomes automatic too soon, a worse version slips in unnoticed. The card solves this with one rule: one decision per week, five lines, three answers—and always a rollback button.

↓ scroll to study

01 Five lines, at most once a week

The decision card is the only thing in the loop that needs your attention. It arrives at most once a week and has five fixed lines: the hypothesis tested; the result with sample size and margin; guardrail results; cost versus the ceiling; and the three options. Nothing else. If it needed more, no one could fit it into Friday.

What follows looks like a receipt, and it is one: each line is a number you already know how to read from previous lessons. This is the card the beauty clinic owner received in cycle four of the real case followed in this course.

CICLO 0004 — hypothesis: open the proposal by mentioning the customer’s review Result: reply rate 19.3% → 29.3% (N = 300 vs. 300, margin OK) Guardrails: complaints, unsubscribes, and average discount—all within tolerance Cost: within the cycle ceiling Decision: [approve] [reject] [wait 2 more weeks]

Read from top to bottom as you would read a test result: first what was tested, then whether the main number moved and with how many runs (300 per side was the required sample), then whether any guardrail failed, then the cost. Only then decide. Sales conversion is not shown because it was not the test target—it was monitored and did not fall.

02 Three answers, and only three

Before the card reaches you, the Evaluator has already given one of three verdicts: B won by a margin, A stays, or insufficient sample. The card translates that into your three options. Approve: the candidate becomes the official version. Reject: the current version stays, and the hypothesis goes into memory as discarded, with the reason. Wait: the test continues for one more period.

What if you do not answer? The safe default applies: nothing is promoted, and the history records “decision pending.” The loop does not sit idle—it keeps running the current version—but it does not decide for you.

The real estate agent received the card during a Friday shift and did not open it. On Monday, the history showed “cycle 0003: pending for 3 days.” The old version kept replying to leads, unchanged. When he opened the card on Tuesday, it was still the same—the loop had made no decision on his behalf.

Common mistake

Approving because the new version “looks better.” This happens because the new text looks nicer, or because you proposed the hypothesis. Avoid it by approving only when line 2 says B won by a margin and the sample reached its minimum. If your preference disagrees with line 2, line 2 wins—the loop stores the hypothesis for a future test, not for forgetting.

03 The real case: it gained ten points and still failed

In the clinic’s second cycle, the hypothesis was “proposals with at most 80 words.” The primary result was excellent: replies rose from 19.3% to 29.7%, with 300 proposals per side—a ten-point gain, not luck. But the guardrail line showed 6 complaints versus 1, and the tolerance complaint tolerance was zero. Verdict: A stays. The hypothesis was discarded.

The system did exactly what it should: applied the rule as written, did not promote, and recorded the result. The rule was wrong. Complaints are rare, about 1%: with 300 proposals, getting 6 on one side and 1 on the other can easily happen by chance. It was noise, not a real difference. Zero tolerance turned noise into rejection.

The lesson is not “ignore the guardrail when the gain is large”—that would destroy the loop’s only strong guarantee. Tolerance is a calculation, not a preference: for a 1% event in 300 runs, normal fluctuation is about one point, so the right tolerance is 0.01, not zero. The clinic owner adjusted the worksheet for future tests. The discarded hypothesis stayed in memory with its reason—and two cycles later, a different hypothesis truly won.

04 Worksheet answer 4: the ceiling

The card’s fourth line—cost—exists because you set a ceiling on the loop worksheet. Answer 4 is how much a cycle may cost and how much of your time it may require each week. The Guardian stops the cycle when the ceiling is reached. Without one, a loop can cost more than the gain it pursues, and no one notices until the bill arrives.

The Meta-agent—the assistant that only observes and reports on the loop—also monitors a second number from the first cycle: cost per hypothesis that became official. If nothing is promoted in five cycles, it proposes one of three options: run less often, use cheaper assistants for reading steps, or load less memory. It proposes; you decide.

In the product’s sample worksheet, the cost line reads “R$ 41 this cycle (R$ 50 ceiling).” The amount is illustrative, but the format is always spend versus limit. The support manager set her ceiling in two parts: up to R$ 60 per cycle and at most one hour of her attention per week—to read the card and evidence, if she chooses.

How to fill in worksheet answers 4 and 5

  1. Money ceiling per cycle. Set it a little above the cost of a typical cycle—the Guardian stops at the ceiling, so a ceiling that is too tight interrupts good tests.
  2. Your time ceiling per week. What you can sustain in month 4, not month 1. One hour is realistic; “whenever I can” is not a ceiling.
  3. Initial approval level: L1. You approve every card. Never start at L2, no matter how much you trust the idea.
  4. Deadline to answer the card. For example, 7 days. If it passes, nothing is promoted; it stays “pending.”
  5. Condition for moving up to L2. Five L1 approvals with no rollbacks. Write down the number; without it, “I will automate this later” never happens.

05 Answer 5: who approves—and the rollback button

Answer 5 on the worksheet is who approves, and it has three levels. L0: the loop only reports; nothing becomes official. L1: the loop proposes; you approve every card. L2: when the verdict is B won, the loop promotes automatically and notifies you—but only after five correct L1 promotions. L3 (the loop redesigning itself) exists, and the last lesson in this track explains why it is still research.

What makes L2 acceptable and L1 safe is the rollback button. In the first cycle after a promotion, if the test metric or any guardrail falls below the previous version beyond tolerance, the previous version returns. At L1, the card alerts you and asks for confirmation. At L2, it rolls back automatically and notifies you. In both cases, the history keeps the reason: nothing disappears or is overwritten.

The real estate agent promoted the new lead reply in cycle 5. In cycle 6, bookings fell from 31% to 19%, below the 22% achieved by the old version. The card had a different line: “drop beyond tolerance; revert to previous version? [yes] [keep].” He reverted. The promoted version remained in the history with the date, numbers, and the label “reverted,” so it would never be tested again as if it were new.

Practice now 0/5 complete

Decide on three cards—and write your ceiling

For each card, write approve, reject, or wait, with a one-line reason. Then fill in answers 4 and 5 on your worksheet. About 12 minutes. This is analysis because a real decision changes a business’s official version; here you practice with prepared cards, two of them real, before deciding for your own business.

Nothing in your business changes until you approve it: cards 1 and 2 are from the real clinic case (already decided); card 3 was created for this exercise. Your answers 4 and 5 go on the worksheet only when you take them to the course’s final lesson.

CARD 1 (real—clinic, cycle 0002)—hypothesis: proposals with at most 80 words. Result: reply rate 19.3% → 29.7% (N = 300 vs 300, margin OK). Guardrails: complaints 6 vs. 1—tolerance 0 · opt-outs and discounts within limits. Cost: within ceiling. Decision: [approve] [reject] [wait]
CARD 2 (real—clinic, cycle 0004)—hypothesis: open by mentioning the client’s review. Result: reply rate 19.3% → 29.3% (N = 300 vs. 300, margin OK). Guardrails: complaints, opt-outs, and discounts—all within limits · conversion 3.7% vs. 3.0% (not significant). Cost: within ceiling. Decision: [approve] [reject] [wait]
CARD 3 (constructed—real estate agent, cycle 0002)—hypothesis: offer two visit times in the first reply. Result: bookings 22% → 31% (N = 140 vs. 138; required sample: 300 per version). Guardrails: no-shows unchanged · average discount unchanged. Cost: within ceiling. Decision: [approve] [reject] [wait]
Annotated answer key

Card 1—reject (A stays). The written rule says zero tolerance and a guardrail worsened; the correct card answer is not to approve. Next, do what the clinic owner did: recognize that 6 versus 1 out of 300 is noise for a rare event, adjust complaint tolerance to 0.01 on the worksheet, and keep the hypothesis in memory as discarded, with the reason, so it can be retested under the right rule. Neither “approve anyway” nor “zero tolerance forever.”

Card 2—approve. B won by a margin, the sample reached its minimum, and no guardrail fell. Conversion did not move significantly—but it was monitored, not targeted, and did not worsen. The version becomes official; rollback is armed for the next cycle. That is what actually happened: it was promoted as v2.

Card 3—wait. 140 is less than half the required sample. A nine-point lead with this sample is “insufficient sample,” with no “but it is winning.” No guardrail fell, so there is no reason to stop: the test continues to 300 per version.

Answers 4 and 5 (acceptable example for the clinic). Ceiling: R$ 50 per cycle, 1 hour of the owner’s time per week. Starting level: L1. Card response deadline: 7 days; no response means nothing is promoted. Move to L2 after 5 L1 approvals with no rollbacks. Your worksheet may have different numbers—that is fine; each line must have a number.

You have just decided on two real cards and one constructed card—and written your loop’s ceiling and approver as numbers.

Summary

  • The decision card is the only attention the loop asks for: five lines, at most once a week, read from top to bottom before deciding.
  • There are only three answers—approve, reject, wait—and silence means “nothing promoted, pending.”
  • In clinic cycle 0002, the system rejected a ten-point gain because a guardrail had zero tolerance; the rule was applied correctly, but the tolerance needed to be calculated.
  • The ceiling (answer 4) is what the Guardian uses to stop a cycle; the approval level (answer 5) starts at L1 and moves to L2 only after five correct approvals.
  • Rollback applies at every level: a drop in the cycle after promotion restores the previous version, with the reason in the history.

Your next step

You can now decide on a real card without being swayed by the size of the gain—and have written down your loop’s ceiling and approver.

Over the next 15 minutes: open the loop worksheet you have been building and write your practice answers on lines 4 and 5, using numbers. Those are the only two left before the worksheet is complete in Track 5.

In the final lesson of this track, about honesty, you will learn what the loop guarantees by design, what it cannot guarantee at all, and how to recognize when someone sells the latter as if it were the former.

Track 04 · Lesson 5

What LOOP-R
does not guarantee

By the end of this lesson, you will be able to state, in one sentence each, the four things the loop guarantees by design and the four it does not—and recognize a false promise when someone sells one.

This entire track has asked for evidence. It would be strange to end by making a promise without any. So this is the honesty lesson: what an improvement loop guarantees because it was built to, and what depends on your volume, your market, and luck. Knowing the difference keeps you from abandoning the loop in month three out of frustration—and from trusting it beyond its limits.

↓ scroll to study

01 Four guarantees, and none of them is “the number goes up”

No learning loop guarantees that a number will rise. Not even this one—the beauty clinic owner heard that on the product’s first screen, before filling in the worksheet. LOOP-R guarantees four things, all by design: they hold because the rules from earlier lessons exist, not because someone promised. A worse version never replaces the current one by system decision. Every change is recorded and reversible. The cycle always runs the same way. Cost never exceeds the ceiling.

Beside each guarantee is what it does not cover. Not replacing a version with a worse one does not mean the next will be better. Having a record does not mean it contains anything useful—it may say “no evidence” for ten cycles in a row. Always running does not mean every cycle produces a good hypothesis. Staying under the ceiling does not mean the cost is worthwhile.

The rule this product follows in every text: the word “guarantee” appears only alongside non-regression, a reversible record, or a consistent cycle. If you read “guarantees more sales” in any material about AI loops—including this one—you are reading a promise no one can make.

Before—the promise being sold

“The AI will learn from every proposal and improve your sales.” No number, no condition, no timeline.

After—what is guaranteed

For the clinic: no version with worse replies or margin worse becomes official; every change has a date, metric, and rollback; the cycle runs every Monday; spending stops at R$ 50.

Takeaway: four statements that can be checked in the history, instead of one that cannot be verified anywhere.

02 The link still unproven: signal → revenue

In lesson 2, you replaced revenue with a signal as the test target because revenue takes too long. That has a cost worth stating: raising the signal does not prove that revenue rises. It is plausible—people who reply are more likely to buy—but plausible is not proven. Until the volume is there, the link between them is an explicit bet, not a result.

The clinic’s real case shows this plainly. In cycle 0004, the reply rate rose from 19.3% to 29.3%, and the version was rightly promoted. Sales conversion went from 3.0% to 3.7%—a figure that, with 300 proposals on each side, cannot be distinguished from luck. The cycle report says exactly that: “the signal → target link remains unproven.” Ten points more replies, and still nothing to claim about sales.

The real estate agent faces the same link under different names: more booked visits have not yet proved more closed sales. His loop optimizes bookings and monitors closings. Once there are 1,500 leads on each side, closings may become the target. Before then, saying “the loop increased sales” would be the same unsupported promise.

Test yourself

Clinic cycle 0004: replies 19.3% → 29.3% (N = 300 vs. 300, significant); conversion 3.0% → 3.7% (not significant). What can you claim?

03 “No evidence” for ten cycles in a row is a possible result

The loop guarantees that a cycle runs every week. It does not guarantee that the week brings a good idea. It is possible—and common in a small business—for the Critic to find only weak observations for several cycles, for the Optimizer to have nothing to propose, and for the history to accumulate “insufficient evidence” line after line. That does not mean the loop is broken. It is telling the truth about its volume.

Two things still matter during those weeks. The memory of discarded hypotheses grows and is as valuable as the memory of promoted ones: it prevents you or a future teammate from testing the same bad idea twice. The Meta-agent, which only observes and reports, tracks cost per promoted hypothesis. If nothing is promoted in five cycles, it suggests running less often instead of pretending to make progress.

The support manager went through six such cycles between March and April: too few tickets per type for the Critic to identify a pattern. Following the Meta-agent’s suggestion, she reduced the loop to every two weeks and let the log spreadsheet accumulate data. In June, with 30 tickets in the most common category, the first moderate hypothesis appeared. Her history has six “no evidence” entries—and she considers all six honest.

04 What is still research, and what is the bottleneck

One part of the LOOP-R idea is not ready for any business: the loop redesigning itself—level L3, where the Meta-agent proposes replacing assistants, changing the cycle order, or rewriting its own rules. That is research. In the product taught in this course, the Meta-agent only reports: cost, hypotheses generated versus promoted, who gets the most Guardian vetoes, and whether the judge is calibrated. No action. Anyone promising a system that “improves itself” is selling something that does not yet exist safely.

And there is a bottleneck no assistant can solve: the log spreadsheet. Without evidence, everything else is theater—and your business’s evidence is in WhatsApp, the reception notebook, or the salesperson’s head. The product connects to one thing: a spreadsheet with columns defined in the worksheet. You fill it in or export data. Each direct connection to another system is its own project; none comes free with the five answers.

For both the clinic and the real estate agent, the conclusion is the same—and it tests any promise: LOOP-R as a process discipline—logging, tests with sufficient sample sizes, and non-regression—has value on its own today. LOOP-R as a “self-learning system” depends on a volume of data most small businesses do not have. An honest product says so on its first screen.

The loop does not promise the number will rise. It promises it will not fall because of a system decision.

Practice now 0/3 complete

True or false: eight claims about guarantees

Mark each statement T or F. For false ones, write one line naming the real guarantee it stretches. About 8 minutes. This is analysis: you cannot “practice” a guarantee; you learn to recognize it when reading sales proposals.

Nothing changes in your business until you approve it: this is reading and marking. Mistakes help—every false statement you mark true is a promise someone may try to sell you.

  • 1 After four cycles, the reply rate must have risen.
  • 2 A worse version never replaces the current one by system decision.
  • 3 Every change can be reversed, and the history records when, with which numbers, and why.
  • 4 If the card is not answered by the deadline, the system promotes by default so the loop does not stall.
  • 5 A cycle’s cost never exceeds the ceiling you wrote on the worksheet.
  • 6 If the cost stayed under the ceiling, the loop was worthwhile.
  • 7 Ten cycles in a row with “no evidence” mean the loop is broken.
  • 8 The cycle runs the same way every week, even if no one has a good hypothesis.
Annotated answer key

1 — F. It stretches non-regression: the loop guarantees the number will not fall by its decision, not that it will rise.

2 — T. This is the core guarantee: the Evaluator promotes only when margin and guardrails remain intact; otherwise, “A stays.”

3 — T. A record with a rollback button. Every promotion and rollback appears in the history with its date, numbers, and reason.

4 — F. It stretches the cycle’s consistency: the loop keeps running, but the default when there is no answer is not to promote and to record “pending.”

5 — T. The Guardian stops the cycle when it reaches the ceiling.

6 — F. It stretches the ceiling: staying within the limit guarantees the spend, not the return. Cost per promoted hypothesis is a separate calculation, reported by the Meta-agent.

7 — F. It stretches consistency: running every week is guaranteed; producing a good hypothesis each week is not. Ten “no evidence” results are honest, and discarded ideas remain in memory.

8 — T. Cycle consistency. It is the cheapest guarantee and the best protection against “we stopped doing it in month 3.”

You have just separated, claim by claim, what the loop can support from what no one can support—the same reading you will apply to the next promise you hear.

Summary

  • By design, the loop guarantees four things: it does not select a worse version, everything is recorded and reversible, the cycle always runs the same way, and cost stays under the ceiling.
  • Each guarantee has a matching limit: none says the next version will be better, that the record will be useful, that each cycle brings a good idea, or that the cost is worthwhile.
  • A higher signal does not prove higher revenue; in the real case, replies rose by ten points while conversion remains indistinguishable from luck.
  • A loop that redesigns itself is still research, and the log spreadsheet is the bottleneck no assistant can solve for you.

Your next step

You can now look at an AI promise and identify, using the table, which of the four guarantees it stretches.

Over the next 15 minutes: reread the loop worksheet you built in earlier tracks and cross out any sentence promising a higher number. Replace it with one of the four guarantees and a number from your worksheet (“the system will not let the reply rate fall below 20%”).

In Track 5, you will follow the clinic’s four real cycles from start to finish, predict the fifth, and submit your completed worksheet with the first decision card filled in.