INEMA.CLUBPROLOOP-R PT · EN · ES

LOOP-R · Track 01 · Diagnosis

Why your business does not improve

Learn to identify where AI does work but learns nothing in your own business. Leave with a rework count, your maturity level marked, and a vague instruction rewritten so its result can be measured.

Track 01 · Lesson 1 · foundation

AI that
ends there

By the end of this lesson, you will be able to look at a task AI does for you today and explain in one sentence why it ends there — and what is missing for the next attempt to start from a better place.

You have already asked AI for hundreds of proposals, replies, and pieces of copy. Each was good enough. And yet today’s proposal starts in exactly the same place as the first: nothing that worked was saved. This lesson names the problem — because without a name, it feels normal.

↓ scroll to study

01 Person, request, response, done — and tomorrow, it starts all over

In most businesses today, AI works in a straight line. Someone writes a request — people in the field call it a prompt —, the AI returns an answer, the person adjusts what is needed, and sends it. Done. Next time starts from the same point: a new request, a blank page, no memory of the last one.

Regina owns an aesthetics clinic and sends about 200 proposals a month over WhatsApp. For each one, she writes something like “draft a skincare package proposal for the customer who asked about the price.” The proposal comes back, she changes two sentences, and sends it. The customer replies or disappears. Proposal 201 starts with no knowledge of the previous 200 — neither which got a reply nor which led to a sale.

The AI did not do a bad job. Every proposal was good. The problem is different: the path has four stops and no way back.

Today, at the clinic

200 proposals a month, each one from scratch. Regina does not know which of the 200 worked or why — and tomorrow’s proposal will not know either.

With a loop

The same 200 proposals, each becoming a row in a log. One decision a week, on a five-line card.

What the LOOP-R reference example shows (simulated data, fictional clinic): across 4 cycles, 1 proven change with 300 sends per version, 1 idea discarded by a rule, 0 worse versions in production.

02 AI remembers nothing from one request to the next — and that is not a flaw

An AI chat is like an excellent temporary assistant who arrives new for every shift: it writes very well and knows the subject, but it was not here yesterday. It did not see whether the customer replied, the lead booked a visit, or the support ticket came back. Everything that happens afterward the response stays outside the conversation.

Marcos, a real estate agent, gets about 40 leads a week through a listing portal — a lead is someone who asked for information but has not decided yet. He replies to each one with AI’s help. When a lead responds well because Marcos mentioned the balcony’s square footage, that teaches something — but only Marcos, if he notices. To the AI, the next lead is the first one it has ever seen.

That is why “using it a lot” does not become “learning.” More use changes nothing until someone records the result and carries it into the next request. The rest of this course builds that feedback loop.

Check your understanding

Marcos used AI 160 times this month to reply to leads. What does it know now about which replies worked?

03 Real improvement takes six things — and none of them is “effort”

When someone says “I want AI to improve,” they usually mean writing a better request. But real improvement does not come from the request. It comes from six ingredients connected in a loop: evidence (what happened, written down), metric (a number that says whether it worked), memory (where past attempts are kept), hypothesis (an idea for a change, with a reason), experiment (the idea tested against the current version), and evaluation (someone determines which one won, based on the number).

Paulo manages support at a small software company: 10,000 tickets a year. For his team, the six ingredients look like this. Evidence: each ticket has its topic recorded. Metric: resolved on first contact, yes or no. Memory: a place where the team records what it has already tried in its canned replies. Hypothesis: “if the password reply includes a direct link, fewer people will come back.” Experiment: half of password tickets get the new reply, half the old one. Evaluation: the count of tickets that “came back within 7 days” determines the winner, not the opinion of the person who wrote the reply.

Notice: none of the six means “ask more forcefully.” If one is missing, the loop does not close — a hypothesis without an experiment is a guess; an experiment without memory is a discovery lost the next week.

04 What a loop guarantees — and what it does not promise

Before we continue, here is the anti-promise, stated in the very first lesson so it cannot be hidden: no improvement loop guarantees that the number will rise. Anyone promising that is selling luck dressed up as a method. A well-built loop guarantees something else, more modest and more valuable:

GuaranteesDoes not guarantee
A worse version will not replace the current one on its own.That the next version will be better.
Every change will be recorded and there will be a rollback button.That the log will contain something useful every week.
The cycle will run every week, in the same way.That every cycle will produce a good idea.
The cost will not exceed the limit.That the cost will be worthwhile.

Notice that all four guarantees are about process, not results. That may sound modest. It is exactly what no promise of “AI that improves itself” can honestly deliver — and it lets you sleep while the cycle runs: nothing enters without evidence, and anything that entered can be removed.

In the LOOP-R reference example — a fictional clinic like Regina’s, using simulated data — the proposal response rate rose from 19.3% to 29.3% in a test with 300 sends per version. Sales conversion did not move: 3.7% versus 3.0%, a difference that could be chance. The loop did not hide that. It recorded it, and the person made the decision with the numbers in front of them. For Regina, that is what matters: not a promised number, but one she can verify.

No loop guarantees that the number will rise. It guarantees that a worse version will not go live on its own — and that you can always roll back.

Practice now 0/3 feito

Identify the tasks that end there — in Regina’s case and in your business

Read a short case, answer three questions, and compare your answers with the key; then list five tasks of your own. ~10 min.

This is only reading and note-taking — nothing changes in your business until you approve it, and no customer receives anything. If your answer differs from the key, note the difference: it is exactly the kind of thing worth discussing with a business partner.

The case. Regina, who owns an aesthetics clinic, uses AI for five things each week: (1) writing package proposals on WhatsApp; (2) answering “how much does it cost?” from people who find her on Instagram; (3) drafting Monday’s post caption; (4) summarizing customer reviews on Google; (5) writing appointment reminders. None of the five tracks what happens afterward. Regina thinks task 4 “already counts as learning” because AI reads what customers said.

Answer on paper or on your phone:

  1. Which of the five tasks end there?
  2. Is task 4 a loop, as Regina thinks? Why or why not?
  3. If you were Regina, which task would you start tracking — and what would you write in each row?
Answer key with explanations (open only after answering)

1. All five. They all follow the path request → response → done: none has a row saying what happened afterward (was the proposal answered? Did the post bring in an appointment? Did the customer show up after the reminder?).

2. No. Reading reviews is a request that ends there, just like the others. It would become a loop if what AI read became a hypothesis (“customers complain about the wait: if the proposal states the session length, will more of them reply?”), were tested against the current version, and the result were recorded. Reading is not learning; changing based on evidence is.

3. The best choice is the most frequent task with an easy-to-see result: proposals (1), with 200 a month and replies arriving within hours. Each row could include: date, proposal version, package type, replied (yes/no), booked (yes/no), complained (yes/no). Task 3 (the post) could work too, but the result takes longer and is harder to connect to the wording.

You can now identify, in one sentence each, the tasks that end there — in Regina’s case and in your business.

Summary

  • Today AI works in a straight line: someone asks, it answers, and that is the end. The next turn starts from the same point.
  • That is not a flaw in the AI; it is the lack of a record of what happened after the reply.
  • Improvement requires six ingredients connected in a loop — and none of them is “asking more forcefully.”
  • An honest loop does not promise a higher number; it promises that things will not get worse without your knowledge, and that there is a rollback button.

Your next step

You can now distinguish “uses AI” from “learns with AI” — and you have seen that almost everything today is the first one.

In the next 15 minutes, open yesterday’s WhatsApp messages or email and mark three replies AI helped write. Next to each one, write: “Do I know what happened afterward? yes / no.”

In the next lesson, you will calculate the cost of starting from scratch — and find out why handling 10,000 support tickets may teach you nothing.

Track 01 · Lesson 2 · foundation

Ten thousand executions and
no learning

By the end of this lesson, you’ll be able to put a number on how many times per month the same task is redone from scratch in your business—and how many of those could become a line in a log.

Volume creates a sense of experience. Ten thousand support tickets handled can feel like ten thousand lessons. But if nobody noted what they were about, those ten thousand are worth no more than the first one. You pay the cost of repeating without getting the benefit of learning—and that bill never shows up on an invoice.

↓ scroll to study

01 Ten thousand tickets, one issue in three—and nobody noticed

Paulo, support manager at a small software company, ends the year with 10,000 tickets handled. Each one was resolved well, one at a time, with AI helping write the reply. The team gets praised. And nobody knows that 37% of those tickets—3,700—are about the same issue: a password screen that confuses everyone.

Why did nobody notice? Because each ticket is a request that ends right there. From inside one ticket, you can’t see the pile. Only someone looking at the whole set—and, for that, needing a “what was it about?” column filled in for each one—can see that one in three points to the same thing.

Here’s the difference between “we handled a lot” and “we learned a lot.” The 10,000 resolved tickets left 10,000 customers served and zero product changes. A single fix to that screen would have prevented 3,700 conversations. The information was there all year; it just wasn’t written down.

02 Starting over has a cost that never appears on the invoice

Marcos, the real estate agent, replies to about 40 leads a week—around 160 a month. Each reply takes about six minutes to request, adjust, and send. That’s 16 hours a month. He doesn’t think it’s expensive because AI “does it fast.” But next month’s 16 hours will buy exactly the same thing as this month’s: 160 isolated replies, none smarter than the last.

Compare that with another scenario using the same amount of time: the same 160 replies, but each leaves a row—date, property type, opening used, whether the lead replied, whether they booked a visit. The effort per reply is the same. The difference is the trail. In one case, all that’s left is fatigue; in the other, there’s material to answer a question Marcos can’t answer today: “Which opening gets more people to book?”

The cost of starting over isn’t the time—you’d spend that time anyway. It’s the learning that passed through your hands and didn’t stay.

Volume without a record isn’t experience. It’s the same first time, repeated ten thousand times.

03 Two companies use the same AI: the difference is who learns faster

For years, a company’s advantage was “having more data than the competition.” That changed. Today, everyone has access to the same AI, with the same capabilities. If two real estate agencies use the same chat, the same listing sites, and serve the same neighborhood, what sets them apart? It isn’t a better-written prompt—you can copy one in a day.

What can’t be copied is the history. One agency records the channel, property type, reply sent, and outcome for each lead. After six months, it has about 1,000 rows to consult before changing anything. The other has 1,000 conversations lost in each agent’s WhatsApp. A competitor can copy the text; they can’t copy six months of “what we’ve already tested and how it went.”

That accumulated history is the asset. It doesn’t come from better technology—it comes from one row per execution, every day, without exception.

Test yourself

Two clinics use the same AI, charge the same prices, and are on the same street. What could one have that the other can’t copy in a month?

04 Where evidence begins: one row per execution

The whole loop depends on one humble step: recording. The place where it happens is the tracking spreadsheet —a regular spreadsheet, with one row for every proposal, reply, or ticket, and a column most people forget: what happened next. Without that column, you have a file of text, not evidence.

For Paulo’s support team, a row looks like this: date · ticket topic · which saved reply was used · resolved on first contact (yes/no) · came back within 7 days (yes/no). Five columns. The question that used to take a week of reading (“what do people complain about most?”) becomes a two-minute count.

One honest point about the method: the evidence source is the bottleneck, not the AI. If the data live in a salesperson’s head or are scattered across conversations, the loop has nothing to read—and a loop without evidence is theater. One caution from the start: 12 rows don’t make a pattern. “Tuesday’s proposals get more replies,” based on 12 proposals, is chance dressed up as a discovery. The loop’s rule is to claim a pattern only with 30 or more items per group.

Paulo’s support team today

10,000 tickets a year, resolved one by one. Answering “what do people complain about most?” takes a week of reading—and nobody has that week.

With the tracking spreadsheet

10,000 rows with the “topic” column filled in. The same question is answered with a count: 37% are about one issue.

Result: same 10,000 replies, same effort per ticket, one extra column—and 3,700 tickets now point to a single product fix.

Practice now 0/4 done

Count how many times the same task was redone from scratch last month

Leave with one written sentence and two numbers: how many times per month the task happens, and how many times it went unrecorded. ~10 min.

This is only a count—nothing in your business changes until you approve it; you won’t alter any reply or tell anyone. If counting everything feels like too much, estimate one week and multiply by four: an approximate scale is enough for this lesson.

You’ve just put a monthly number to a feeling—and that number is the first line of your business diagnosis.

Recap

  • Ten thousand tickets resolved one by one teach as much as the first one if nobody noted what they were about.
  • The cost of starting over is invisible on the invoice and only shows up when you count: the same task, dozens of times a month, with no trail.
  • With the same AI available to everyone, the hard-to-copy company is the one that builds a history of what each execution produced.
  • Evidence begins in a simple spreadsheet, with a column for what happened next—and only becomes a pattern from 30 per group onward.

Your next step

You’ve just turned “we do this all the time” into a monthly number, with the unrecorded part counted.

In the next 15 minutes, create a spreadsheet with five columns—date · what was sent · to whom · what happened · notes—and fill in the last three times the task occurred.

In the next lesson, you’ll find out which of the five maturity levels your business is at today—and why jumping straight from 1 to 5 never works.

Track 01 · Lesson 3 · foundation

The five levels: from “responds” to
a a learning company

By the end of this lesson, you’ll be able to identify which of the five maturity levels your business is at today using observable signs, not impressions—and the only next level that makes sense.

Everyone who uses AI thinks they’re “advanced”—and almost everyone is at level 1. Without a yardstick, you buy a level-4 tool to solve a level-1 problem and pay for theater. This lesson’s yardstick is here to keep you from wasting that money.

↓ scroll to study

01 Level 1 — Responds: you ask, it delivers, you check

Level 1 is where almost every business is, and there’s nothing wrong with that. It’s AI answering a request. You write, it delivers text, you check and use it. All the process intelligence lives in your head: what to ask, what to adjust, what to send.

Regina, at the clinic, is a pure level 1. The signs are easy to spot: each use starts with her writing; the result of each proposal lives in her memory, nowhere else; if she goes on vacation, the AI “forgets” the clinic—because it never knew anything beyond that day’s request.

Level 1 isn’t bad. Getting a good answer a thousand times already saves many hours. Just don’t confuse that with learning: a thousand good answers are still a thousand first times.

Giving a good answer a thousand times is level 1. Level 1 isn’t a delay; it’s where almost everyone starts.

02 Level 2 — Executes a goal: the single-purpose assistant

At level 2 comes what practitioners call an agent—in this course, a single-purpose assistant. You no longer dictate each step; you give it a goal. It plans, uses a tool (a calendar, a spreadsheet, WhatsApp), and delivers the result.

Marcos, the real estate agent, has one: the assistant reads the lead that came through the listing site, finds the property in his spreadsheet, writes a reply, and suggests two times for a viewing. Marcos just checks and approves. It does more than reply—it executes a goal from start to finish.

Here’s the catch: the assistant does more things and learns exactly as much as at level 1—nothing. Nobody records whether the viewing happened. Level 2 expands reach; it doesn’t close the loop.

Test yourself

A support assistant reads a ticket, finds the right saved reply, and answers on its own—without anyone noting whether the customer came back. What level is this?

03 Level 3 — The loop: execute, measure, evaluate, improve, execute again

Level 3 is the first one where the company learns. Each execution becomes a row; every week, someone (or an assistant with that role) reviews the rows, proposes a hypothesis, tests the new version against the current one, and decides based on the numbers. If it wins with proof, it becomes the official version. If not, the result is recorded so no one unknowingly tries it again.

In LOOP-R’s reference example (a fictional clinic with simulated data), the loop ran four cycles between February and July. One hypothesis—“open by mentioning the review the client left”—was tested with 300 proposals on each side: response rate went from 19.3% to 29.3%, without worsening anything being monitored (complaints, discount); it was approved on the card and promoted to the official version. Another, “up to 80 words,” rose by almost as much and was discarded by rule. Both remain in memory.

Within level 3 there are three degrees of autonomy. L0the loop only suggests; you do everything. L1you approve every change on a five-line card, at most once a week. L2it promotes on its own, alerts you, and keeps the rollback button—and unlocks only after you correctly approve five changes at L1. This course sets yours up at L1.

Level 1 — Regina today

200 proposals a month. “Did the AI improve?” Nobody knows, and there’s no way to answer the question.

Level 3 — with the loop

The 200 proposals become 200 rows. Each cycle, one hypothesis is tested in two versions, and Regina gets a card to decide.

Reference example result: over 4 cycles, 2 hypotheses tested with 300 sends per version, 1 approved and promoted, 1 discarded—and no change went live without an approved card.

04 Levels 4 and 5—and why your only next step is level 3

Level 4 is the system that observes its own improvement system: the cost of each cycle, how many hypotheses became versions, which assistant gets vetoed most. It sounds like magic, and honesty matters here: today, in the product, this level only reports. It doesn’t redesign anything on its own. A system that changes itself with little data multiplies every error from the levels below—which is why autonomous redesign remains research, not a product.

Level 5 is the whole company: every department has its loop, with another loop above them connecting the lessons. Paulo’s support team shows the path: tickets logged → cause found (the password screen) → new reply tested against the old one → help text updated → product change suggested. Support stops just answering and starts improving the product. But that only works because, underneath, every ticket became a row.

That leads to the rule: you don’t skip a step. Someone at level 1 who buys “level 5” is buying theater, because the loop above has no evidence to learn from. From levels 1 and 2, the only useful next step is level 3—and level 3 starts with one more column in the spreadsheet and a test with numbers.

Practice now 0/3 done

Maturity self-check: identify the level of three businesses—and then your own

Read the case, mark each person’s level, check the answer key, then mark the signs that apply to your business and find your level. ~12 min.

This is a paper diagnosis—nothing in your business changes until you approve it, and nobody needs to know the result. If you’re unsure between two levels, choose the lower one: the next step is the same, and being conservative costs nothing.

The case. Regina asks AI for each proposal and edits it by hand; the outcome stays in her head. Marcos has an assistant that reads the lead, checks the property spreadsheet, and suggests a viewing on its own; nobody records whether the viewing happened. Paulo logs each ticket with its topic and “resolved on first contact?”, tests a new reply against the old one every month, and decides on a card. All three say they’re “pretty advanced with AI.”

Answer:

  1. What level is each of the three at?
  2. Which one has the cheapest next step—and what is it?
  3. Now your business: mark the signs below that apply today—and find your level in the answer key.
  • Every AI use starts with someone writing the request, and the result stays in that person’s head.
  • An assistant carries out a goal end to end and uses a tool (calendar, spreadsheet, WhatsApp) on its own.
  • There is one row per execution, with a column for “what happened next.”
  • A new version has been tested against the old one using a number on both sides.
  • There is a recorded decision to approve or reject a change, with a reason.
  • Someone reviews the cost per cycle and how many hypotheses became versions.
  • More than one department has a loop, and what one learns reaches the others.
Annotated answer key (open only after marking)

1. Regina: level 1 (responds). Marcos: level 2 (executes a goal; no measurement). Paulo: level 3 (one row per execution, a test with numbers, recorded decision).

2. Regina and Marcos tie on cost: both need the same step—a result column for each execution (“did they reply?”, “did the viewing happen?”). Marcos has one advantage: his assistant already produces a standardized execution, so the row is easier to fill in. Paulo’s next step is different: keep the cycle consistent and monitor cost per promoted hypothesis.

3. Reading the signs. Only A: level 1. B without C: level 2. C + D + E: level 3. F on top of C-D-E: level 4 (reporting). G: level 5. If you marked C without D, you’re between levels 2 and 3—you have evidence but lack an experiment; this is the most common position for someone who started logging. If you marked B and G without C, be skeptical: without one row per execution, the “cross-department loop” is just an impression.

You’ve just put a number on your level using signs anyone on the team can check—and identified the only next step.

Recap

  • Level 1 responds; level 2 executes a goal; level 3 closes the loop with measurement and testing; levels 4 and 5 are still largely research.
  • You identify the level by signs anyone can verify—is there one row per execution? Was there a test with numbers? Is there a recorded decision?—never by the feeling that “we’re advanced.”
  • Almost every small business is at level 1 or 2, and that’s the normal starting point, not a delay.
  • The useful next step is always level 3, starting at L1: you approve every change on a card.

Your next step

You’ve just identified your level and the only next step—with signs, not impressions.

In the next 15 minutes, message your business partner (or yourself): “We’re at level N because [sign]. The next step is [a result column / a test with numbers].” One message, three lines.

In the next lesson, you’ll see the most common mistake when trying to reach level 3—writing “AI, get better”—and learn to replace it with an instruction you can measure.

Track 01 · Lesson 4 · foundation

“AI, get better”—the mistake
that sounds smart

By the end of this lesson, you’ll be able to take a vague instruction (“get better,” “be more persuasive”) and rewrite it with evidence, a numeric goal, a “without worsening” guardrail, and a testable hypothesis—using any AI chat.

It’s the first thing almost everyone does: open the assistant and write “analyze your work and improve every time.” It sounds like a smart command. But it has no number, no data, and no way to know whether it worked—and the assistant will say it did improve. This lesson replaces that sentence with a request you can verify.

↓ scroll to study

01 “Analyze your work and keep getting better” isn’t a system—it’s a sentence

Paulo from support wrote this in the assistant that replies to tickets: “analyze your replies and keep improving.” A month later, he asked whether it had improved. The assistant confidently said yes. There was no way to know if that was true—and that’s what makes the sentence dangerous: it creates the feeling of improvement without producing any.

Three questions the sentence doesn’t answer are missing. Better than what (which version?). Measured by whom (the assistant itself, which has every reason to say it improved?). Compared how (against what, and using how many cases?). Without recorded evidence, a number, and a previous version, “improved” is just the expected answer to the question you asked.

This isn’t an AI limitation; it’s a request limitation. The course’s anchor sentence exists to replace exactly this habit:

Don’t ask AI to improve. Make every execution produce evidence, every piece of evidence generate a hypothesis, and every hypothesis become an experiment—and only what is proven goes into the next version.

02 A goal needs a number, a direction, and “without worsening”

The first change is to replace “improve” with “from X to Y, without worsening Z”. For Regina: “proposal response rate from 20% to 30%, without worsening complaints or average discount.” For Marcos: “viewing booking rate from 20% to 30%, without worsening no-show visits.” The “without worsening” part has a name— guardrail —and it’s not decoration.

Without a guardrail, every goal becomes a target on its own, and the assistant learns to cheat: ask for “more clicks” and it writes clickbait; ask for “less time per ticket” and Paulo’s team starts closing tickets quickly without resolving them; ask for “more replies” and messages start provoking “stop sending me this”—which counts as a reply. A goal without “without worsening” is this track’s quietest mistake, because the number goes up while the business gets worse.

Two cautions about the number. First, choose a metric customers respond to quickly—reply rate, bookings—and monitor a money metric (conversion, margin) as a guardrail. Second, goals near 3% are almost impossible to prove in a small business: proving 3% → 5% takes about 1,500 executions on each side, which at 200 a month takes 15 months. Proving 20% → 30% takes ~300 on each side. Start where proof is possible.

03 A hypothesis is a sentence with IF, THEN, and BECAUSE—and only one change

The second change is to replace “idea” with a hypothesis. A hypothesis has a fixed form: IF [one change] THEN [expected effect on the goal] BECAUSE [reason]. In LOOP-R’s reference example, the fictional clinic wrote: “IF the proposal is up to 80 words, THEN the reply rate rises from 19% to about 30%, BECAUSE long WhatsApp proposals aren’t read to the end. Change: length only.”

The hardest part is the last one: one change only. Marcos wants to “cut the text, change the tone, and send it in the morning”—three changes at once. If it works, he won’t know which one worked; if it fails, he won’t know which one hurt. Three ideas are three hypotheses, tested one at a time or recorded for later.

For Regina, the same form applies: “IF the proposal opens by mentioning the review the client left, THEN more clients reply, BECAUSE they recognize the message was written for them”—that was, in fact, the hypothesis that was promoted in the example. A good hypothesis also protects you: if it fails, you know exactly what to discard and what to record so nobody unknowingly tests it again.

Notice what happened to Paulo’s vague instruction after these two changes:

Vague instruction

“Assistant, answer tickets better and keep improving over time.”

Verifiable request

“Log each ticket (topic, reply used, resolved on 1st contact?). Goal: resolved on 1st contact from 50% to 65%, without worsening 7-day reopenings. Hypothesis: IF the password reply includes the direct link THEN fewer people come back BECAUSE customers get lost on the screen. Test: half with each version, ~170 per version.”

Result: the first produces nothing verifiable; the second produces a record, a target number, a guardrail, and a test with a sample size—four things to discuss with a business partner.

04 The person who decides whether it worked isn’t the person who came up with the idea

The third change is to separate the proposer from the judge. The judge can only answer one of three things: B won by a clear margin, A stays, or not enough data—keep testing. The third answer is the most common and must be acceptable, because the alternative is promoting chance. That “300 on each side” has a name— sample size that proves it —and it’s calculated before the test, not after.

This is where the second quiet mistake lives: stopping a test halfway because the new version is “in the lead.” Marcos, with 40 leads a week, needs about 15 weeks to reach 300 per version. In week 2, the new version has 12 replies out of 40 versus 7 out of 40 for the old one, and he wants to promote it. Twelve versus seven out of 40 is exactly the kind of difference that appears and disappears on its own. Stop early for one reason only: a guardrail is falling. Never because of a partial lead.

Test yourself

Week 2 of the test: new version gets 12 replies out of 40; old version gets 7 out of 40. What should the judge say?

05 “Looks better” isn’t “is better”: the verdict comes from the number and the guardrails

Back to the reference example. In cycle two, the “up to 80 words” hypothesis was tested with 300 proposals on each side: replies rose from 19.3% to 29.7%—ten points, with a proven sample size. And it wasn’t promoted. The “complaints” guardrail showed 6 versus 1, and the agreed tolerance was zero. The rule was applied as written; the version stayed out; everything was recorded.

Then comes the second lesson, as important as the first: the rule was poorly calibrated. Six complaints versus one out of 300 is noise—it can happen by chance. Zero tolerance for a rare event rejects good ideas because of bad luck. The team adjusted the tolerance to 1% in later tests, and the rejected hypothesis stayed in memory with the reason. In cycle four, “open by mentioning the client’s review” rose from 19.3% to 29.3%, stayed within the guardrails, and became the official version. Sales conversion didn’t move—and that was recorded too.

The same rule applies to Paulo: if “resolved on first contact” rises but “reopened within 7 days” rises too, A stays. The system doesn’t promote something because it looks better—and at L1, you’re the one who approves it on the card, with the numbers in front of you.

Common mistake

Promoting because it looks better. The new proposal “is clearly prettier,” the rate went up in the first week—and it becomes official. This happens because the human eye confuses taste with results, and a good week with a pattern. How to avoid it: nothing becomes official without a verdict based on the required sample size and checked guardrails; and the decision goes through your card, not the impression of the person who wrote it.

Practice now 0/4 done

Rewrite one of your vague instructions with evidence, a goal, and a hypothesis

Leave with a four-part text—evidence, a goal with “without worsening,” an IF-THEN-BECAUSE hypothesis, and a test—generated in an AI chat from a vague instruction in your business. ~10 min.

You’re only writing a request in a chat—nothing in your business changes until you approve it, and no customer reply goes out from here. If the answer includes numbers you didn’t provide, delete them and ask again, adding “don’t make up numbers.”

This is a ready-to-paste request for an AI chat: the parts between < > are the only ones you replace—the rest stays as is.

I have an AI assistant that does this task in my business: <describe the task—for example, reply to real estate portal leads>.
Today, the instruction I gave was vague: “<paste the vague instruction—for example, be more persuasive and keep improving>.”
The current number I know is: <for example, 20% of leads reply—or “I don’t know”>.

Rewrite the instruction in four parts, without making up numbers I haven’t given you:
1. EVIDENCE—what to record per execution, one row at a time, in up to 5 columns (one must be “what happened next”).
2. GOAL—in the format “from X to Y, without worsening Z,” using only the number I provided. If I said “I don’t know,” say the first task is to measure for 2 weeks and suggest what to measure.
3. HYPOTHESIS—one sentence: “IF [one change only] THEN [effect on the goal] BECAUSE [reason].”
4. TEST—how to split into two versions (current and new) and how many executions per version to collect before deciding (for goals between 20% and 30%, use about 300 per version).
End by listing what must NOT change without my approval.

What you should see: a response with four numbered sections; a GOAL with a “without worsening”; a HYPOTHESIS with just ONE change; a TEST with a sample size per version; and a short list of what’s out of scope. If the hypothesis contains two or three changes, reply: “separate this into one hypothesis per change.”

You’ve just turned a vague command into a request you can verify—with a record, a target number, a guardrail, and a test.

Recap

  • “Get better” is a sentence, not a system: without evidence, a number, and a comparison, the assistant will only say it improved.
  • A useful goal has a direction, a number, and “without worsening”—because every goal by itself teaches the assistant to cheat.
  • A hypothesis is IF-THEN-BECAUSE with one change only; the numbers judge, using the sample size agreed in advance—not the person who had the idea.
  • A version can rise by ten points and still be rejected—and that means the system is working, not failing.

Your next step

You’ve just turned a vague command into a request you can verify—the first draft of your loop worksheet.

In the next 15 minutes, take the GOAL AI wrote and check the current number against real data: count the last 20 executions and see how many succeeded. Adjust X in “from X to Y.”

In Track 2, you’ll build the LOOP-R step by step, starting with L—Locate: one process, one number—and fill in the first question on the loop worksheet with what you just wrote.