LOOP-R turns a business process into a loop: execute, measure, critique, propose, test — and only what is proven becomes the new version. Everything is recorded, with a rollback button.

Today AI works like this: request → response → done. The next proposal starts from scratch. LOOP-R closes the loop: every completed task informs the next version of how the work gets done.
Each run becomes a row in the spreadsheet. The system states only what the numbers support — and says when there is not enough data to conclude anything.
Executor, Observer, Critic, Optimizer, Guardian, Experimenter, Evaluator, Memory, and Meta-agent. None evaluates its own work — this separation keeps the system from convincing itself.
Every process version stays in the history. Promotion is one command; rollback is another. The system never replaces the current version with a worse one.
Do not ask AI to improve. Make every run produce evidence, every piece of evidence generate a hypothesis, and every hypothesis become an experiment — only proven results enter the next version.
Locate what to improve, operate, observe the result, propose a hypothesis, reinforce what won. Then repeat. Inside, an eight-step cycle assigns one assistant to each step.
They read and count the spreadsheet entries, then report what worked and failed — with evidence strength marked as strong, moderate, or weak.
One writes testable hypotheses (IF… THEN… BECAUSE…). The other vetoes anything that violates what you defined as untouchable. The veto is final.
They design an A/B test with the sample size needed for evidence and return one of three outcomes: B won, A stays, or the sample is insufficient.
They record everything — including failed ideas, so they are not tested again — and monitor the loop’s own cost.
| Guarantees | Does not guarantee |
|---|---|
| No regression — a worse version never enters production | that the number will rise |
| Auditability — every version, test, and decision is recorded; rollback takes one command | that every cycle produces a good hypothesis |
| Consistency — the cycle runs the same way every time | that the cost is worthwhile without your review |
| Cost ceiling — stops when reached | — |
In version 0.1, the loop runs inside Claude Code, and data comes from a spreadsheet exported as CSV — one row per run. No server, database, or integration.
The runner is a skill; the nine assistants are text files. If you have Claude Code, clone the repository and run it.
# install / check claude --version
It provides history and rollback: each process version is a commit; promotion and rollback use Git.
git --versionWhere you already record the result of each proposal, support ticket, or post. Run the command iniciar to generate the column schema for you to fill in or export.
# one row per run
id,data,versao,variante,resposta,resultado,...Five plain-language answers set up the whole loop. After that, one command each week.
The repository includes the runner, nine assistants, and a complete example that was actually run.
git clone https://github.com/inematds/loop-r && cd loop-r claude
These are the only things a system cannot safely infer. LOOP-R proposes the rest, and you confirm.
/loop-r iniciar # 1. Which process, and what number should move from what value to what value? # 2. Where is each run’s result recorded? # 3. What must AI NEVER change on its own? What must not get worse? # 4. How much can each cycle spend? Weekly or monthly? # 5. Do you approve every change (L1), or only want to see proposals (L0)?
Before creating anything, the system calculates how many runs your test needs — and says whether it should first optimize a faster metric.
"With 50 proposals per week, proving 3% → 5% takes ~60 weeks. I will optimize the RESPONSE RATE first (~12 weeks) and monitor conversion. Confirm?"
The nine assistants run in sequence and leave a manifest of what each read, decided, and why. A cycle without a complete manifest does not count.
/loop-r ciclo # → ciclos/0001/ (evidence, critique, hypotheses, veto, experiment, verdict)
When version B wins by a margin without worsening what you monitor, you receive a decision card. With no response in seven days, the default is not to promote.
CYCLE 0004 — hypothesis: "open with something specific from the review"
Result: response rate 19.3% → 29.3% (N = 300 vs 300)
Monitored: margin ok · complaints ok · opt-out ok
Cost: ~R$ 45 this cycle (R$ 50 ceiling)
Decision: approve / reject / wait for more data
/loop-r decidir
Promotion creates a new version and points production to it in a commit. Rollback points back. The old version is never deleted: it is part of the memory.
/loop-r promover # versoes/v2 becomes official · commit "promote v2" /loop-r reverter # rolls back to v1 · ledger entry with the reason /loop-r status
Aesthetics clinic, WhatsApp proposals, 1,500 synthetic runs. The assistants really ran; the data did not — it was generated to reflect a real effect, and the README in the example explains exactly what is real and what is simulated. What happened teaches more than a perfect case would.
The Optimizer proposed closing with “would you rather start this week or next?” The Guardian vetoed it: a forced choice creates pressure and complaints. The second hypothesis (a message of up to 80 words) went to testing — after the Experimenter calculated that the original test would take 40 weeks.
The response rate rose from 19% to 30% (N=300/300). But complaints were 6 versus 1 with zero tolerance, so the Evaluator returned A stays. The rule was wrong (6 vs. 1 out of 300 is noise), not the system: a person adjusted the tolerance for later tests, and the hypothesis was recorded as discarded.
A hypothesis approved in the first cycle waited in the queue for thirteen weeks — only one test can run at a time. The Meta-agent pointed out the cost of waiting and suggested an automatic trigger. It went onto the roadmap.
Open the proposal by mentioning something specific from the customer’s review: responses rose from 19.3% to 29.3%, with guardrails within tolerance. Decision card, approval, versoes/v2. The loop’s first proven learning.
# memoria/ledger.md — excerpt (one line per event, never deleted)
| 2026-02-02 | ciclo 0001 | H1 (fechar com pergunta de escolha forçada) | — | — | — | VETADA |
| 2026-04-27 | ciclo 0002 | H2 (teto de 80 palavras, E0001) | A_SEGUE — resposta A=0,1933 vs B=0,2967, p=0,0033 | reclamacoes Δ+0,0167 tol 0,0 → FALHA | aceito | DESCARTADA |
| 2026-05-04 | ciclo 0003 | H4 (valor logo após a saudação) | — | — | — | VETADA |
| 2026-05-04 | ciclo 0003 | H3 (abrir citando a avaliação, E0002) | AMOSTRA_INSUFICIENTE — N=0/212 | — | — | EM TESTE |
| 2026-07-27 | ciclo 0004 | H3 (abrir citando a avaliação, E0002) | B_GANHOU — resposta A=0,1933 vs B=0,2933, p=0,0043 | todas OK | aprovado | PROMOVIDA v2 |
Without evidence, everything else is theater. Each phase delivers something that works on its own, and the next starts only after a real user has three cycles in the history.