Framework · Self-improving systems and businesses

Your business does a thousand tasks with AI. Why does nothing change?

LOOP-R turns a business process into a loop: execute, measure, critique, propose, test — and only what is proven becomes the new version. Everything is recorded, with a rollback button.

LOOP-R — continuous improvement cycle with AI
What it is

It is not a better prompt. It is a system that learns from every run.

Today AI works like this: request → response → done. The next proposal starts from scratch. LOOP-R closes the loop: every completed task informs the next version of how the work gets done.

🧪 Evidence, not opinion

Each run becomes a row in the spreadsheet. The system states only what the numbers support — and says when there is not enough data to conclude anything.

👥 Nine assistants, one role each

Executor, Observer, Critic, Optimizer, Guardian, Experimenter, Evaluator, Memory, and Meta-agent. None evaluates its own work — this separation keeps the system from convincing itself.

↩️ Rollback button

Every process version stays in the history. Promotion is one command; rollback is another. The system never replaces the current version with a worse one.

Do not ask AI to improve. Make every run produce evidence, every piece of evidence generate a hypothesis, and every hypothesis become an experiment — only proven results enter the next version.

How it works

Five steps that form a loop

Locate what to improve, operate, observe the result, propose a hypothesis, reinforce what won. Then repeat. Inside, an eight-step cycle assigns one assistant to each step.

Locate→ Operate→ Observe→ Propose→ Reinforce→ ↻
1

Observer and Critic

They read and count the spreadsheet entries, then report what worked and failed — with evidence strength marked as strong, moderate, or weak.

2

Optimizer and Guardian

One writes testable hypotheses (IF… THEN… BECAUSE…). The other vetoes anything that violates what you defined as untouchable. The veto is final.

3

Experimenter and Evaluator

They design an A/B test with the sample size needed for evidence and return one of three outcomes: B won, A stays, or the sample is insufficient.

4

Memory and Meta-agent

They record everything — including failed ideas, so they are not tested again — and monitor the loop’s own cost.

What it guarantees — and what it does not

GuaranteesDoes not guarantee
No regression — a worse version never enters productionthat the number will rise
Auditability — every version, test, and decision is recorded; rollback takes one commandthat every cycle produces a good hypothesis
Consistency — the cycle runs the same way every timethat the cost is worthwhile without your review
Cost ceiling — stops when reached—
Prerequisites

A spreadsheet and Claude Code

In version 0.1, the loop runs inside Claude Code, and data comes from a spreadsheet exported as CSV — one row per run. No server, database, or integration.

Claude Code

The runner is a skill; the nine assistants are text files. If you have Claude Code, clone the repository and run it.

# install / check
claude --version

Git

It provides history and rollback: each process version is a commit; promotion and rollback use Git.

git --version

A process spreadsheet

Where you already record the result of each proposal, support ticket, or post. Run the command iniciar to generate the column schema for you to fill in or export.

# one row per run
id,data,versao,variante,resposta,resultado,...
User guide · step by step

From zero to your first cycle

Five plain-language answers set up the whole loop. After that, one command each week.

1

Clone and open in Claude Code

The repository includes the runner, nine assistants, and a complete example that was actually run.

git clone https://github.com/inematds/loop-r && cd loop-r
claude
2

Answer the five questions

These are the only things a system cannot safely infer. LOOP-R proposes the rest, and you confirm.

/loop-r iniciar
# 1. Which process, and what number should move from what value to what value?
# 2. Where is each run’s result recorded?
# 3. What must AI NEVER change on its own? What must not get worse?
# 4. How much can each cycle spend? Weekly or monthly?
# 5. Do you approve every change (L1), or only want to see proposals (L0)?
3

Read the honest estimate and confirm

Before creating anything, the system calculates how many runs your test needs — and says whether it should first optimize a faster metric.

"With 50 proposals per week, proving 3% → 5% takes ~60 weeks.
 I will optimize the RESPONSE RATE first (~12 weeks) and monitor conversion. Confirm?"
4

Fill in the spreadsheet and run a cycle

The nine assistants run in sequence and leave a manifest of what each read, decided, and why. A cycle without a complete manifest does not count.

/loop-r ciclo     # → ciclos/0001/ (evidence, critique, hypotheses, veto, experiment, verdict)
5

Decide with the five-line card

When version B wins by a margin without worsening what you monitor, you receive a decision card. With no response in seven days, the default is not to promote.

CYCLE 0004 — hypothesis: "open with something specific from the review"
Result:      response rate 19.3% → 29.3%  (N = 300 vs 300)
Monitored:   margin ok · complaints ok · opt-out ok
Cost:        ~R$ 45 this cycle (R$ 50 ceiling)
Decision:    approve / reject / wait for more data

/loop-r decidir
6

Promote — or roll back

Promotion creates a new version and points production to it in a commit. Rollback points back. The old version is never deleted: it is part of the memory.

/loop-r promover   # versoes/v2 becomes official · commit "promote v2"
/loop-r reverter   # rolls back to v1 · ledger entry with the reason
/loop-r status
Worked example

Four assistant-run cycles on synthetic data

Aesthetics clinic, WhatsApp proposals, 1,500 synthetic runs. The assistants really ran; the data did not — it was generated to reflect a real effect, and the README in the example explains exactly what is real and what is simulated. What happened teaches more than a perfect case would.

Cycle 0001 — the Guardian vetoed the first idea

The Optimizer proposed closing with “would you rather start this week or next?” The Guardian vetoed it: a forced choice creates pressure and complaints. The second hypothesis (a message of up to 80 words) went to testing — after the Experimenter calculated that the original test would take 40 weeks.

Cycle 0002 — it won on the metric, failed the guardrail

The response rate rose from 19% to 30% (N=300/300). But complaints were 6 versus 1 with zero tolerance, so the Evaluator returned A stays. The rule was wrong (6 vs. 1 out of 300 is noise), not the system: a person adjusted the tolerance for later tests, and the hypothesis was recorded as discarded.

Cycle 0003 — the hypothesis that was waiting

A hypothesis approved in the first cycle waited in the queue for thirteen weeks — only one test can run at a time. The Meta-agent pointed out the cost of waiting and suggested an automatic trigger. It went onto the roadmap.

Cycle 0004 — promotion

Open the proposal by mentioning something specific from the customer’s review: responses rose from 19.3% to 29.3%, with guardrails within tolerance. Decision card, approval, versoes/v2. The loop’s first proven learning.

# memoria/ledger.md — excerpt (one line per event, never deleted)
| 2026-02-02 | ciclo 0001 | H1 (fechar com pergunta de escolha forçada) | — | — | — | VETADA |
| 2026-04-27 | ciclo 0002 | H2 (teto de 80 palavras, E0001) | A_SEGUE — resposta A=0,1933 vs B=0,2967, p=0,0033 | reclamacoes Δ+0,0167 tol 0,0 → FALHA | aceito | DESCARTADA |
| 2026-05-04 | ciclo 0003 | H4 (valor logo após a saudação) | — | — | — | VETADA |
| 2026-05-04 | ciclo 0003 | H3 (abrir citando a avaliação, E0002) | AMOSTRA_INSUFICIENTE — N=0/212 | — | — | EM TESTE |
| 2026-07-27 | ciclo 0004 | H3 (abrir citando a avaliação, E0002) | B_GANHOU — resposta A=0,1933 vs B=0,2933, p=0,0043 | todas OK | aprovado | PROMOVIDA v2 |
Roadmap

Start with the data, not the agents

Without evidence, everything else is theater. Each phase delivers something that works on its own, and the next starts only after a real user has three cycles in the history.

v0.1
One loop, one spreadsheet, one cycle — readyFive questions → loop profile → nine assistants → CSV cycle, with Git as history. The Meta-agent only reports. Worked example.
v0.2
Measure for realYes/no checklist for the work product, judge using shuffled pairs, calibration with 15 examples you evaluate. Complete course (5 tracks).
v1
Runs without youRunner independent of Claude Code, scheduling, a trigger when the sample reaches its minimum, Google Sheets and WhatsApp connectors, and automatic promotion with rollback (L2) after five correct promotions.
v2
Multiple loopsSales, support, and content in the same repository; a meta-loop reads all histories. Research depends on months of real data.