Understand why the agent thinks it’s finished
A long task is something like migrating spreadsheets, building a website, or reviewing a hundred files. Along the way, the agent loses details and starts thinking it’s already finished. It writes "all done" with complete conviction.
Recipe R6 changes the rules of the game: the agent only finishes when the criteria prove that finished, not when it thinks it did. A script gives the verdict, the runtime/scripts/verificar.mjs.
🆕 New here? Four words from this module
- Goal — a text file with the result you want and the list of completion criteria.
- Criterion — one line that says "run this command; this text must appear in the output". Yes or no.
- Output 0 / output 1 — every command ends with a number. 0 means “it worked”; any other number means “it failed.” Scripts and agents read this number.
- Human gate — a point in the goal where the agent stops and waits for your yes.
How to read the diagram: at the top, the guesswork path: the task ends when the agent convinces itself. Below, the R6 path: between the work and the end is the purple box, and the dashed arrow sends the agent back to work after each failure.
✗ Ready based on a guess
- ✗ “I reviewed everything and it looks great”
- ✗ No one can check it later
- ✗ Depends on the model’s mood at the time
- ✗ You discover the error only when you use it
✓ Proven ready
- ✓
4/4 critérios OKand exit code 0 - ✓ You run it again tomorrow and get the same result
- ✓ Applies to Claude Code and Codex
- ✓ The failure appears with a name and line number
result + criteria
command → output
the judge of done
only with your yes
Write criteria as command → output
Each criterion is a list item with two parts enclosed in backticks: the command and the text that must appear in the output, separated by an arrow.
The kit includes a ready-to-use goal about the bridge between Clara’s calendar and Sônia’s sales summary. Here’s the entire file, exactly as it appears in the kit:
# Goal de exemplo — ponte da agenda ## Resultado A ponte MCP lê a agenda e o resumo de vendas, e o kit continua saudável. ## Critérios de pronto - [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `tools: 2` - [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `2026-10-06 09:00 · Dra. Ana` - [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `TOTAL: R$ 856.00` - [ ] `node runtime/scripts/doctor.mjs` → `PRONTO` Rode da raiz do kit: `node runtime/scripts/verificar.mjs runtime/exemplos/goal-exemplo.md`
✓ Good criterion
- ✓ Can be checked with a command
- ✓ The answer is yes or no
- ✓ It’s impossible to get through without doing the work
- ✓ E.g., Sônia’s total,
TOTAL: R$ 856.00, checked by hand
✗ Not a criterion
- ✗ “Make it good”
- ✗ “Review carefully”
- ✗ Text that appears even when the work hasn’t been done
- ✗ Something only you can tell by looking
💡 How the script decides
Through the code for verificar.mjs, a criterion passes only if both things happen: the command finishes with exit 0 e the output contains the expected text. Correct text with a broken command doesn't pass.
list row
between backticks
the arrow separates
output text
Run verify on the sample goal
Before handing the goal to an agent, run the verifier yourself. It doesn’t call a model: it only runs the commands in the criteria. You can run it as many times as you like.
Always run it from the kit root, because the goal commands use paths that start with runtime/.
In the terminal, at the root of the kit folder:
node runtime/scripts/verificar.mjs runtime/exemplos/goal-exemplo.md
Real output (10/05/2026, Linux), output 0:
OK node runtime/pontes/mcp-modelo/server.mjs --selftest OK node runtime/pontes/mcp-modelo/server.mjs --selftest OK node runtime/pontes/mcp-modelo/server.mjs --selftest OK node runtime/scripts/doctor.mjs 4/4 critérios OK
4/4 critérios OK. If you see FALHA in doctor, go back to module 1.2; if it’s in selftest, go back to module 2.2.| Output line | Criterion verified | From whom |
|---|---|---|
| 1st OK | tools: 2: the bridge has both tools | ponte-modelo |
| 2nd OK | 2026-10-06 09:00 · Dra. Ana: available time read | Clara's calendar |
| 3rd OK | TOTAL: R$ 856.00: total sales | Sônia’s ERP |
| 4th OK | PRONTO: the kit stays healthy | doctor |
What to look for in the table: the first three lines of the output show the same command. What distinguishes them is the expected text for each criterion, which OK doesn’t repeat.
💡 Sônia’s total was checked by hand
10×18,50 + 25×5,20 + 40×5,20 + 6×18,50 + 12×18,50 = 185 + 130 + 208 + 111 + 222 = 856. A criterion only counts if the expected value is correct. Check the number once, carefully, before trusting it forever.
where to run it
doesn’t call a model
the last line
everything passed
Intentionally fail a criterion
A judge that always says "OK" is useless. Before trusting verificar, watch it reject something. Work in a copy, so the example goal stays intact.
It’s the same test the kit ran in the version 0.2.0 exam.
Create the copy
Create meu-goal.md in the kit root with the same content as goal-exemplo.md, in your text editor or by asking the agent.
Change an expected value
In the copy, replace the text between backticks after an arrow. For example, Sônia’s total of R$ 856.00 for another value.
Run verify on the copy
Use the command in the box below, the same one the R6 prompt tells the agent to run.
Undo the change
Put the correct value back in the copy. It becomes the starting point for your own goal.
In the terminal, at the root of the kit, after changing an expected value in meu-goal.md:
node runtime/scripts/verificar.mjs meu-goal.md
Result confirmed in CHANGELOG 0.2.0:
with an incorrect expected value: FAIL, exit code 1
FALHA instead of OK. In the script's code, right below it comes the text esperado: and the last lines of the actual output, so you can compare them.How to read the diagram: the drawing repeats for each line of the goal. The purple box is the script's entire rule. Just one line falling to the right, below, makes the output number 1, and that's the number the agent reads to know whether it can stop.
meu-goal.md
changed on purpose
with the expected result and output
the agent doesn't stop
Run the agent until it passes
With the goal written and the verification step tested, hand the task to the agent. The R6 request says three things: meet the goal, run verification after each step, and stop only when everything is OK or at a human checkpoint.
Works in Claude Code and Codex because the same script acts as the judge for both.
🆕 New here? What is /goal
Words that start with a slash, typed inside the session, are agent commands. O /goal gives the agent an objective to pursue to completion, in several steps, instead of just one response.
Open claude (or codex) in the kit folder and paste:
/goal Complete the goal in meu-goal.md. After each step, run node runtime/scripts/verificar.mjs meu-goal.md. Stop only when everything is OK or at a human checkpoint in the goal.
meu-goal.md. The last line must show that all criteria are OK. If it doesn’t, you’re not done.The agent performs a step
Creates, fixes, or adjusts what the goal asks for.
Run verify
Reads the OK and FALHA lines and the exit code.
Output 1: go back to step 1
The FAILURE line says what was missing, with the expected result and the actual output.
Output 0 or gate: stop
Truly finished, or reached a point where only you can decide.
💡 Without a screen, for hours
On Linux, to run headlessly with time and memory limits, the kit points to the complete long-running execution method: github.com/inematds/execucao-longa. Start with the request above, with you watching.
goal through completion
do it, verify it, fix it
the same judge
the final round
Add human checkpoints and log failures
An agent that runs for hours can't decide on its own what is irreversible. R6 requires human approval gates written in the goal itself, for four types of action: spending, sending, deleting, and publishing.
And it asks for a habit: when something breaks, add one line to FALHAS.md; when the environment blocks you, one line in the LIMITES.md.
| Gate in the goal | Limit in POLITICA | Example |
|---|---|---|
| Cost | N1: prepares, you execute | sign up for a new service |
| Sending | N2: asks every time | Clara sending the confirmation to the patient |
| Delete | N1 + backup | clean up old ERP exports |
| Publish | N2 (it's a send action: push, post) | Sônia publishing the report for the client |
What to look for in the table: no gate goes beyond N2. In a long goal, write the gate phrase at the exact step where it belongs, for example "before sending, stop and ask me".
| data | o que quebrou | menor correção | prompt \| infra | |---|---|---|---| | 2026-10-05 | `doctor.mjs` dizia "codex sem login" com o Codex logado | ler stdout **e** stderr (`codex login status` responde no stderr) | infra |
⚠️ A gate that isn't written down doesn't exist
If the goal doesn't say "stop before sending," an agent pursuing the OK may treat sending as just another step. Before launching a long goal, read it looking for spending, sending, deleting, and publishing. Each one needs its own stop phrase.
Quick test (optional): which of these lines is a criterion that the verificar.mjs accept?
N1
N2
N1 + backup
one line each
🎓 Module summary
Next learning path:
Track 4 — Govern and learn (starts at 4.1, Autonomy policy)