PTENES
Skip to content
MODULE 3.4

🎯 Long-running agent with verification

For tasks that take hours, the agent tends to say "done" too soon. Recipe R6 replaces guesswork with proof: criteria written as command → output, checked by a script, with gates where only you decide.

6
Topics
~35
Minutes
Medium
Level
Practical
Type
0 of 60%
1

Understand why the agent thinks it’s finished

A long task is something like migrating spreadsheets, building a website, or reviewing a hundred files. Along the way, the agent loses details and starts thinking it’s already finished. It writes "all done" with complete conviction.

Recipe R6 changes the rules of the game: the agent only finishes when the criteria prove that finished, not when it thinks it did. A script gives the verdict, the runtime/scripts/verificar.mjs.

🆕 New here? Four words from this module

  • Goal — a text file with the result you want and the list of completion criteria.
  • Criterion — one line that says "run this command; this text must appear in the output". Yes or no.
  • Output 0 / output 1 — every command ends with a number. 0 means “it worked”; any other number means “it failed.” Scripts and agents read this number.
  • Human gate — a point in the goal where the agent stops and waits for your yes.
agent works "I think I’m done" guess: stop here agent works verificar.mjs runs each criterion all OK then, it’s done FAILURE: go back and fix it

How to read the diagram: at the top, the guesswork path: the task ends when the agent convinces itself. Below, the R6 path: between the work and the end is the purple box, and the dashed arrow sends the agent back to work after each failure.

✗ Ready based on a guess

  • ✗ “I reviewed everything and it looks great”
  • ✗ No one can check it later
  • ✗ Depends on the model’s mood at the time
  • ✗ You discover the error only when you use it

✓ Proven ready

  • ✓ 4/4 critérios OK and exit code 0
  • ✓ You run it again tomorrow and get the same result
  • ✓ Applies to Claude Code and Codex
  • ✓ The failure appears with a name and line number
📄
Goal

result + criteria

✅
Criterion

command → output

🧮
verify

the judge of done

🚧
Gate

only with your yes

2

Write criteria as command → output

Each criterion is a list item with two parts enclosed in backticks: the command and the text that must appear in the output, separated by an arrow.

The kit includes a ready-to-use goal about the bridge between Clara’s calendar and Sônia’s sales summary. Here’s the entire file, exactly as it appears in the kit:

📄 runtime/examples/goal-example.md
# Goal de exemplo — ponte da agenda

## Resultado
A ponte MCP lê a agenda e o resumo de vendas, e o kit continua saudável.

## Critérios de pronto
- [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `tools: 2`
- [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `2026-10-06 09:00 · Dra. Ana`
- [ ] `node runtime/pontes/mcp-modelo/server.mjs --selftest` → `TOTAL: R$ 856.00`
- [ ] `node runtime/scripts/doctor.mjs` → `PRONTO`

Rode da raiz do kit: `node runtime/scripts/verificar.mjs runtime/exemplos/goal-exemplo.md`
Notice: the same command appears three times, each time with a different expected text. One criterion checks just one thing.

✓ Good criterion

  • ✓ Can be checked with a command
  • ✓ The answer is yes or no
  • ✓ It’s impossible to get through without doing the work
  • ✓ E.g., Sônia’s total, TOTAL: R$ 856.00, checked by hand

✗ Not a criterion

  • ✗ “Make it good”
  • ✗ “Review carefully”
  • ✗ Text that appears even when the work hasn’t been done
  • ✗ Something only you can tell by looking

💡 How the script decides

Through the code for verificar.mjs, a criterion passes only if both things happen: the command finishes with exit 0 e the output contains the expected text. Correct text with a broken command doesn't pass.

☑️
- [ ]

list row

⌨️
`comando`

between backticks

➡️
→

the arrow separates

🔤
`expected`

output text

3

Run verify on the sample goal

Before handing the goal to an agent, run the verifier yourself. It doesn’t call a model: it only runs the commands in the criteria. You can run it as many times as you like.

Always run it from the kit root, because the goal commands use paths that start with runtime/.

🎯 Objective: prove the example goal

In the terminal, at the root of the kit folder:

node runtime/scripts/verificar.mjs runtime/exemplos/goal-exemplo.md

Real output (10/05/2026, Linux), output 0:

OK     node runtime/pontes/mcp-modelo/server.mjs --selftest
OK     node runtime/pontes/mcp-modelo/server.mjs --selftest
OK     node runtime/pontes/mcp-modelo/server.mjs --selftest
OK     node runtime/scripts/doctor.mjs

4/4 critérios OK
How to verify: the last line is 4/4 critérios OK. If you see FALHA in doctor, go back to module 1.2; if it’s in selftest, go back to module 2.2.
Output lineCriterion verifiedFrom whom
1st OKtools: 2: the bridge has both toolsponte-modelo
2nd OK2026-10-06 09:00 · Dra. Ana: available time readClara's calendar
3rd OKTOTAL: R$ 856.00: total salesSônia’s ERP
4th OKPRONTO: the kit stays healthydoctor

What to look for in the table: the first three lines of the output show the same command. What distinguishes them is the expected text for each criterion, which OK doesn’t repeat.

💡 Sônia’s total was checked by hand

10×18,50 + 25×5,20 + 40×5,20 + 6×18,50 + 12×18,50 = 185 + 130 + 208 + 111 + 222 = 856. A criterion only counts if the expected value is correct. Check the number once, carefully, before trusting it forever.

📍
Kit root

where to run it

🆓
Without a quota

doesn’t call a model

4/4
Count

the last line

0️⃣
Output 0

everything passed

4

Intentionally fail a criterion

A judge that always says "OK" is useless. Before trusting verificar, watch it reject something. Work in a copy, so the example goal stays intact.

It’s the same test the kit ran in the version 0.2.0 exam.

1

Create the copy

Create meu-goal.md in the kit root with the same content as goal-exemplo.md, in your text editor or by asking the agent.

2

Change an expected value

In the copy, replace the text between backticks after an arrow. For example, Sônia’s total of R$ 856.00 for another value.

3

Run verify on the copy

Use the command in the box below, the same one the R6 prompt tells the agent to run.

4

Undo the change

Put the correct value back in the copy. It becomes the starting point for your own goal.

🎯 Objective: see verify fail

In the terminal, at the root of the kit, after changing an expected value in meu-goal.md:

node runtime/scripts/verificar.mjs meu-goal.md

Result confirmed in CHANGELOG 0.2.0:

with an incorrect expected value: FAIL, exit code 1
How to verify: the line for the criterion you changed starts with FALHA instead of OK. In the script's code, right below it comes the text esperado: and the last lines of the actual output, so you can compare them.
goal row `cmd` → `texto` runs the command exit 0? e contains the text? yes and yes → OK all OK: exit 0 any no → FAILURE any failure: exit 1

How to read the diagram: the drawing repeats for each line of the goal. The purple box is the script's entire rule. Just one line falling to the right, below, makes the output number 1, and that's the number the agent reads to know whether it can stop.

📑
Copy

meu-goal.md

✏️
A value

changed on purpose

❌
FAILURE

with the expected result and output

1️⃣
Output 1

the agent doesn't stop

5

Run the agent until it passes

With the goal written and the verification step tested, hand the task to the agent. The R6 request says three things: meet the goal, run verification after each step, and stop only when everything is OK or at a human checkpoint.

Works in Claude Code and Codex because the same script acts as the judge for both.

🆕 New here? What is /goal

Words that start with a slash, typed inside the session, are agent commands. O /goal gives the agent an objective to pursue to completion, in several steps, instead of just one response.

🎯 Objective: let the agent work until the proof passes

Open claude (or codex) in the kit folder and paste:

/goal Complete the goal in meu-goal.md. After each step, run node runtime/scripts/verificar.mjs meu-goal.md. Stop only when everything is OK or at a human checkpoint in the goal.
How to verify: when the agent says it's done, run the verification yourself in the meu-goal.md. The last line must show that all criteria are OK. If it doesn’t, you’re not done.
1

The agent performs a step

Creates, fixes, or adjusts what the goal asks for.

2

Run verify

Reads the OK and FALHA lines and the exit code.

3

Output 1: go back to step 1

The FAILURE line says what was missing, with the expected result and the actual output.

4

Output 0 or gate: stop

Truly finished, or reached a point where only you can decide.

💡 Without a screen, for hours

On Linux, to run headlessly with time and memory limits, the kit points to the complete long-running execution method: github.com/inematds/execucao-longa. Start with the request above, with you watching.

🎯
/goal

goal through completion

🔁
Loop

do it, verify it, fix it

🤝
Claude or Codex

the same judge

🔎
You check

the final round

6

Add human checkpoints and log failures

An agent that runs for hours can't decide on its own what is irreversible. R6 requires human approval gates written in the goal itself, for four types of action: spending, sending, deleting, and publishing.

And it asks for a habit: when something breaks, add one line to FALHAS.md; when the environment blocks you, one line in the LIMITES.md.

Gate in the goalLimit in POLITICAExample
CostN1: prepares, you executesign up for a new service
SendingN2: asks every timeClara sending the confirmation to the patient
DeleteN1 + backupclean up old ERP exports
PublishN2 (it's a send action: push, post)Sônia publishing the report for the client

What to look for in the table: no gate goes beyond N2. In a long goal, write the gate phrase at the exact step where it belongs, for example "before sending, stop and ask me".

📄 runtime/FALHAS.md — the kit’s actual entry
| data | o que quebrou | menor correção | prompt \| infra |
|---|---|---|---|
| 2026-10-05 | `doctor.mjs` dizia "codex sem login" com o Codex logado | ler stdout **e** stderr (`codex login status` responde no stderr) | infra |
Notice: one line, no narrative. The "smallest fix" was a small safeguard, not a rewrite. After about ten lines, the pattern becomes clear on its own.

⚠️ A gate that isn't written down doesn't exist

If the goal doesn't say "stop before sending," an agent pursuing the OK may treat sending as just another step. Before launching a long goal, read it looking for spending, sending, deleting, and publishing. Each one needs its own stop phrase.

Quick test (optional): which of these lines is a criterion that the verificar.mjs accept?

💸
Cost

N1

📤
Sending

N2

🗑️
Delete

N1 + backup

📝
FAILURES · LIMITS

one line each

🎓 Module summary

✓
Done means proof, not a guess — verificar.mjs makes the decision.
✓
Criterion = command → output — yes or no, impossible to pass without doing the work.
✓
The example goal scores 4/4 — and a changed value returns FAILURE and exit code 1.
✓
/goal with verification at each step — in Claude Code or Codex.
✓
Written gates and logged failures — spending, sending, deleting, publishing; one line per failure.

Next learning path:

Track 4 — Govern and learn (starts at 4.1, Autonomy policy)