Codex /goal Β· Claude Code /goal Β· Oct/2026

Agents that work for hours, in long runs

It's not about telling the agent to "work for 10 hours". It's giving it a goal with a verifiable done condition, keeping state in files and setting caps. It works in cycles until it finishes.

Long Runs banner: agents that work for hours and days without losing their way
What it is

A method + templates for long runs

This repository brings together the source material on the long session in Codex, a sourced survey of what changed in Jul–Oct/2026, a usage plan and ready-to-copy templates.

The six pillars: /goal, state files, continuous loop, guardrails, metrics and phases

🎯 Verifiable goal

The /goal needs Result, Constraints and Verification: commands whose output proves it's done. No "until it looks good".

πŸ—‚οΈ State in files

goal, plan, state, progress, failures and decisions.md hold the task; canal.md holds what compaction loses (facts, glossary, pitfalls). After compacting or resuming, the agent rereads the files instead of relying on the conversation's memory.

πŸ›‘οΈ Caps and gates

Time, tokens and memory get a limit. Credit spend, APIs, production and irreversible actions become human gates.

How it works

A loop that only stops when the criterion is met

The agent always picks the next useful action, tests, records and keeps going. Three cycles with no measurable progress = stop and call the human.

Goal→ Read state→ Next action→ Execute→ Test→ Fix→ Save state + commit↺

Persistent session

Logical continuity of the work: same goal, same history.

Compaction

Summarizes old history to fit the window. It changes the prefix and the cache drops right after β€” that's expected.

Prompt cache

Reuses the identical prefix. In continuous sessions the hit rate exceeds 95%. Measure cached Γ· input.

Prerequisites

What you need

A coding agent with a goal mode and a project with some automated test to act as the oracle.

Codex CLI

/goal lives inside the TUI (it doesn't show up in --help).

# recent version
codex --version

or Claude Code

It also has /goal, plus /loop, background agents and workflows.

claude --version

The templates

Clone this repository to copy the state files and the prompts.

git clone https://github.com/inematds/execucao-longa
User guide Β· step by step

From the first line of the goal to "done"

Use the templates in templates/. Each run gets its own folder inside the project.

1

Create the run's state folder

One folder per run, with seven files (the six state files + canal.md).

execucao-longa/tools/novo-longrun.sh . meu-objetivo   # creates the folder with the seven files and dates goal.md

# or by hand:
mkdir -p longrun/2026-10-01-meu-objetivo
cp execucao-longa/templates/{goal,plan,state,progress,failures,decisions,canal}.md \
   longrun/2026-10-01-meu-objetivo/
2

Write goal.md with verifiable criteria

Each criterion is a command and its expected output. Also list the human gates.

## Done criteria (verifiable)
- [ ] npm test  β†’ 0 failures
- [ ] npm run build  β†’ exits with code 0
## Human gates (stop and ask)
- credit spend / paid API / production deploy
3

Start the goal

In Codex, paste a filled-in templates/prompt-goal-codex.md. If the goal is still vague, run /plan first.

codex
/goal RESULT: ... VERIFICATION: ... STATE: longrun/2026-10-01-meu-objetivo/
# in Claude Code: the condition must show up in the output
/goal the npm test output shows 0 failures and state.md says "done"
4

Follow along without interrupting

Check progress, pause and resume. Use fork for alternative paths.

/goals          # lists goals and status
/goal pause     # /goal resume to continue
/side           # status question without stopping the work
/fork           # branches the session
5

No UI: the headless loop

Each cycle is one codex exec with closed stdin, timeout, memory cap and flock. The test decides whether to continue: the loop reverts changes to protected tests, commits a checkpoint and stops on stagnation.

# loop.env + prompt.md na pasta (copie de longrun/2026-10-01-medir-sessao/)
execucao-longa/tools/loop-longrun.sh longrun/2026-10-01-meu-objetivo
# stops on its own: DONE, 3 cycles without progress, or the cycle cap
6

Wrap up and measure

Check it yourself (run the final test, look at the test hash) and record in progress.md what medir-sessao.py shows: duration, compactions, tokens, cache and tool output.

python3 execucao-longa/tools/medir-sessao.py <sessΓ£o.jsonl>              # summary: duration, compactions, cache
python3 execucao-longa/tools/medir-sessao.py <sessΓ£o.jsonl> --json --por-turno  # per-turn curve
Getting started

How to use it all

Install once per machine. Then every long goal follows one of two paths: interactive (you watch it) or headless (it runs on its own).

1 Β· Once per machine

  1. Clone the repo into ~/projetos/execucao-longa.
  2. Paste templates/AGENTS-long-run.md into your global CLAUDE.md and AGENTS.md.
  3. In ~/.claude/settings.json, hook tools/hook-longrun.sh to PreCompact, SessionStart (matcher compact|resume), PostToolUse and UserPromptSubmit: it reminds to save before compacting, says to reread afterwards and warns at the context bands (50/70/85%). Set "cleanupPeriodDays": 365 so Claude does not delete transcripts after 30 days.
  4. Turn on the watchdog: copy tools/systemd/longrun-vigia.* to ~/.config/systemd/user/ and run systemctl --user enable --now longrun-vigia.timer.
  5. On cron jobs that call agents: flock -n + timeout.

2A Β· Interactive β€” you watch it

  1. tools/novo-longrun.sh <project> <slug>
  2. Write goal.md with level-3 criteria.
  3. /goal in Codex or Claude Code with the prompt from templates/.
  4. The hook says to reread state after compacting or resuming; the watchdog warns if it stalls.
  5. At the end: medir-sessao.py and note it in progress.md.

2B Β· Headless β€” runs on its own

  1. tools/novo-longrun.sh <project> <slug>
  2. goal.md + frozen tests (the test is the judge).
  3. prompt.md + loop.env β€” copy from the example longrun/2026-10-01-medir-sessao/.
  4. tools/loop-longrun.sh <folder> (in tmux or in the background).
  5. It stops on its own: done, stagnation or cap. Check independently and measure.

Maintenance

To find something said in any session (Codex or Claude): recall "term" --projeto X --desde YYYY-MM-DD. Watchdog alerts go to ~/.local/state/execucao-longa/alertas.log (and notify-send). tools/arquivar-sessoes.py shows how much space old sessions take; --aplicar compresses them and --restaurar brings one back. Full run example: longrun/2026-10-01-medir-sessao/.

Done criteria

How to tell a good success criterion

A good criterion can be checked by an outsider without trusting the agent, and the agent cannot meet it through a shortcut. For long runs, level 3 is the minimum.

0 Β· Vague

"Make the site good". Nobody knows when it is done.

1 Β· Subjective

"Code reviewed and clean". The agent approves itself.

2 Β· Gameable

"0 failures", "20 pages". It can delete tests or generate empty pages.

3 Β· Protected βœ…

Command + lock: "0 failures and β‰₯ 48 tests and tests/ untouched". Minimum accepted.

4 Β· Independent

Level 3 + external check: real e2e, separate evaluator, human sample. When mistakes are costly.

5 questions (each "no" lowers the level)

  1. Can it be checked with a command (command β†’ expected output)?
  2. Is the answer yes/no, with no "improved"?
  3. Is it impossible to meet without doing the work? If not, add a lock: minimum count, frozen file, validator.
  4. Does the proof show up in the output? Claude's /goal evaluator only reads the conversation.
  5. Does it cover function (does what it should), regression (did not break the rest) and limits (did not touch what it should not)?

Example: "translate the guide into English"

❌ Level 0: the sentence itself. ⚠️ Level 2: "guia/en/index.html exists". βœ… Level 3–4: same <section id> as PT, no Portuguese left, internal links return 200, screenshot checked.

Signs of a bad criterion: "good", "clean", "complete"; relies on the agent saying it is done; counting only; does not say what must not change; verification that takes hours. templates/goal.md already includes the scale.

Long sessions Β· pitfalls

Context rot, caching and queue mode

/goal works, but sessions with many compactions lose the thread. Details and sources in docs/pesquisa-goal-contexto-fila-2026-10.md (in Portuguese).

🧠 Does /goal degrade?

It is not /goal, it is repeated compaction. The objective survives, but "what is done / what is left" gets lost: the agent keeps saying "I will finish and commit", reopens work and never converges (issue openai/codex #34095).

πŸ›‘οΈ How to avoid it

Context bands, before the automatic one (~90–95%): ~50% β†’ note facts, learnings and pitfalls in canal.md; ~70% β†’ update state and run /compact; ~85% or 3rd compaction β†’ /session-handoff, fresh session and /prime. In Claude Code a hook measures the % and warns the agent once per band. One goal per feature.

⏱️ Cache between cycles

If the orchestrator sleeps between cycles, wake it before the cache expires: OpenAI 30 min (wake every ~25), Claude 1 h (~55). Losing the cache costs 12.5x to 25x. Waking only to keep the cache warm pays off if there is work in the next few hours.

πŸ“‹ Queue mode

A growing backlog: the orchestrator takes the next task, dispatches it in a short session, checks it and closes it. It stops when no task is open (OpenAI's Symphony pattern, on top of Linear). Start with a local file-based queue.

πŸ”Ž Memory across sessions

claude-mem works and is good for remembering what was already done in a project, but it only covers Claude Code and stores summaries, not what was said. To search word for word across everything (Codex and Claude), use recall "term" (SQLite FTS5 index, F8). The cleanup with backup (F9) cleared the old queue and closed the stuck sessions. Analysis in docs/pesquisa-memoria-claude-mem-2026-10.md (in Portuguese).

Queue mode rules

  • One task per file in longrun/<run>/tasks/, with a status and an evidencia: field.
  • Closing requires evidence (test output, file, commit hash) β€” otherwise "close without doing" becomes the shortcut.
  • The agent creates at most 5 tasks per run; beyond that they become proposta and wait for approval.
  • Empty queue β†’ stop workers, log it in progress.md and stop. Linear/GitHub Issues only with authorization to use the API.
Research Β· Jul–Oct/2026

What changed and what is myth

Summary of the sourced research in docs/pesquisa-web-2026-10.md (in Portuguese).

βœ… Confirmed

/goal, /compact, /resume and /fork exist in Codex (official docs). Goals resume after a usage limit (07/20) and after a daemon restart (09/17). Claude Code has had /goal since May/2026, with an evaluator after each turn.

πŸ†• New models

GPT-6 Astra (09/03) keeps notes across context windows. Claude Opus 5.5 (09/22) has 1M context and cache reads at US$0.20/MTok.

❓ No source

The 11-day / 573-turn / 1.3 GB session has no public source. It's plausible: there are reports of session logs from 0.7 to 2 GB.

πŸ“ Not an official standard

The state files are a split of the same idea as PLANS.md (OpenAI) and PROGRESS.md + git (Anthropic).

Source material

Where it started

The infographic and the cover that gave rise to the project, kept in docs/origem/ together with the four texts.

Codex infographic β€” Long Runs
Infographic: /goal, continuous loop, state files, autonomy, gates and metrics.
Cover: Long run, 11 days in the same session
Experiment cover: 11 days in the same session (number with no public source).
Roadmap

Plan phases

Details in docs/PLANO-EXECUCAO-LONGA.md (in Portuguese).

F0 βœ“
Documentation and templatesSource material, research, plan and templates published.
F1 βœ“
PilotA real long run (medir-sessao): level-3 goal, 14 frozen tests, headless loop β€” done in the 1st cycle, in 3 min, and checked independently.
F2 βœ“
Measurementtools/medir-sessao.py: reads 1.8 GB in ~9 s; matches the oracle on 2 large sessions and on a third unseen one.
F3 βœ“
RulesLong-run block in Claude's and Codex's global instructions; a fresh session follows it without being reminded.
F4 βœ“
Safety netflock + timeout on cron jobs that call agents, --max-turns, a hook that says to reread after compacting or resuming, a structured compaction summary in Codex.
F5 βœ“
Agent watchdogtools/vigia.py every 10 min (systemd timer): flags runs that stopped, stopped silently or went idle.
F6 ◐
Hygienetools/arquivar-sessoes.py ready and tested (compresses with verification, restores); not applied β€” only 0.04 GB were older than 90 days.
F7 ◐
"How far can it go" experimentRetrospective curve done: cache does not drop with compactions; without compacting, cost per turn grows ~5x. The prospective single-session Γ— queue-mode experiment is pending.
F8 βœ“
History searchrecall "term" command: SQLite FTS5 index of everything said with Codex and Claude (no tool output), reindexed hourly. Built by a long run in 1 cycle; 90k snippets, search in 0.02–0.04 s.
F9 βœ“
claude-mem cleanupDone with a backup: 7,904 old pending messages and 3,860 stuck sessions (automated batches from July) handled without reprocessing; logs from 640 to 129 MB; search and capture still work.
Related

Other INEMA projects on the topic