It's not about telling the agent to "work for 10 hours". It's giving it a goal with a verifiable done condition, keeping state in files and setting caps. It works in cycles until it finishes.

This repository brings together the source material on the long session in Codex, a sourced survey of what changed in JulβOct/2026, a usage plan and ready-to-copy templates.
The /goal needs Result, Constraints and Verification: commands whose output proves it's done. No "until it looks good".
goal, plan, state, progress, failures and decisions.md hold the task; canal.md holds what compaction loses (facts, glossary, pitfalls). After compacting or resuming, the agent rereads the files instead of relying on the conversation's memory.
Time, tokens and memory get a limit. Credit spend, APIs, production and irreversible actions become human gates.
The agent always picks the next useful action, tests, records and keeps going. Three cycles with no measurable progress = stop and call the human.
Logical continuity of the work: same goal, same history.
Summarizes old history to fit the window. It changes the prefix and the cache drops right after β that's expected.
Reuses the identical prefix. In continuous sessions the hit rate exceeds 95%. Measure cached Γ· input.
A coding agent with a goal mode and a project with some automated test to act as the oracle.
/goal lives inside the TUI (it doesn't show up in --help).
# recent version codex --version
It also has /goal, plus /loop, background agents and workflows.
claude --versionClone this repository to copy the state files and the prompts.
git clone https://github.com/inematds/execucao-longa
Use the templates in templates/. Each run gets its own folder inside the project.
One folder per run, with seven files (the six state files + canal.md).
execucao-longa/tools/novo-longrun.sh . meu-objetivo # creates the folder with the seven files and dates goal.md # or by hand: mkdir -p longrun/2026-10-01-meu-objetivo cp execucao-longa/templates/{goal,plan,state,progress,failures,decisions,canal}.md \ longrun/2026-10-01-meu-objetivo/
Each criterion is a command and its expected output. Also list the human gates.
## Done criteria (verifiable) - [ ] npm test β 0 failures - [ ] npm run build β exits with code 0 ## Human gates (stop and ask) - credit spend / paid API / production deploy
In Codex, paste a filled-in templates/prompt-goal-codex.md. If the goal is still vague, run /plan first.
codex /goal RESULT: ... VERIFICATION: ... STATE: longrun/2026-10-01-meu-objetivo/ # in Claude Code: the condition must show up in the output /goal the npm test output shows 0 failures and state.md says "done"
Check progress, pause and resume. Use fork for alternative paths.
/goals # lists goals and status /goal pause # /goal resume to continue /side # status question without stopping the work /fork # branches the session
Each cycle is one codex exec with closed stdin, timeout, memory cap and flock. The test decides whether to continue: the loop reverts changes to protected tests, commits a checkpoint and stops on stagnation.
# loop.env + prompt.md na pasta (copie de longrun/2026-10-01-medir-sessao/) execucao-longa/tools/loop-longrun.sh longrun/2026-10-01-meu-objetivo # stops on its own: DONE, 3 cycles without progress, or the cycle cap
Check it yourself (run the final test, look at the test hash) and record in progress.md what medir-sessao.py shows: duration, compactions, tokens, cache and tool output.
python3 execucao-longa/tools/medir-sessao.py <sessΓ£o.jsonl> # summary: duration, compactions, cache python3 execucao-longa/tools/medir-sessao.py <sessΓ£o.jsonl> --json --por-turno # per-turn curve
Install once per machine. Then every long goal follows one of two paths: interactive (you watch it) or headless (it runs on its own).
~/projetos/execucao-longa.templates/AGENTS-long-run.md into your global CLAUDE.md and AGENTS.md.~/.claude/settings.json, hook tools/hook-longrun.sh to PreCompact, SessionStart (matcher compact|resume), PostToolUse and UserPromptSubmit: it reminds to save before compacting, says to reread afterwards and warns at the context bands (50/70/85%). Set "cleanupPeriodDays": 365 so Claude does not delete transcripts after 30 days.tools/systemd/longrun-vigia.* to ~/.config/systemd/user/ and run systemctl --user enable --now longrun-vigia.timer.flock -n + timeout.tools/novo-longrun.sh <project> <slug>goal.md with level-3 criteria./goal in Codex or Claude Code with the prompt from templates/.medir-sessao.py and note it in progress.md.tools/novo-longrun.sh <project> <slug>goal.md + frozen tests (the test is the judge).prompt.md + loop.env β copy from the example longrun/2026-10-01-medir-sessao/.tools/loop-longrun.sh <folder> (in tmux or in the background).To find something said in any session (Codex or Claude): recall "term" --projeto X --desde YYYY-MM-DD. Watchdog alerts go to ~/.local/state/execucao-longa/alertas.log (and notify-send). tools/arquivar-sessoes.py shows how much space old sessions take; --aplicar compresses them and --restaurar brings one back. Full run example: longrun/2026-10-01-medir-sessao/.
A good criterion can be checked by an outsider without trusting the agent, and the agent cannot meet it through a shortcut. For long runs, level 3 is the minimum.
"Make the site good". Nobody knows when it is done.
"Code reviewed and clean". The agent approves itself.
"0 failures", "20 pages". It can delete tests or generate empty pages.
Command + lock: "0 failures and β₯ 48 tests and tests/ untouched". Minimum accepted.
Level 3 + external check: real e2e, separate evaluator, human sample. When mistakes are costly.
command β expected output)?/goal evaluator only reads the conversation.β Level 0: the sentence itself. β οΈ Level 2: "guia/en/index.html exists". β
Level 3β4: same <section id> as PT, no Portuguese left, internal links return 200, screenshot checked.
Signs of a bad criterion: "good", "clean", "complete"; relies on the agent saying it is done; counting only; does not say what must not change; verification that takes hours. templates/goal.md already includes the scale.
/goal works, but sessions with many compactions lose the thread. Details and sources in docs/pesquisa-goal-contexto-fila-2026-10.md (in Portuguese).
It is not /goal, it is repeated compaction. The objective survives, but "what is done / what is left" gets lost: the agent keeps saying "I will finish and commit", reopens work and never converges (issue openai/codex #34095).
Context bands, before the automatic one (~90β95%): ~50% β note facts, learnings and pitfalls in canal.md; ~70% β update state and run /compact; ~85% or 3rd compaction β /session-handoff, fresh session and /prime. In Claude Code a hook measures the % and warns the agent once per band. One goal per feature.
If the orchestrator sleeps between cycles, wake it before the cache expires: OpenAI 30 min (wake every ~25), Claude 1 h (~55). Losing the cache costs 12.5x to 25x. Waking only to keep the cache warm pays off if there is work in the next few hours.
A growing backlog: the orchestrator takes the next task, dispatches it in a short session, checks it and closes it. It stops when no task is open (OpenAI's Symphony pattern, on top of Linear). Start with a local file-based queue.
claude-mem works and is good for remembering what was already done in a project, but it only covers Claude Code and stores summaries, not what was said. To search word for word across everything (Codex and Claude), use recall "term" (SQLite FTS5 index, F8). The cleanup with backup (F9) cleared the old queue and closed the stuck sessions. Analysis in docs/pesquisa-memoria-claude-mem-2026-10.md (in Portuguese).
longrun/<run>/tasks/, with a status and an evidencia: field.proposta and wait for approval.progress.md and stop. Linear/GitHub Issues only with authorization to use the API.Summary of the sourced research in docs/pesquisa-web-2026-10.md (in Portuguese).
/goal, /compact, /resume and /fork exist in Codex (official docs). Goals resume after a usage limit (07/20) and after a daemon restart (09/17). Claude Code has had /goal since May/2026, with an evaluator after each turn.
GPT-6 Astra (09/03) keeps notes across context windows. Claude Opus 5.5 (09/22) has 1M context and cache reads at US$0.20/MTok.
The 11-day / 573-turn / 1.3 GB session has no public source. It's plausible: there are reports of session logs from 0.7 to 2 GB.
The state files are a split of the same idea as PLANS.md (OpenAI) and PROGRESS.md + git (Anthropic).
The infographic and the cover that gave rise to the project, kept in docs/origem/ together with the four texts.


Details in docs/PLANO-EXECUCAO-LONGA.md (in Portuguese).
recall "term" command: SQLite FTS5 index of everything said with Codex and Claude (no tool output), reindexed hourly. Built by a long run in 1 cycle; 90k snippets, search in 0.02β0.04 s.