INEMA.CLUBPROAI Development v6.2

AI Development v6.2 · 6 modules · 30 lessons of about 15 minutes each

One does the work, the other checks

Claude and Codex together, from briefing to approval: one plans or implements, the other critiques the actual artifact, and the tests and you decide. With the model and effort chosen for each step and the cost measured.

A developer and an operations analyst review a printed sheet with handwritten marks together, beside two open laptops.

Module 1 · Validation is the work

Replace “looks right” with “I checked it this way.”

Module 2 · Plan and critique

The plan is in a file, critiqued by the other model using that same file.

Module 3 · Build and review the diff

One builds in stages; the other reviews the diff against the briefing and the tests.

Module 4 · Verify and decide

Relevant tests, real behavior, and human judgment.

Module 5 · Continuity: handoff and prime

Switch sessions or models without explaining everything again.

Module 6 · Models, effort, and budget

Choose a model and effort level for each stage, and measure cost × results.

Glossary · 22 terms

AI Development v6.2

Glossary

The course’s technical terms in simple words. Each term links to the lessons where it appears.

AGENTS.md

Text file at the project root with the rules and reading order. Codex reads it automatically when starting; other assistants read it when someone tells them to.

Appears in: Lesson 23

API

An access point that lets a program use the AI service directly, billed by usage and opened with a secret key.

Appears in: Lesson 29

branch

A separate test line in Git. You test changes there without changing the version that already works.

Appears in: Lesson 11 Lesson 12 Lesson 15

briefing

A short file with the task’s goal, scope, acceptance criteria, and limits. Both models read the same one.

Appears in: Lesson 2 Lesson 5 Lesson 6 Lesson 7 Lesson 8 Lesson 9 Lesson 10 Lesson 11 Lesson 12 Lesson 13 Lesson 14 Lesson 15 Lesson 16 Lesson 17 Lesson 18 Lesson 20 Lesson 22 Lesson 23 Lesson 24 Lesson 25 Lesson 26 Lesson 27 Lesson 30

Claude Code

Anthropic’s coding assistant that reads and edits files in a folder on your computer.

Appears in: Lesson 1 Lesson 8

CLI

A program used through commands typed in the terminal, without windows or buttons.

Appears in: Lesson 8

Codex

OpenAI’s coding assistant that reads and edits files in a folder on your computer.

Appears in: Lesson 1 Lesson 8

commit

A version saved in Git, with a date and a sentence describing what changed.

Appears in: Lesson 11 Lesson 12 Lesson 13 Lesson 15

acceptance criterion

A sentence written before the work that says how you’ll check that it’s done. E.g.: 'the test message arrives in the reception inbox'.

Appears in: Lesson 1 Lesson 2 Lesson 3 Lesson 6 Lesson 10 Lesson 12 Lesson 16 Lesson 20 Lesson 26 Lesson 30

diff

The list of what changed between two versions of a file: lines that were removed (with −) and lines that were added (with +).

Appears in: Lesson 11 Lesson 12 Lesson 13 Lesson 15

reasoning effort

A control for how much the model analyzes before answering. More effort usually takes longer and uses more of your account’s allowance.

Appears in: Lesson 27

Git

A version control program: it saves every change to files in a folder so you can see what changed and go back. Module 3 teaches what you need.

Appears in: Lesson 8 Lesson 11 Lesson 12 Lesson 13 Lesson 15 Lesson 24

handoff

A short file that records decisions, pending items, and the next action so another conversation or model can pick up where you left off.

Appears in: Lesson 5 Lesson 21 Lesson 22 Lesson 23 Lesson 24 Lesson 25 Lesson 27

model

The AI program that reads your request and writes the response, such as Claude or GPT. Another model = another AI; a new conversation with the same assistant also works as a second opinion.

Appears in: Lesson 1 Lesson 2 Lesson 5 Lesson 6 Lesson 7 Lesson 10 Lesson 13 Lesson 17 Lesson 20 Lesson 21 Lesson 23 Lesson 24 Lesson 25 Lesson 26 Lesson 27 Lesson 28 Lesson 29 Lesson 30

plugin

An add-on that gives a program a new function.

Appears in: Lesson 8 Lesson 9

prime

The first request in a session: the assistant reads the project files and the handoff without changing anything, then returns a summary with sources.

Appears in: Lesson 22 Lesson 23 Lesson 24 Lesson 25

cross-review

One model does the work, and another model or conversation examines the actual result and points out flaws with evidence.

Appears in: Lesson 2 Lesson 7 Lesson 13 Lesson 20 Lesson 25 Lesson 28 Lesson 30

review round

A back-and-forth between a plan and a critique: the reviewer points out findings, and the author responds to each one and writes the new version.

Appears in: Lesson 9

sandbox

A protected area where the assistant works with limited permissions. In read-only mode, it can read the files but cannot change anything.

Appears in: Lesson 8 Lesson 24

terminal

A text window where you type commands for the computer to run and read its response.

Appears in: Lesson 4 Lesson 8 Lesson 11 Lesson 12 Lesson 13 Lesson 19 Lesson 24 Lesson 27 Lesson 30

automated test

A check that the computer runs by itself and reports “passed” or “failed,” the same way every time. AI can write the test for you.

Appears in: Lesson 16 Lesson 17

tokens

A piece of text that the model reads or writes. AI use is measured and billed in tokens.

Appears in: Lesson 4 Lesson 19 Lesson 28 Lesson 29

Module 1 · Lesson 1 of 5

Generating is easy. Checking is not.

A developer holds a printed webpage and circles an error in red pen, with an open laptop beside them.

You can look at a result made with AI and separate what you checked from what you simply accepted.

AI writes code, text, and spreadsheets in minutes, and with a lot of confidence. The problem is that confidence is not proof. The person who delivers the work is responsible for it, not the AI.

In 1 minute

  1. AI says "done" with the same confidence when it gets things right and when it gets them wrong.
  2. Checking means looking at the actual result, not the summary AI made of it.
  3. The decision to deliver is yours. AI helps you check, but it doesn’t sign off.

1AI says "done" in the same voice when it gets things wrong

Code assistants like Claude Code and Codex work quickly. At the end, they write a summary of what they did.

That summary is AI's opinion of its own work. Sometimes it’s right. Sometimes it says "tested" without having tested anything.

Rafael, an independent developer, asked for a contact form for a clinic's website. The response came in two minutes, with "tested and working".

Code assistant

RafaelCreate the contact form for the clinic's page. Send submissions to the reception email.

AIDone. I created the form with name, phone number, and message fields, and set submissions to go to reception. Tested and working.

It seems complete. But what was tested, and how?

RafaelHow exactly did you test sending it to reception?

AII checked that the page opens and the form appears. I didn’t send any real messages. To test sending, submit a message through the page and check the reception inbox.

One question revealed that "tested" only meant the page opened.

Tap or click the two buttons in the panel: the same task, before and after asking.

2Checking means looking at the actual result

Accepting means trusting the summary. Checking means opening what was made and seeing if it does what it’s supposed to.

Rafael opened the page, filled out the form, and clicked Submit. The screen showed "message sent," but nothing arrived in the reception inbox.

Accepted

What he did: read "tested and working" and sent the link to the clinic.

Result: reception discovers the error days later, with patients still waiting for a response.

Checked

What he did: sent a test message and checked the reception inbox.

Result: found the error before delivery and asked for a fix.

Net gain: five minutes of checking instead of a week of lost messages.

3Three ways to check

There isn’t just one way to check. This course uses three, always together.

The first is an acceptance criterion, written beforehand. The second is to open and use the result. The third is a second look: another model, a new conversation, or another person.

Carla, an operations analyst, uses AI to generate the weekly delivery report. She checks the total number of orders against the original spreadsheet before sending it to management.

Carla’s report check
1 Criterion: order total matches the spreadsheet
2 Open the report and add up the order column
3 Ask another AI, or a new conversation, for the numbers without sources
  1. 1Written before requesting the report.
  2. 2Done by her, in the actual file.
  3. 3An outside perspective, which module 2 teaches.

Stuck here? That's normalDoes this seem like too much work for every request? It isn’t for every request. Start with what you’ll send to someone else: a client, your boss, your team.

Test yourself

The AI finished the task and wrote "all tests passed." What is this?

4AI helps you check; you make the decision

A second model finds flaws the first one didn’t see. Even so, the decision to deliver is still yours.

Carla asks another model to review the report. It points out two cities that were switched. She decides whether the report goes out today.

What the AI does

Creates, reviews, and points out flaws with evidence.

What’s yours

Set the criterion, look at the actual result, and decide whether to deliver it.

Both cards are right: one is the AI’s job, the other is yours.

Practice now 0/3

Checked or just accepted?

You’re done when you’ve marked each item in the case as "checked" or "accepted." About 8 minutes, on paper or in a notes app.

This is just reading a case. If you’re unsure about an item, mark it "accepted": if you’re unsure, it wasn’t checked.

The case. Rafael asked AI for a pricing page. It replied: "Page created, all three plans are shown, the button leads to payment, and the text has been reviewed." Rafael opened the page and saw the three plans. He didn’t click the button. He didn’t read the text. He sent the link to the client.

Show answer key

Three plans: checked, he saw them. Button: accepted, no one clicked it. Text reviewed: accepted, no one read it. In one minute: click the button and see whether it opens the right payment page, then read the page text and look for an incorrect price or name.

You already separate what was seen from what was only said.

Lesson cheat sheet

Check

  1. AI summaryis a statement, not proof.
  2. Three checkscriteria written beforehand, result opened, second look.
  3. Decisionwhether to deliver it is always yours.

Your next step

You already know how to separate what you checked from what you only accepted.

For the next AI result you share with someone else, check one claim by opening the actual result. It takes five minutes.

In the next lesson: what if another model says everything is okay? Two AIs agreeing still isn't proof.

Lesson 1 · Dev with AI v6.2 · INEMA.CLUB

Module 1 · Lesson 2 of 5

Two AIs agreeing isn't proof

An operations analyst manually checks a line in a printed report with a desk calculator, after two models had already approved it.

You can say what completes a work review beyond the opinion of a second model.

Asking another model to review is a good step. The risk is stopping there: "the other one approved it too, so it's correct." Two calculators with the same wrong number entered give the same wrong result.

In 1 minute

  1. A different model spots flaws the author doesn't see.
  2. But both can be wrong together if they read the same thing incorrectly.
  3. What completes the review: acceptance criteria, a test on the actual result, and your decision.

1People tend to approve what they wrote

When you ask the same model, in the same conversation, to review its own work, it usually agrees with itself. A different model, or a new conversation, looks at it with less attachment.

Only have one assistant? Open a new conversation with it: that's already a less attached look. Another model is even better. The course calls this cross-review: one does it, the other checks it.

Rafael asked Claude for a plan for a password-protected area of the clinic's website. Then he asked Codex to read the same plan and point out flaws.

Plan review

RafaelReview the plan you just wrote.

ClaudeThe plan is complete and covers the requirements. I don't see any adjustments needed.

The person who wrote it, in the same conversation, confirmed their own plan.

RafaelRead the plan below and point out flaws, with the passage that shows each one.

CodexFlaw 1: the plan doesn't say what happens when someone enters the wrong password several times. Where: the plan's "Entry" section lists only email, password, and the sign in button.

An outside perspective found a case the author hadn't anticipated.

Tap or click the two buttons in the panel and compare the two reviews.

2Both can be wrong together

Cross-review finds flaws. It doesn't guarantee that none remain. If both models read the same data incorrectly, they'll both agree with the error.

Carla asked a model for the formula for the average shipping time, from when an order leaves the warehouse until delivery. The second model reviewed and approved it. Both counted from the order date, not the date it left the warehouse.

Two approved

What Carla had: the formula and both models saying "it's correct."

What happened: the shipping time came out two days longer than it really was.

Checked against the real result

What Carla did: calculated three orders by hand and compared them with the formula.

What happened: the difference showed up on the first order.

Net gain: three calculations by hand caught what two reviews missed.

Stuck here? That's normalSo, is review no use? It is, and very much so. It just isn't the final word. Think of it as another pair of eyes, not the final stamp of approval.

3What closes the check

Three things close the loop, and none of them is a model's opinion. The acceptance criterion written beforehand. The test against the real result. Your decision.

In the clinic's access area, Rafael's criterion was "after five incorrect passwords, a message appears telling you to wait ten minutes." He tested it by entering the wrong password five times.

What closes the loop
1 Acceptance criterion, written before the work
2 Test against the real result, done by someone
3 Human decision: does it go in or not?
  1. 1Says what to check.
  2. 2Shows whether it passed.
  3. 3Takes responsibility.

Test yourself

Claude wrote it, and Codex reviewed and approved it. What still needs to happen before you deliver it?

4The reviewer gets the original, not a summary

For the review to be useful, the second model gets the same briefing and the actual work: the file, plan, or spreadsheet. Never a summary of what the first model said it did.

Carla started pasting the original request and attaching the spreadsheet using the chat's paperclip. Before, she pasted only the first AI's response. In a coding assistant, just refer to the file by name.

Summary

"The other model said the formula is correct. Check it."

Original

The initial request, spreadsheet, and formula, with: "point out flaws and include the passage that shows each one."

With the summary, the reviewer checks a sentence. With the original, they check the work.

Practice now 0/3

What still needs to happen?

You're done when you have answered the case's three questions. About 8 minutes, on paper or in your notes app.

This is just reading a case. If two answers seem right, choose the one that includes a test against the real result.

The case. Carla asked Claude for a routine that combines the delivery spreadsheets for the week. She pasted the answer into Codex with “is this right?” Codex replied, “yes, it looks correct.” Carla sent the report to management.

Show answer key

1. Only Claude’s answer, without the spreadsheets or the original request: the material needed to check it was missing. 2. Something like “the total number of requests in the report equals the sum of the spreadsheets for the week.” 3. Add up the requests column in the spreadsheets and compare it with the report total.

You already know what’s missing when “they both agreed.”

Lesson cheat sheet

Cross-review

  1. A fresh perspectivea different model or a new conversation spots what the author doesn’t see.
  2. It’s not prooftwo models can make the same mistake.
  3. Who signs offthe criterion, the real-world test, and your decision.

Your next step

You already know how to use a second model without treating it as the final stamp of approval.

The next time you ask for a review, send the original and ask for “issues with the passage that shows each one.” It takes two extra minutes.

Next lesson: if the acceptance criterion closes the loop, how do you write one that works?

Supplementary material · Where this lesson comes fromFurther reading on the topic. Not included in the lesson time.

What the sources say

The INEMA summary from September 27, 2026 puts it this way: cross-review helps reveal problems, but agreement between models does not prove that the system works. Acceptance criteria, tests, and human verification close the loop.

The Use Both kit, summarized in the Codex + Claude area of Eventos INEMA, says the same thing at level 1: one AI writes the plan, the other critiques the same file, read-only, with evidence-backed findings and severity, in no more than two rounds. “Two AIs agreeing is not independent proof.”

Why self-review fails

In the view of Mark Kashef, author of the video that accompanies the kit, the model that wrote the plan, especially in the same session, tends to approve it. A different model or a new session gets a different result. This is the author’s account, without published measurements.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 2 · Dev with AI v6.2 · INEMA.CLUB

Module 1 · Lesson 3 of 5

The acceptance criterion comes before the work

An operations analyst handwrites a short checklist on a paper card before asking AI for the report.

You can write three acceptance criteria for a real task of yours, each answerable with yes or no.

Without a criterion written first, “done” becomes whatever the AI thinks done means. Then checking depends on memory and mood. With a criterion, anyone can check it the same way.

In 1 minute

  1. A wish is “looks good and works.” A criterion is “the test message arrives.”
  2. A good criterion can be answered with yes or no by someone who didn’t do the work.
  3. The criterion goes in the request, along with what must not change.

1A wish can’t be checked; a criterion can.

The acceptance criterion turns a wish into something you can check. "Easy to find" becomes "the Contact link appears in the menu on every page."

Rafael rewrote the request for the clinic’s form. He removed "beautiful" and "working." He added three sentences he can check.

Wish

"A beautiful, working contact form."

Criterion

1. The test message arrives in the front desk inbox.
2. The phone field rejects letters.
3. On a phone, the page doesn’t scroll sideways.

Net gain: three things anyone can check in five minutes.

2The yes-or-no test

Read each criterion and ask: could someone who didn’t do the work answer yes or no? If the answer is "it depends," the criterion is still a wish.

Carla wrote "clear report." It didn’t pass the test. She changed it to "each city appears once, with its total number of requests for the week."

Carla’s sheet · weekly report
1 Total requests match the sum of the week’s spreadsheets
2 Each city appears once, with its total
3 Below each table, the name of the spreadsheet it came from
  1. 1Check by adding them up.
  2. 2Check by reading the list of cities.
  3. 3Check by reading the line below each table.

Stuck here? That's normalCan't come up with any criteria? Think about how the work could go wrong for the client. Each possible error, turned around, becomes a criterion.

3Put the criterion in the request

Writing the criterion just for yourself helps. Adding it to the request helps more: the AI now works to meet it. And you ask it to say how it will check each one.

Rafael added the three criteria to the end of the request. The response began listing how to check each one instead of just saying "tested."

Code assistant

RafaelCreate the clinic’s contact form. Acceptance criteria: 1. The test message arrives in the front desk inbox. 2. The phone field rejects letters. 3. On a phone, the page doesn’t scroll sideways. At the end, say how to check each criterion.

AIForm created. To check it: 1. Send a test message and check the front desk inbox. 2. Type letters in the phone field: sending should be blocked. 3. Open the page on a phone and try to scroll sideways.

The AI gives you a checklist, not just "working."

CarlaCombine the week’s delivery spreadsheets into a new report. Acceptance criteria: 1. Total requests match the sum of the spreadsheets. 2. Each city appears once, with its total. 3. Below each table, the name of the spreadsheet it came from. At the end, say how to check each criterion.

AIReport created. To check it: 1. Add up the Orders column in the spreadsheets and compare it with the report total. 2. Read the city list and look for duplicate names. 3. Check that each table has the spreadsheet name underneath it.

The same way of making a request works for a report.

Tap or click the two buttons in the box. The worksheet is for you to follow; it does not replace checking the results.

4Also say what must not change

A good request has an outer boundary: what the AI must not touch. Without it, the AI “improves” things nobody asked for, and checking takes longer.

Carla added: “don’t change the original spreadsheets; save the report in a new file.” She no longer needs to check whether the source data changed.

Inside

Create the weekly report in a new file, using the three criteria.

Outside

Do not change the original spreadsheets. Do not change the report format.

“Inside” and “Outside” go in the same request. “Outside” reduces what you need to check.

Practice now 0/3

Three criteria for one of your tasks

You’re done when you have three criteria that pass the yes-or-no test and one line saying what’s out of scope. About 10 minutes, on paper or in a notes app.

You only write; nothing is sent to anyone. If you don’t have a coding task, use one involving text or a spreadsheet: the criterion works the same way.

Task: <what you will ask the AI to do>
Acceptance criteria:
1. <something you can check with yes or no>
2. <another one>
3. <another one>
Outside: <what the AI must not touch>

Carla’s example:
Task: weekly delivery report
1. Total orders matches the sum of the week’s spreadsheets
2. Each city appears once, with its total
3. Under each table, the name of the spreadsheet it came from
Outside: do not change the original spreadsheets

You’re already turning a wish into criteria someone else can check.

Lesson cheat sheet

Acceptance criterion

  1. Beforewrite the criterion before asking for the work.
  2. Yes or nosomeone who didn’t do the work must be able to answer.
  3. In the requestpaste the criteria and what’s out of scope.

Your next step

You already write down what “done” means before you start.

Use the three criteria from the practice in your next real request. Paste them at the end and ask for the checking steps.

Next lesson: without a limit, criteria can turn into endless work. How do you set a real limit?

Lesson 3 · Dev with AI v6.2 · INEMA.CLUB

Module 1 · Lesson 4 of 5

“Stop after two hours” is a request, not a limit

A developer places a kitchen timer and a paper card beside the laptop before starting an AI task.

You can replace a vague limit with three limits you can check: attempts, files, and time or cost.

A long task with AI uses your time and your account’s usage. A phrase like “stop after two hours” sounds like a limit, but the AI might not follow it. The limit that protects you is one that a person or something else can check.

In 1 minute

  1. A sentence in the request is a request. A limit is what stops the work even if the AI doesn’t want to.
  2. Three verifiable limits: attempts, files, and time or spending.
  3. When it hits the limit, the AI stops and describes what blocked it. It doesn't keep trying.

1A sentence in the request isn't a safeguard

The AI reads "stop in two hours" like any other sentence. It may not have a clock available, and it may think there's still time. The sentence helps, but it doesn't guarantee anything.

The risk grows in automatic mode, where the assistant approves its own actions and continues without asking for your confirmation. Each step uses your account.

Rafael left the assistant in automatic mode to fix the form, "in no more than one hour." He came back from lunch, and it was on its tenth attempt.

Request

"Fix the form submission. Stop in one hour."

What happened: ten attempts, and the work continued past the hour.

Limit and safeguard

In the request: "no more than two attempts, only the form file." Outside the request: he stayed nearby, with a 30-minute alarm.

What happened: it stopped on the second attempt and described what was missing.

2Three limits you can check

Each limit answers a question you can check afterward. How many attempts did it make? Which files did it change? How much time or usage did it spend? An attempt is each time the AI says it will try a different way.

Carla had the AI work in a folder of copies. Then she opened the folder with the original spreadsheets in Details view (List on a Mac) and checked the "Date modified" column: none of the dates were today.

Task limits
1 Attempts: no more than 2
2 Files: only those listed in the request
3 Time or spending: a safeguard outside the AI
  1. 1Check by counting how many times the AI said in the conversation that it would try a different way.
  2. 2Check the folder's "Date modified" column.
  3. 3Check the clock and the account usage page.

3The safeguard is outside the AI

You can check a limit in the request afterward. A safeguard stops the work right away. For time, if you stay nearby, an alarm is enough: when it rings, click the assistant's stop button (a square, in the app and extension) or press Esc in the terminal.

To check spending, go to the site where you subscribed, open your account settings, and look for "Usage." That's where you measure it. A spending safeguard exists only when the account allows you to set a cap, which is common with pay-as-you-go accounts.

Carla found her account usage page and noted how much a report task used. Now she knows how many tasks fit in a week.

Time safeguard

Set a 30-minute alarm. When it rings, click the stop button and ask for a report of what was done.

Measure spending

Open the account usage page before and after the task. Set a spending cap if the account offers one.

The alarm depends on you being nearby. To stop without you, only the automatic safeguard in the card below will work.
Automatic safeguard for terminal users

In the Linux terminal, the timeout command ends another program when time runs out. On a Mac, it comes with the coreutils package (brew install coreutils) and is called gtimeout.

Terminal · Linux
$ timeout 30m claude -p "Corrija o envio do formulário. No máximo 2 tentativas. Altere só contato.html."
...
$ echo $?
124

124 means time ran out and the program was stopped. Another number is the program's own response: zero if it finished successfully.

With Codex, replace claude -p with codex exec. Stopping in the middle can leave a file half finished: check what changed before continuing.

Stuck here? That's normalCouldn't find your account's usage page? Start with the alarm only. Look for the usage page later, when you have time, just once.

4You hit the limit: stop and describe what's blocking you

The limit only works if you say what to do when it is reached. Ask it to stop and describe what it tried, what failed, and what it would need to continue.

Carla received a clear warning after two attempts: one spreadsheet had its columns in a different order. She fixed the spreadsheet in two minutes.

Code assistant

CarlaCombine the spreadsheets in the "entregas-copia" folder. No more than 2 attempts. If you can't, stop and say what you tried, what failed, and what's missing.

AII stopped after 2 attempts. I tried combining the spreadsheets based on the column order. It failed: one spreadsheet has the "cidade" column in a different position. Missing: tell me whether I can reorder the columns in that spreadsheet in the copy.

A described blocker is worth more than a tenth attempt.

The "entregas-copia" folder is a copy Carla made earlier: right-click the original folder › Copy, then Paste beside it.

Practice now 0/3

Put three limits in your request

You're done when your request has an attempt limit, a list of files, a stop instruction, and the agreed alarm. About 10 minutes, in a notes app.

You only write the request; you don't need to send it now. Didn't do lesson 3? Use any task you would ask AI to do this week.

<your request, with acceptance criteria>

Limits:
- No more than 2 attempts.
- Change only: <file or folder>.
- If you hit a limit, stop and say what you tried, what failed, and what's missing.

You already know how to turn "don't take too long" into limits you can check.

Lesson cheat sheet

Real limits

  1. A request isn't a lockthe sentence helps, but it doesn't stop anything.
  2. Three limitsattempts, files, and time or cost.
  3. Stopwhen it hits the limit, AI describes what's blocking it.

Your next step

You already know how to set a lock that works even if AI doesn't agree.

Go to your AI assistant's website, find the account usage page, and save the link to your bookmarks. It takes five minutes.

Next lesson: put the task, criteria, and limits in one file, and see the whole cycle.

Supplementary material · Goal with a stop conditionFurther reading on this topic. Not included in the lesson time.

Where this lesson comes from

The Use Both kit, an open guide to using Claude and Codex together, calls this a "goal with a stop condition": an objective with observable success (which tests, what result) and limits on files, attempts, and cost. Asking "stop in two hours" is a request, not a guaranteed limit.

In the video that comes with the kit, the author, Mark Kashef, says that Codex’s goal mode (where it pursues a goal on its own) takes 3 to 5 times more time and tokens. This is his account, with no published measurement.

What INEMA’s synthesis recommends

Set verifiable limits for long tasks. Check the prices, quotas, and tools available in your account before using paid resources. Prices and limits change; what matters is what your account shows that day.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 4 · Dev with AI v6.2 · INEMA.CLUB

Module 1 · Lesson 5 of 5

The five-step cycle and your pilot briefing

A developer and an operations analyst pin paper cards in sequence on a corkboard, connecting them with arrows to map out a task.

You can put together your pilot briefing: the goal, what’s in and out of scope, acceptance criteria, and limits, all in one file.

In the previous lessons, you saw the pieces separately: checking, cross-review, criteria, and limits. On their own, they get lost from one request to the next. Together in one file, they become the start of a method that any model can follow.

In 1 minute

  1. The cycle: define, plan and critique, build and review, verify, record.
  2. It all starts in one file: the briefing that both models read.
  3. The pilot is a small task, lasting one or two hours, not the entire system.

1The whole cycle on one sheet

This course follows five steps, one module for each. You’ve already practiced the first: define the task. The others come in the next modules.

The last step is the handoff: a record that lets the next conversation continue without you explaining everything again.

Rafael put the five steps on the office wall. Every AI task he does for a client goes through them.

The course cycle
1 Define: goal, scope, criteria, limits
2 Plan and critique: one writes the plan, the other challenges it
3 Build and review: one does the work, the other reads what changed
4 Verify: tests, actual result, your decision
5 Record: handoff for the next conversation
  1. 1Module 1, this one.
  2. 2Module 2.
  3. 3Module 3.
  4. 4Module 4.
  5. 5Module 5. Module 6 measures the cost of everything.

2Who does what is an initial route

A common starting point: Claude writes the plan and Codex critiques it. Then one builds and the other reviews. This is one way to start, not a ranking of which model is better.

Only have one assistant? Start by having it do the work, then open a new conversation with it to critique. The method is the same.

Carla has Codex and an AI chat. She uses Codex to do the work and the chat, in a new conversation, to critique it.

Starting route

Plan: Claude writes, Codex critiques.
Build: one does the work, the other reviews what changed.

What matters

The quality of the work for your task. Models, access, and billing vary from account to account.

Start with the starting route and stick with whichever gives you the best result for your task.

3The briefing is the file everyone reads

The briefing brings together in one text file what you wrote in lessons 3 and 4. Both models receive the same file. That way, no one works from a summary.

The .md extension means plain text in Markdown format: # marks headings. Assistants read this format well.

Rafael saves briefing.md in the client's project folder. In Claude Code and Codex, he types @ in the request and selects the file. In a chat, he pastes the text.

briefing.md · clinic form
1 Goal: patients send messages through the website
2 Included: contact page · Excluded: home page and colors
3 Acceptance: message reaches reception; phone number rejects letters; opens on a phone
4 Limits: 2 attempts; contact.html only; 30-minute timer each time the AI works
  1. 1One sentence: what it is for.
  2. 2What can and cannot change.
  3. 3The criteria from lesson 3.
  4. 4The limits from lesson 4.

4The pilot: a real, small task

The entire course revolves around a pilot of your own. Choose a real, small task that fits into one or two hours of your work. If it's too big, the cycle gets lost. If it's made up, no one checks it for real.

Carla chose the routine that combines the week's delivery spreadsheets. It's small, she does it every week, and she knows how to tell when it's right.

Too big

"Automate all the company's reports."

Good pilot

"Combine the week's delivery spreadsheets into a report, checked against the total."

Stuck here? That's normalCan't think of a task that's small enough? Take one piece of a larger task: just one page, one report, or one formula.

Practice now 0/4

Write your pilot briefing

Done when the briefing.md file has all four parts filled in. About 12 minutes, on a computer.

It's your text file; nothing is sent. If you don't know a limit, write "to be determined" and return to lesson 4. No computer right now? Write it in your phone's notes app and create the file later.

# Brief: <pilot name>

## Objective
<one sentence: what it is for>

## In scope and out of scope
In scope: <what can change>
Out of scope: <what cannot change>

## Acceptance criteria
1. <yes or no>
2. <yes or no>
3. <yes or no>

## Limits
- Attempts: <no more than 2>
- Files: <which ones>
- Time or cost: <where the limit is>
Create the file in Windows
1 Pilot folder
2 View › Show › File name extensions
3 Right-click › New › Text Document
4 briefing.md
  1. 1Open the folder you created for the pilot.
  2. 2Turn on file extensions. In Windows 10: select the View tab, then check File name extensions.
  3. 3Right-click an empty space in the folder.
  4. 4Rename it to briefing.md and confirm. To open it: right-click › Open with › Notepad.
On Mac: TextEdit › Format › Make Plain Text, paste the template, and save as briefing.md; if prompted, choose "Use .md". Having trouble with the extension? A briefing.txt works too.

You already have the briefing that both models will read in module 2.

Lesson cheat sheet

Cycle and briefing

  1. Five stepsdefine, plan and critique, build and review, verify, document.
  2. Briefa file with the objective, what is in scope and out of scope, criteria, and limits.
  3. Pilota real, small task that you know how to check.

Your next step

You finished module 1 with your pilot briefing ready.

Reread the briefing tomorrow with a fresh mind. Revise one criterion that still depends on opinion.

In module 2: one model writes the plan based on your briefing, and the other tries to find flaws in it.

Lesson 5 · Dev with AI v6.2 · INEMA.CLUB

Module 2 · Lesson 1 of 5

The plan lives in a file, not in the conversation

A developer marks up a home's technical drawing in pencil on a blue sheet spread across the table before construction begins.

You can ask AI for a plan-v1.md with scope, assumptions, acceptance criteria, and risks, without it changing anything yet.

When the plan stays in the middle of the conversation, it gets mixed in with questions, corrections, and code. No one can critique it as a whole. In a file, the plan becomes something another AI and you can read from start to finish.

In 1 minute

  1. Ask for the plan in a file: plan-v1.md, in the pilot folder.
  2. Four parts: scope, assumptions, acceptance criteria, and risks.
  3. Write "don't implement anything yet." Critiquing a plan costs less than undoing code.

1 A plan in the conversation disappears; a plan in a file stays

The plan is the first thing the AI writes. If it only appears in the chat, it disappears when you scroll. In a file, it stays in the folder, with a name and version.

That way, the second model reads exactly the same text you read.

Rafael asked for a plan for the clinic form. The first time, it came in the chat, mixed in with code. The second time, he asked for a file.

Code assistant

RafaelPlan the clinic's contact form.

AISure. First I'll create the form. I've already created the contato.html file with the fields. Next…

The plan turned into implementation halfway through the response. There was nothing left to critique.

RafaelRead briefing.md. Write the plan in plan-v1.md, with scope, assumptions, acceptance criteria, and risks. Don't implement anything yet.

AII saved plan-v1.md in the pilot folder. I didn't change any other files.

A file anyone, or another AI, can read in full.

Tap or click the two buttons in the panel and compare the two requests.

2The four parts of a plan

A plan you can critique has four parts. Scope: what is included and what is left out. Assumptions: what the AI is taking for granted. Acceptance criteria: how to check. Risks: what could go wrong.

The acceptance criteria come from your briefing. The plan says how each one will be tested.

In Rafael's plan-v1.md, the risks section warned that the clinic's hosting might block email delivery.

plan-v1.md · clinic form
1 Scope: contact page with name, "phone field, free-text field" and message; "send to the reception email". Excluded: home page and colors
2 Assumptions: reception uses a single email address
3 Acceptance criteria: the 3 from the briefing, each with its test
4 Risks: hosting might block email delivery
  1. 1What is included and what is left out.
  2. 2What the AI took for granted.
  3. 3How each criterion will be checked.
  4. 4What could go wrong.

3"Don't implement anything yet"

Coding assistants like to get started right away. Without the phrase "don't implement anything yet," the plan turns into implementation. And it's more costly to undo the wrong implementation than to fix the wrong plan.

The sentence is a request. Claude Code has a real lock: planning mode, which prevents you from editing files. Press Shift+Tab until it appears; the plan returns in the conversation, and you save it yourself in plan-v1.md.

Carla asked for a plan for the spreadsheet routine and forgot the sentence. The assistant had already created three new files in the folder. She had to delete everything and ask again.

Without the sentence

Request: "Plan the spreadsheet routine."

Result: three files created before anyone read the plan.

With the sentence

Request: "Write the plan in plan-v1.md. Don't implement anything yet."

Result: one file, ready to read in five minutes.

4Written assumptions are errors you can find

The assumptions section is the most valuable. Errors often hide in what no one wrote down. When AI writes down what it took for granted, you and the reviewer can disagree.

Carla's plan-v1.md said: "all the spreadsheets have the same columns, in the same order." She knew that wasn't true for one city. The error surfaced before any attempt.

Hidden assumption

The plan says to "combine the spreadsheets by column order." No one notices what this assumes.

Written assumption

"I assume all the spreadsheets have the same columns, in the same order." Carla reads it and corrects it right away.

Stuck here? That's normalWas the plan too long? Ask: "summarize it in one page, keeping the four parts". A short plan is easier to critique.

Practice now 0/3

Ask for the plan-v1.md for your pilot

You're done when there's a plan-v1.md in the pilot folder with the four parts and no other new files. About 10 minutes, on your computer.

The request says not to implement anything. If the assistant creates other files anyway, stop it (Esc key in Claude Code; stop button in the app) and delete only what it created. Didn't do lesson 5? Write three lines of briefing in the request itself.

Read briefing.md. Write the plan in plan-v1.md, in the same folder.
Include four parts: scope (in and out), assumptions,
acceptance criteria (each with the test that will check it), and risks.
Don't implement anything yet. Don't change any other files.

You already have a plan in a file, ready for another AI to critique.

Lesson cheat sheet

Plan in a file

  1. plan-v1.mdthe plan stays in the pilot folder, not in the conversation.
  2. Four partsscope, assumptions, acceptance criteria, and risks.
  3. No implementation"don't implement anything yet" goes in the request.

Your next step

You can now turn "plan this" into a file that can be critiqued.

Read your plan-v1.md quietly, just the assumptions section. Write down one question you'd ask the person who wrote it.

In the next lesson: who will review this plan without mercy, and how to ask for useful criticism.

Supplementary material · Where plan-v1.md comes fromFurther reading on the topic. Does not count toward lesson time.

The Use Both kit request

The first flow in the Use Both kit, “plan and then challenge,” starts like this: create plan-v1.md for the task, without implementing anything, including scope, assumptions, acceptance criteria (each with its own test), and risks. The same file then goes to the other model for a read-only critique. The numbered version in the name lets you keep the original plan when the revised version arrives.

Sources: Use Both Guide · Codex + Claude Area — INEMA Events

Lesson 6 · Dev with AI v6.2 · INEMA.CLUB

Module 2 · Lesson 2 of 5

Read-only critique: finding with evidence

An operations analyst reads a printed document and writes notes in the margin with a red pen, pointing to specific passages.

You can ask another AI to critique your plan so that each finding includes a problem, evidence, impact, and the smallest fix.

“Looks good” and “you could consider more cases” don't help anyone. A critique is useful only when it points to the passage, explains what happens if nobody fixes it, and proposes the simplest fix.

In 1 minute

  1. The reviewer only reads. They don't change the plan or any other file.
  2. Each finding has four fields: problem, evidence, impact, and smallest fix.
  3. The reviewer gets the real briefing and plan, never a summary.

1The reviewer only reads

In a critique, you want findings, not a new plan written over yours. That's why the request says: read, point things out, don't change files. That sentence is still a request; the real safeguard that prevents changes comes in lesson 8.

The reviewer can be another model or a new conversation with the same assistant. This applies the cross-review from lesson 2 to the plan.

Rafael sent plan-v1.md from the form to Codex with a request for a read-only critique. He received three findings, and no files were changed.

Reviewer

RafaelRead briefing.md and plan-v1.md. List the findings in the plan. For each finding: problem, passage from the plan, impact, and the smallest fix. Don't change files.

CodexFinding 1. Problem: the plan doesn't say what happens if sending fails. Passage: “send to the reception email.” Impact: the patient sees “sent,” and the message gets lost. Smallest fix: show an error message and record the failure.

A finding Rafael can check and fix without asking anything.

2The four fields of a finding

A flaw is what is wrong or missing. Evidence is the excerpt that shows it. Impact is what happens if no one fixes it. The smallest fix is the simplest change that solves it.

Also ask for a severity: high, medium, or low. That way, you know where to start.

Carla asked the AI chat to critique the spreadsheet plan in a new conversation. She asked for the findings in order of severity.

review.md · finding 1 from Carla’s plan
1 Flaw: nothing handles a spreadsheet with columns in a different order
2 Evidence: "combine the spreadsheets by column order"
3 Impact: totals for one city are added in the wrong column
4 Smallest fix: combine by column name
  1. 1What is wrong.
  2. 2Where it is, using the plan’s own words.
  3. 3Why it matters.
  4. 4The simplest fix.

3A vague finding is not a finding

Without all four fields, the critique becomes generic advice. You don’t know whether you agree because you don’t know what it’s about.

The first critique Rafael received, without the template, said "consider validating the fields more carefully." With the template, it became a finding about the phone number.

Vague

"Consider validating the form fields more carefully."

Finding

Failure: the plan doesn’t handle phone numbers with letters. Evidence: “phone field, free text.” Impact: violates criterion 2 in the briefing. Smallest fix: accept numbers only.

The finding cites the briefing criterion that would be violated.

Stuck here? That's normalDid the reviewer return findings that seem wrong? Great: you don’t have to accept them all. In lesson 9, you’ll learn to respond to each one with "accepted," "rejected with evidence," or "open."

4The reviewer gets the briefing and the plan

Without the briefing, the reviewer judges the plan based on personal preference. With the briefing, the reviewer judges it based on what you asked for. Send both actual files.

Carla pasted the briefing into the chat, then the plan, and only then the request for a critique. In the coding assistant, it’s enough to mention both files by name.

Plan only

The reviewer suggests replacing the spreadsheet with a database. Out of scope.

Brief and plan

The reviewer points out that the plan forgot the criterion "each city appears once".

Practice now 0/3

Ask for a critique of your plan-v1.md

You’re done when you have a review.md with at least two findings in all four fields. About 10 minutes, on a computer.

This request is read-only: no files change. Use another assistant or a new conversation with the same one. Haven’t done lesson 6? Ask for a critique of any plan you have as text.

Read briefing.md and plan-v1.md. Critique the plan by reading only.
List the findings. For each finding, write:
- Flaw:
- Evidence (the exact excerpt from the plan):
- Impact (what happens if no one fixes it):
- Smallest fix:
- Severity (high, medium, or low):
Order them by severity. Do not change any files.

You now get feedback you can verify, not generic advice.

Lesson cheat sheet

Useful feedback

  1. Read-onlythe reviewer points things out; it does not rewrite or change files.
  2. Four fieldsissue, evidence, impact, and smallest fix.
  3. Shared contextthe briefing and plan go together to the reviewer.

Your next step

You now ask for feedback with evidence, not opinions.

Save the practice template in a file called pedido-critica.md in the pilot folder. You’ll use it again in module 3.

Next lesson: call the other model without copying and pasting, with a single command.

Lesson 7 · Dev with AI v6.2 · INEMA.CLUB

Module 2 · Lesson 3 of 5

Connect Claude and Codex: three ways and the manual method

A developer, seated between a laptop and a larger monitor, passes a printed sheet from one side of the desk to the other.

You can ask another assistant to review the plan with one command, or use the manual method, then open the review.md that comes back.

Copying the plan from one window and pasting it into another works, but it gets tiring and leads to mistakes: a section is missing, another one is extra. There are ways for one assistant to call the other and have the response go straight into a file.

In 1 minute

  1. Three automatic ways: the official plugin, a command that calls the other assistant, and a review of a published change. Plus the manual method.
  2. With the command, the reviewer runs in read-only mode, and the response goes into a file.
  3. No terminal? Copying and pasting into a new conversation still works.

1Three ways to connect the two, plus the manual method

Claude Code and Codex can communicate in three automatic ways. The first is an official plugin from Codex for Claude Code. The second is a command that calls the other assistant and saves the response in a file.

The third is a review of a published change, which appears in module 3. All methods, including the manual one, follow the module’s rule: the reviewer gets the briefing and the actual file.

Rafael uses the second method: one command, run in the pilot folder.

Three ways and the manual method
1 Official plugin: openai/codex-plugin-cc
2 One command calls the other; response in review.md
3 Review of a published change (module 3)
4 Manual: copy and paste into a new conversation
  1. 1Works as an add-on in Claude Code.
  2. 2Runs in the terminal. Requires the other assistant’s CLI, connected to your account.
  3. 3Useful when code has already been changed.
  4. 4Works with any chat.

2Codex critiques Claude’s plan with a command

In the terminal, inside the pilot folder, the codex exec command makes a request and exits. It is Codex’s CLI.

The --sandbox read-only option puts Codex in a read-only sandbox. The -o option saves the last response to review.md.

Rafael’s pilot folder does not use Git yet. So he adds --skip-git-repo-check, which Codex requires when you run it outside a Git repository.

Terminal · pilot folder
$ codex exec --skip-git-repo-check --sandbox read-only -c model_reasoning_effort=medium -o review.md 'Leia briefing.md e plan-v1.md. Liste os achados do plano, cada um com falha, evidência, impacto e a menor correção. Não altere arquivos.'
...
$ ls
briefing.md  plan-v1.md  review.md

ls shows that only review.md appeared. The plan is still the same.

-c model_reasoning_effort=medium sets the effort to medium. Module 6 covers this again.

3The reverse: Claude critiques

If Codex wrote the plan, Claude can review it. The claude -p command makes a request, prints the response, and exits. The > symbol saves that response to a file.

The safeguard is the --permission-mode plan option: it is planning mode, where Claude reads the files but does not change any of them. Without it, “don’t change files” would just be a request.

Rafael reversed the roles in a second project: Codex planned, and Claude reviewed.

Terminal · pilot folder
$ claude -p --permission-mode plan 'Leia briefing.md e plan-v1.md. Liste os achados do plano, cada um com falha, evidência, impacto e a menor correção. Não altere arquivos.' > review-claude.md

Nothing appears on screen: the entire response went to review-claude.md.

To choose a model and effort level, add --model and --effort; the names depend on your account. Module 6 covers this again.

Stuck here? That's normalNever used the terminal? Use the manual path, item 4 on the chart: open a new conversation, paste the briefing, the plan, and the critique request from lesson 7, then save the response as review.md. The result is the same.

4Manual or command: what changes

The critique’s content does not change. The hands-on work and the risk of pasting the wrong thing do. With the command, the reviewer reads the files directly from the folder.

Carla uses Codex through the app and an AI chat. She chose the manual path and saved the critique request in a file so she can always paste the same thing.

Manual

Start a new conversation, paste the briefing, plan, and request. Save the response as review.md. It works in any chat.

By command

One command in the folder. The reviewer reads the files, and the response goes into review.md. Less copying, fewer paste errors.

Both approaches work. Choose the one you already use.

Practice now 0/3

Get feedback on the plan in a file

Done when a review.md exists alongside plan-v1.md and the plan is unchanged. About 12 minutes, on a computer.

The reviewer runs in read-only mode. If the command returns an error, check that you’re in the pilot folder and that the assistant is connected to your account. If that doesn’t fix it, use the manual approach. To open the terminal in the folder: on Windows 11, right-click the folder › Open in Terminal; on Mac, type cd and a space, then drag the folder into the window. On Windows, use PowerShell: single quotes don’t work in Command Prompt. Haven’t done lessons 5 and 6? Use any plan in text.

codex exec --skip-git-repo-check --sandbox read-only -o review.md 'Read briefing.md and plan-v1.md. List the findings about the plan, each with the issue, evidence, impact, and smallest fix. Do not change any files.'

Only have Claude Code? Use this:

claude -p --permission-mode plan 'Read briefing.md and plan-v1.md. List the findings about the plan, each with the issue, evidence, impact, and smallest fix. Do not change any files.' > review.md

You now get feedback from another assistant without changing the plan.

Lesson cheat sheet

Connect the two

  1. codex execwith --sandbox read-only and -o review.md.
  2. claude -pwith --permission-mode plan and > review-claude.md.
  3. Manualnew conversation, paste everything, save the response.

Your next step

You now have one assistant review the other, with the response in a file.

Read review.md and mark next to each finding whether you agree. It takes ten minutes.

Next lesson: what if the two disagree forever? When to end the debate.

Supplementary material · Commands and the pluginMore on this topic. Not included in the lesson time.

What the sources say

The Codex + Claude section of Eventos INEMA lists three ways to connect the two: the official Codex plugin for Claude Code (openai/codex-plugin-cc), one CLI calling the other with the response saved in a .md file, and a pull request (a published change for review) that the other model reviews. The two commands in this lesson were checked in the INEMA environment; options change with new versions, so run codex exec --help and claude --help to see the ones available on your machine.

Why --skip-git-repo-check

By default, codex exec only runs inside a folder with version control. Outside one, this option allows it to run. In module 3, the pilot folder will have version control, so the option won’t be needed.

Sources: Codex + Claude Area — INEMA Events

Lesson 8 · Dev with AI v6.2 · INEMA.CLUB

Module 2 · Lesson 4 of 5

Two review rounds and that's it

An operations analyst closes a folder of documents with two sticky notes attached to the cover, with the gesture of someone who has decided to end the discussion.

You can respond to every finding in the critique, decide when to end the discussion, and know when it's worth automating it.

The plan and critique can go back and forth forever. Each round takes time and uses up your account, and after a point they just agree out of exhaustion. Tests, not another round, decide what stays unresolved.

In 1 minute

  1. Each finding gets a response: accepted, rejected with evidence, or unresolved.
  2. At most two review rounds. Agreement isn't proof.
  3. What stays unresolved becomes a named test. For a large plan: claudex automates it.

1Each finding gets a response

A review round is a back-and-forth: the critique arrives, and each finding gets a response. Accepted: it goes into the plan. Rejected with evidence: you show why it doesn't. Unresolved: no one knows yet.

Carla responded to all three findings in the spreadsheet plan. One accepted, one rejected, and one unresolved.

review.md · Carla's responses
1 Match by column name → accepted
2 Replace the spreadsheet with a database → rejected: outside the briefing
3 Spreadsheet with a blank city → unresolved
  1. 1It goes into the plan.
  2. 2It doesn't go in, and the evidence is written down.
  3. 3It becomes a test, in step 3 of this lesson.

2Why stop after two rounds

The first round finds most of the problems. The second checks whether the fixes are good. After that, each round finds less and costs the same. The two models also start agreeing, which lesson 2 showed isn't proof.

Rafael let the discussion about the form go to a fifth round. The last three only rearranged the wording.

Five rounds

Rounds 3, 4, and 5: wording changes. Both assistants finish "in agreement." No named tests.

Two rounds

Round 1 finds problems. Round 2 checks the fixes. What's left becomes a test before you build.

Test yourself

After the second round, both models agree that the plan is ready. What does that guarantee?

3What stays unresolved becomes a named test

An unresolved finding isn't settled by discussion. It becomes a test written into the plan, to run when there's something to test. That way, the discussion ends, but the question doesn't disappear.

The finding about a blank city became a test. Run it with the Juiz de Fora spreadsheet, which has a blank city, and check where the request appears.

Assistant

CarlaFinding 3 is still open: I don’t know what to do with a blank city field. Turn it into a test for the plan, without deciding for me.

AIProposed test: run the routine with a spreadsheet that has a row with a blank city field. Check whether the request appears in the overall total and where it is listed. You decide what the right behavior is.

The test says what to check. The business rule is still Carla’s decision.

4Large plan: claudex automates the debate

For a small plan, two manual rounds are enough. For a large plan, there’s claudex, a plugin for Claude Code. With the command /claudex:plan --rounds 2 <topic>, Claude writes PLAN.md and Codex critiques it, with a different reviewer each round. The cycle repeats until approval or until it reaches the number of rounds. The default is 3; set it to 2 to follow this lesson’s rule.

Rafael uses the two manual rounds for the form. He saved claudex for the clinic’s much larger appointment scheduling system.

Small task

Two manual rounds. You read each finding and respond.

Large plan

/claudex:plan --rounds 2, with the number of rounds set. At the end, you still read the plan and name the tests.

Automating the debate doesn’t automate the decision.

Stuck here? That's normalDon’t use Claude Code? Ignore claudex for now. The two manual rounds work with any assistant, even in a new conversation with the same one.

Practice now 0/3

Respond to the findings

You’re done when you’ve responded to every finding in the case and turned the open one into a test. About 8 minutes, on paper or in a notes app.

This is a practice case. If a finding seems both accepted and rejected, mark it open and write the test that would decide.

The case. Rafael’s briefing asks for a contact form for the clinic, only on the contact page. The critique found three issues. A: “the plan doesn’t limit the message length.” B: “change the colors on the home page to make the contact option stand out.” C: “it’s unclear whether the front desk wants copies of messages on a phone.”

Show answer key

A: accepted; the size limit is part of the plan. B: rejected; the briefing puts the home page and colors under “Out of scope.” C: open; turn it into the question “Does reception want a copy on a cell phone?” for the clinic to answer before building.

You can end the debate once you’ve responded to every finding and captured the open question in a test.

Lesson cheat sheet

End the debate

  1. Three responsesaccepted, rejected with evidence, open.
  2. Two roundsafter that, agreement doesn’t add proof.
  3. Openbecomes a named test; big plan, claudex.

Your next step

You already know how to close a review without getting stuck in an endless debate.

Next to each finding in your review.md, write: accepted, rejected, or open. It takes ten minutes.

In the next lesson: combine the responses in the pilot's plan-v2.md and see why “Claude plans” is only a beginning.

Supplementary material · Reconcile and stopFurther reading on this topic. Doesn't count toward the lesson time.

What the Use Both kit asks you to do

In the “plan, then challenge” flow, the kit asks you to reconcile findings as accepted, rejected with evidence, or open (the kit says “unresolved”), save plan-v2.md and review.md, and limit review to two rounds. And remember: agreement is not proof; name the tests or evidence needed before implementing.

claudex × Use Both

According to the Codex + Claude area of Eventos INEMA, the two are complementary. claudex is software that automates debate about a big plan. Use Both is a method for everything else: reviewing a change, setting a goal with a stopping condition, switching sessions.

Sources: Use Both Guide · Codex + Claude Area — INEMA Events

Lesson 9 · Dev with AI v6.2 · INEMA.CLUB

Module 2 · Lesson 5 of 5

Starting path, not a ranking, and the pilot's plan-v2.md

A developer and an operations analyst compare a document full of pen marks and the clean new version side by side on the desk.

You can finish your pilot's plan-v2.md, with a response to every finding and a test linked to every acceptance criterion.

Receiving the critique is half the journey. The other half is combining what was accepted into a new version. Without losing the original plan, and without forgetting what remains open. And without treating the kit's path as absolute truth.

In 1 minute

  1. “Claude plans, Codex critiques” is a path to get started, not a ranking.
  2. plan-v2.md is a new file; plan-v1.md stays saved alongside it.
  3. Before building, every acceptance criterion in the briefing needs a test in the plan.

1The route table is a starting point

The Use Both kit includes a table showing who does what. For planning: Claude writes and Codex critiques the assumptions. The finish line is a plan with scope and acceptance criteria, each with its own test. These are initial preferences recorded by the author, not a measured ranking of models.

Rafael started with the path in the table. On the second project, he switched the roles to see which critique found more real problems in his task.

Who does what · starting path
1 Plan: Claude writes · Codex critiques the assumptions
2 Build: Claude builds · Codex reviews what changed
3 Review a document: either one writes · the other checks
  1. 1This module.
  2. 2Module 3.
  3. 3Works for text, reports, and plans.
Keep the route that produces better evidence for your task.

2The evidence from your task is what decides

Models, access, and billing change from account to account and from month to month. So the useful question isn’t “which one is best?” It’s “for this task, which review found real issues, with evidence?”

Carla compared two reviews of the same plan: one from Codex, the other from a new chat conversation. She counted how many findings had evidence she checked.

By reputation

“People say model X is the best for planning.” Carla would use only that one, without checking.

By the evidence

Review 1: 3 findings, 2 checked. Review 2: 5 findings, 1 checked. For this task, review 1 was more useful.

3The plan-v2.md is a new file

Don’t write over plan-v1.md. The new version is a sibling file: plan-v2.md. That way, you can compare the two and see what changed because of the review.

The plan-v2.md contains the accepted findings, a list of the rejected ones with the evidence, and the open ones as tests.

Rafael’s pilot folder ended up with four files, each with a purpose.

Pilot folder · clinic form
1 piloto-formulario
briefing.md
2 plan-v1.md
3 review.md
4 plan-v2.md
  1. 1One folder per pilot.
  2. 2The original plan, saved.
  3. 3The review, with your responses.
  4. 4The version that goes into development.

4Check plan-v2.md against the briefing

Before building, check that every acceptance criterion in the briefing has a test in the plan. A criterion without a test is a criterion no one will check.

Rafael asked a new conversation to do this check. One criterion had no test.

Assistant

RafaelRead briefing.md and plan-v2.md. List each acceptance criterion from the briefing and the plan test that checks it. Point out any criterion without a test. Don’t change any files.

AICriterion 1, the message arrives at reception: real send test. Criterion 2, the phone field rejects letters: typing test. Criterion 3, it opens on a phone: no test in the plan.

The gap showed up before the first line of code.

Stuck here? That’s normalDoes your review.md have too many findings? Put only the high- and medium-severity ones in plan-v2.md. For low-severity ones, respond “open (low)” and leave them in review.md; don’t add them to plan-v2.md.

Practice now 0/4

Finish plan-v2.md for your pilot

You’re done when plan-v2.md is in the folder, next to plan-v1.md, with every acceptance criterion linked to a test. About 12 minutes, on a computer.

plan-v1.md does not change: you create a new file. Haven’t done lessons 6 to 9? Use the template with any plan and critique you have in text.

Read briefing.md, plan-v1.md, and review.md (with my responses:
accepted, rejected, or open). Write plan-v2.md as a new file:
- incorporate the accepted findings;
- list the rejected findings with the evidence;
- turn the open findings into named tests.
Do not change plan-v1.md or review.md. Do not implement anything.

You’ve closed the pilot plan: critiqued, answered, and with a test for every criterion.

Lesson cheat sheet

Plan closed

  1. Starting routeis a way to begin; the task evidence decides.
  2. Sibling filenew plan-v2.md, plan-v1.md kept.
  3. Criterion with a testno briefing criterion without a planned check.

Your next step

You’ve finished module 2 with the pilot plan critiqued and ready to become a build.

Reread plan-v2.md tomorrow and underline the first construction step. Just that one.

In module 3: someone builds from the plan, and the other reads every line that changed.

Supplementary material · The route cardFurther reading on the topic. It doesn’t count toward lesson time.

Where the table comes from

The Use Both kit’s route card, summarized in the Codex + Claude area of Eventos INEMA, presents initial preferences recorded in Mark Kashef’s video and translated into tasks. The kit itself says these are editorial choices, not a measured ranking: experiment and stick with the route that produces better evidence. The video mentions Claude Opus 5.5 and GPT-6 Astra; models, tools, limits, and charges vary by account, and the flow is more portable than those names.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 10 · AI Dev v6.2 · INEMA.CLUB

Module 3 · Lesson 1 of 5

Just enough Git: see what changed

A developer compares two nearly identical printed sheets and marks the lines that changed from one to the other in green and red.

You can create a test branch and read a git diff line by line, saying what was removed and what was added.

In module 2, someone critiqued the plan. Now AI will change real files. To review what it did, you need to see exactly what changed, not what it says changed.

In 1 minute

  1. Git stores versions of files in a folder and shows the difference between them.
  2. A branch separates the test: you can make changes there without damaging the working version.
  3. In git diff, a line with − was removed, and a line with + was added.

1Git shows what changed, not what was said

Git stores versions of files in a folder. With it, you compare the earlier version with the current one, line by line.

For reviewing AI work, this is what matters. The summary says "I changed the form." The diff shows which lines, in which files.

Rafael asked for a correction to the phone field on the clinic website. The summary mentioned one file. The diff showed two: the AI had also changed the footer.

Without Git

Question: "what did you change?"

Answer: the AI summary, based on its memory.

With Git

Question: "what changed?"

Answer: Git lists the files that changed, including new ones, and shows the lines.

Net gain: Rafael found the change in the footer, which no one had requested.

2A branch is a test line

Before the AI makes changes, create a branch. If the change isn't good, the working version stays intact.

In the terminal, inside the project folder, one command creates the branch and switches to it. git status confirms where you are.

Rafael created the formulario-contato branch before asking the AI for anything. The published version of the website stayed on the main branch.

Terminal · website folder
$ git switch -c formulario-contato
Switched to a new branch 'formulario-contato'
$ git status
On branch formulario-contato
nothing to commit, working tree clean

"Switched to a new branch" = the branch was created and you're already on it. "nothing to commit" = nothing has changed yet.

Git responds in English. Read just the beginning of each line.

3Reading the diff: − removed, + added

After the AI is done, run git status first: it lists all the files that changed, including new ones. A new file doesn't appear in git diff until you use git add, as lesson 13 shows.

Next, git diff shows the diff. Each changed file is shown with its name. A line that starts with − was removed; a line with + was added. The other lines are context, with no changes.

In Rafael's diff, one line in the phone field was removed and another took its place. He read both and saw the new rule: numbers, spaces, parentheses, and hyphens only.

Terminal · website folder
$ git diff
diff --git a/contato.html b/contato.html
index 3f2a1c0..8b7d4e2 100644
--- a/contato.html
+++ b/contato.html
@@ -14,3 +14,3 @@
   <label for="telefone">Telefone</label>
-  <input type="text" id="telefone" name="telefone">
+  <input type="tel" id="telefone" name="telefone" pattern="[0-9 \(\)\-]+">
   <label for="mensagem">Mensagem</label>

The first five lines say which file and which section. Then, − the old line and + the new one. A long diff opens one page at a time: Space moves forward, q exits.

Read the file names first: a file outside the request is the first sign of a problem.

Test yourself

In a diff, a line starts with −. What does that mean?

4No terminal: ask the assistant for the diff

If you use the assistant in the app, you don't need to open the terminal. Ask it to show the complete diff of what changed. The assistant runs the command and pastes the result.

Carla uses Codex in the app. After each request, she writes: "show the complete git diff". That way, she sees the lines, not just the sentence "I changed juntar.py".

An .xlsx spreadsheet is different: in the diff, it appears only as "Binary files differ," without the cells. For a spreadsheet, ask the assistant for a summary of the cells that changed, or work with the CSV copy, where the diff shows each line.

Code assistant

CarlaAdd up the requests by city in the juntar.py file.

AIDone. I adjusted the sum by city.

Where did you adjust it? How many lines? You can't tell.

CarlaShow me the complete git diff of what you changed.

AIdiff --git a/juntar.py b/juntar.py --- a/juntar.py +++ b/juntar.py @@ -8,1 +8,1 @@ -total = pedidos.sum() +total = pedidos.groupby("cidade").sum()

Now Carla can see the line that was removed and the one that was added.

Tap or click the two buttons in the panel and compare.

Stuck here? That's normalDoes the diff look like alphabet soup? Read just three things: the file name, the lines with −, and the lines with +. The rest is address and context.

Practice now 0/4

Your first diff, in the pilot folder

You're done when you've read a diff of your briefing.md and written down one line that was removed and one that was added. About 12 minutes, on a computer.

Everything happens in the pilot folder, with the briefing from lesson 5; haven't made one? Create any text file there. If "git: command not found" appears (on Windows: "The term 'git' is not recognized"), install Git from the official site, git-scm.com, and close and reopen the terminal. If the commit asks for a name and email, run the two git config commands it shows, replacing the example name and email with yours. Prefer not to use the terminal? Ask your assistant for each step, in the pilot folder.

git init
git add briefing.md
git commit -m "briefing inicial"
git switch -c teste-diff
git diff

You can now see what changed in a file in Git itself.

Lesson cheat sheet

Git for review

  1. git switch -ccreates a test branch and switches to it.
  2. git statustells you where you are and lists everything that changed, including new files.
  3. git diffshows the lines: − removed, + added.

Your next step

You can now see the actual change without relying on the AI summary.

For your next request to a coding assistant, add this at the end: "show the complete git diff". It takes ten seconds.

In the next lesson: if AI does everything at once, the diff turns into a book. How do you ask for it in stages?

Lesson 11 · Dev with AI v6.2 · INEMA.CLUB

Module 3 · Lesson 2 of 5

One step at a time, with the criteria right there

An operations analyst assembles a shelf in stages: secures the first shelf and checks it with a level before picking up the next piece.

You can request just the first step of your pilot, with the briefing criteria included and a stop at the end.

Requesting the whole task at once creates a huge change. No one can review three hundred lines well. In steps, each change is small, checked, and reviewed before the next one.

In 1 minute

  1. Break the plan into small steps, each with its own criterion.
  2. Request one step, with the criteria included and “stop when done.”
  3. Did you check the step? It goes to the lesson 13 review; the commit comes later.

1Doing everything at once creates a diff no one reads

When AI does everything in a single request, the diff gets long. You skim it, get tired, and approve it. That’s how errors get in.

Rafael requested “the entire form, with submission and validation.” Changes came back in seven files. The next time, he requested only the fields: one file, a few lines.

All at once

Request: complete form.

Diff: seven files. Rafael approved without reading everything.

One step

Request: just the form fields.

Diff: one file, which he read line by line.

Net gain: a small change is one you actually check.

2Each step has its own criterion

Take the plan from module 2 and split it into steps that fit in one request. Each step inherits from the briefing the acceptance criterion it must meet.

Carla split the spreadsheet routine into three steps. The first only reads the spreadsheets for the week. Its criterion: the total number of rows read matches the sum of the rows in the spreadsheets.

Steps in Carla’s pilot
1 Read the spreadsheets · criterion: total rows read match
2 Add up by city · criterion: each city appears once
3 Save the report · criterion: new file, originals untouched
  1. 1Start with what the other steps need.
  2. 2One new thing per step.
  3. 3The last step delivers the result.

3The step request: include the criteria and a stop

The request has four parts: where to read the context, which step to do, its criteria, and the instruction to stop at the end. With that stop, AI won’t continue to the next step before you check.

Rafael names the files using @ and requests step 1. The response lists what changed and how to check it.

Code assistant

RafaelRead @briefing.md and @plan-v2.md. Do only step 1: the form fields. Criteria for this step: 1. Name, phone, and message appear on the page. 2. The phone field rejects letters. Change only contato.html. At the end, stop, list the changed files, and say how to check each criterion.

AIStep 1 done. File changed: contato.html. To check: 1. open the page and see the three fields; 2. type letters in the phone field and try to submit. I stopped here. Step 2 is submission.

One file, two criteria, and the AI stopped and waited for you.

Stuck here? That's normalDon't have the plan-v2.md for module 2? Use only the briefing from lesson 5 and write step 1 yourself in the request. The rest works the same way.

4Bad step? Undo it before saving

The version saved in Git is called a commit. In this course, it comes after the review in lesson 13. Until then, the change is only in the files, and you can undo it.

If the step didn't work, git restore . returns the files Git already tracks to the last saved version. It won't remove a new file the AI created: check git status and delete it manually.

Carla didn't like the first attempt at step 1. She asked Codex in the app: "undo this step's changes with git restore and show git status". Everything went back to how it was before.

Terminal · Carla's pilot folder
$ git status
On branch etapas-relatorio
Changes not staged for commit:
	modified:   juntar.py
$ git restore .
$ git status
On branch etapas-relatorio
nothing to commit, working tree clean

"modified" = file changed and not yet saved. After git restore ., "nothing to commit" = it went back to the last saved version.

In the terminal or by asking the assistant. If you've already run the git add from lesson 13, use git restore --staged --worktree . to also undo what was staged.

Practice now 0/3

Request step 1 of your pilot

You're done when the request for step 1 has context, a step, criteria, and a stopping point, and you've checked the response. About 10 minutes, on your computer.

The AI changes only the file you list, inside the test branch from lesson 11 (didn't do it? ask the assistant: "create a test branch and switch to it"). If it goes beyond step 1, press the stop button (Esc in the terminal) and ask it for a report. In a regular chat, paste the briefing text in place of @.

Read @briefing.md <and @plan-v2.md, if you have it>.
Do only step 1: <what step 1 is>.
Criteria for this step:
1. <yes or no>
2. <yes or no>
Change only: <file>.
At the end, stop, list the changed files, and say how to check each criterion.

Carla's example:
Read @briefing.md.
Do only step 1: read the week's spreadsheets.
Criteria for this step:
1. The total number of rows read equals the sum of the rows in the spreadsheets.
2. The original spreadsheets don't change.
Change only: juntar.py.
At the end, stop, list the changed files, and say how to check each criterion.

You already ask for work in steps you can check one at a time.

Lesson cheat sheet

Steps

  1. Smallone step fits in a diff you can read all the way through.
  2. Requestcontext, step, criteria, and "stop at the end".
  3. Undogit restore . before the commit, which comes after the review.

Your next step

You already turn a plan into small, checkable requests.

Write steps 2 and 3 of your pilot, each with one criterion. It takes five minutes.

Next lesson: the step is done. Who checks it, and what does that person need to receive?

Lesson 12 · Dev with AI v6.2 · INEMA.CLUB

Module 3 · Lesson 3 of 5

The reviewer gets the diff, not the summary

A developer hands a colleague a folder with three sheets separated by paper clips: the request, the list of changes, and the result of the checks.

You can ask the other model to review a step with the right package: briefing, diff, and the results of your checks.

In the previous lesson, AI delivered a step. If the reviewer gets only "I did this," they review one sentence. With the right package, they review the work and point to the exact line with the problem.

In 1 minute

  1. The reviewer’s package has three parts: briefing, diff, and check results.
  2. Ask for findings that say where the issue is, what fails, the impact, and the smallest fix.
  3. The reviewer only reads: they point things out, but don’t change files.

1The reviewer’s package has three parts

A cross-review of a step needs three things. The briefing, so the reviewer knows what was requested. The diff, so the reviewer can see what changed. And what you checked, with the results.

Rafael writes down the step’s checks in a conferencias.txt file: "3 fields appear: yes. Phone rejects letters: yes." This sheet goes with the package to the reviewer.

Pilot folder · review package
1 briefing.md · what was requested
2 mudanca.diff · what changed, line by line
3 conferencias.txt · what you checked and the results
  1. 1The same briefing the person who did the work received.
  2. 2The actual change, generated by Git.
  3. 3Evidence, not opinion.

2With the package, the reviewer finds what’s missing

With the summary, the reviewer tends to agree. With the package, they compare the request with the change and with what was checked. That’s when the criterion no one tested becomes clear.

Rafael sent the stage 1 package to Codex. The review noticed that conferencias.txt said “yes” for the phone number without saying how it was checked.

Stage 1 review

RafaelClaude said it completed the form fields. Is that good?

CodexIt seems adequate. Form fields are a simple change.

The reviewer approved a sentence.

RafaelRead briefing.md, mudanca.diff, and conferencias.txt. List the findings. For each finding: where it is, what fails, the impact, and the smallest fix. Do not change any files.

CodexFinding 1. Where: conferencias.txt, the line "Phone rejects letters: yes". Failure: it doesn't say how it was checked. Impact: medium. Smallest fix: type letters in the phone field, try to submit, and note what appeared.

The finding points to the file, the line, and the fix.

Tap or click the two buttons in the panel and compare the two reviews.

3Assemble and send the package

Before the commit, the saved version, ask Git for the diff in a file, using the terminal. Then ask the other model in read-only mode, with the response saved in a file.

In the app, the steps are the same in plain language. Ask the person who did the work: "run git add -A and save git diff --staged to mudanca.diff". Then open a new conversation with the other assistant or the same one, and mention the three files. In the app, "do not change files" is a request, not a lock: afterward, check with git status that nothing changed.

Carla does everything through the app. Codex did the stage; the AI chat, in a new conversation, gets the three files attached with the paperclip. If the chat refuses the .diff, she renames it to mudanca.txt.

Terminal · pilot folder
$ git add -A
$ git diff --staged --output=mudanca.diff
$ codex exec --sandbox read-only -o review.md "Leia briefing.md, mudanca.diff e conferencias.txt. Liste os achados. Para cada achado: onde, o que falha, impacto e a menor correção. Não altere arquivos."

The first two lines prepare everything, including new files, and save the diff. The third asks Codex only to read, and the response is saved in review.md. In Claude: claude -p --permission-mode plan "…same request…" > review-claude.md. On Windows, prefer PowerShell 7 or the app for this >, which can mess up accents in the older PowerShell.

"read-only" and "plan" are locks: the reviewer can't change any files.

Stuck here? That's normalDoes the long command feel intimidating? Copy it, change only the file names if they're different, and paste it. Or follow the steps in the app; you'll get to the same place.

4A useful finding has four parts

“Could be improved” doesn’t help anyone. Ask for each finding to include where it is, what’s failing, the impact, and the smallest fix. That way, you can quickly decide what to fix.

Carla received a vague finding about the summing step. She asked again, using the four parts, and got the exact issue: the summing line, where the city name with and without an accent was treated as two cities.

Vague finding

“The sum by city may have inconsistencies.”

Useful finding

Where: juntar.py, summing line. Failure: “São Paulo” and “Sao Paulo” are treated as two cities. Impact: total split. Fix: remove accents before summing.

Practice now 0/3

Ask for a review of your step 1

You’re done when you have a review with at least one finding covering all four parts, or the phrase “no issues found” along with what was read. About 10 minutes, on a computer.

The reviewer only reads. Do you have only one assistant? Start a new conversation with it to review. Didn’t do the step from lesson 12? Review the briefing diff from lesson 11.

Read briefing.md, mudanca.diff, and conferencias.txt.
Review step <number> against the briefing criteria.
For each issue: where it is, what’s failing, impact (high, medium, low), and the smallest fix.
If you find no issues, say what you read to reach that conclusion.
Do not change files.

You already ask for reviews of the actual work, with evidence for each finding.

Lesson cheat sheet

Diff review

  1. Packagebriefing, mudanca.diff, and conferencias.txt.
  2. Read-onlyread-only in Codex, plan in Claude; in the app, check git status.
  3. Four partswhere it is, failure, impact, smallest fix.

Your next step

You already prepare the package that makes cross-review useful.

Create conferencias.txt in the pilot folder and write, one line per criterion, what you checked today.

Next lesson: what if the change is an image, not a line of text?

Supplementary material · Where this lesson comes fromFurther reading on the topic. Not included in the lesson time.

Level 3 of the Use Both kit

“Split building and reviewing”: one assistant builds separately from the working version; the other reviews the diff with the briefing and test results. Tests and your approval decide what goes in. The reviewer gets the same briefing and the actual artifact, never a vague summary of what the first assistant did.

Three ways to connect Claude and Codex

According to the video that comes with the kit: the official Codex extension for Claude Code; one command line calling the other, with the response saved to a file; or a change request on a code site that the other model reviews.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 13 · Dev with AI v6.2 · INEMA.CLUB

Module 3 · Lesson 4 of 5

Images: check first, inspect afterward

An operations analyst holds the printed illustration beside the monitor showing the same small image, comparing how it looks at its actual size in use.

You can check an AI-generated image in four ways: the tool and cost beforehand, the actual file, the size you’ll use, and a companion version.

An image is also a change to the project, and it can also be confidently wrong. It may cost you credits without your noticing, arrive only as a link, or have made-up letters around the edges.

In 1 minute

  1. Before: does the session have an image tool? How does it charge?
  2. Done means the file is saved in the project and opened at the size it will be used.
  3. Never overwrite: save the new version alongside it and note which one the project uses.

1Before: is there a tool, and how much does it cost?

Having a coding assistant doesn’t guarantee it can generate images. When it can, each image may use credits from your account. Check both before you ask.

So, write in the request: if there’s no tool, stop and let me know, without using another paid service. Without that instruction, it may switch on its own.

Rafael wanted a banner for the clinic’s contact page. First, he asked the assistant whether the session had an image tool and how it was charged.

Code assistant

RafaelBefore generating: does this session have an image tool? How is it charged? If there isn’t one, stop and tell me; don’t use another service.

AII couldn’t find an image tool in this session. I stopped here without generating anything. I can write the banner description for you to use where you have access.

With the instruction written in the request, it stopped and let him know instead of looking for another service.

2Done means the file is in the project

An image description isn’t an image. A temporary link doesn’t count either: it disappears. Only the file saved in the project folder, with a clear name, counts.

Carla asked for a cover for the weekly report. The first response included only a link. She asked again: "save the file as imagens/capa-relatorio-v1.png".

Doesn’t count

An image description, or a link that expires.

Counts

The file imagens/capa-relatorio-v1.png, inside the project folder.

3Open it at the size it will be used

Almost any image looks good as a thumbnail. Open it at the actual size you’ll use and check three things: text, logos, and cropping. Made-up letters often show up around the edges.

On the small screen, Carla’s cover looked great. Printed on a full sheet, it showed a nonsense word in the corner and a person whose face was cut off.

Cover check at actual size
1 Text: are any letters or words made up?
2 Logos: is there a company logo that isn’t yours?
3 Cropping: is a face or object cut off at the edge?
  1. 1Carefully check all four edges.
  2. 2Someone else’s logo can’t go to the client.
  3. 3Check it in its final format: webpage, phone, or print.

Test yourself

The image looked great as a thumbnail. What do you do before using it?

4A sibling version, never over the original

Need to change the image? Save the new one alongside it, with a different number, and keep the original. Note in the project which version is in use. That way, you can go back and compare.

Rafael generated the banner where he had access, in three versions. He saved all three and noted in the briefing (briefing.md): "banner in use: v2".

images · clinic website
1 images
banner-contato-v1.png
2 banner-contato-v2.png
banner-contato-v3.png
  1. 1A folder for the project's images.
  2. 2The version in use, noted in the briefing.

Stuck here? That's normalYour assistant doesn't generate images? That's okay: the checks in steps 2 to 4 apply to any image, made by AI, by you, or by a designer.

Practice now 0/4

Check a real image

You're done when you have an image from the pilot checked in all three ways and a sibling version saved alongside it. About 10 minutes, on a computer.

Nothing is generated or charged: use an image you already have, from the project or one of your own photos. If the assistant says it will generate something, stop first.

Before generating any image: does this session have an image tool?
How is it charged? Don't generate anything now, and don't use another service.

You now check an AI image the same way you check any other change.

Lesson cheat sheet

AI image

  1. BeforeIs there a tool? How is it charged? Tell it to stop if there isn't one.
  2. Filesaved in the project and opened at actual size.
  3. Versionsv1, v2, v3 side by side; note which one is in use.

Your next step

You now know how to prevent an image from costing you without warning or disappearing from the project.

Open your pilot folder and create the imagens folder, even if it's empty. It takes a minute.

Next lesson: bring everything together in a pilot change, with the review recorded, and switch roles.

Supplementary material · Image with Codex at INEMAFurther reading for this topic. Doesn't count toward the lesson time.

Level 2 of the Use Both kit

Only if an image tool is available in the session; check access and charges first. You're done when the real file is in the project and opened at the size it will be used. Keep the original and save revisions as sibling files, recording which version the project uses.

How INEMA does it today

An automatic routine asks Codex’s image generator for an image and saves the file in a folder. Each generation uses subscription credits, so it runs only when authorized; without authorization, the default is a free local generator. After generation, check the image: text along the edges may be made up.

Sources: Codex + Claude Area — INEMA Events

Lesson 14 · Dev with AI v6.2 · INEMA.CLUB

Module 3 · Lesson 5 of 5

Switch roles and record the review

A developer and an operations analyst, seated at the same desk with two laptops, review each other's work.

You can deliver a change from your pilot with the cross-review recorded in a file: findings, what you fixed, and what you rejected.

A review that stays only in the conversation gets lost. Tomorrow, no one remembers what was found or why a point was rejected. And if you always use the same reviewer, you won’t find out whether the other route works better for your task.

In 1 minute

  1. The person who does the work and the reviewer can switch roles. Try the other route for your task.
  2. Fix the finding that has evidence; reject the rest, with a reason.
  3. Record everything in revisao-etapa-N.md, alongside the code.

1The person who does the work and the reviewer can switch roles

“Claude does the work, Codex reviews” is a route to start with, not a ranking. In one step, try the reverse and compare using your own criteria: which route found more real issues?

Only have one assistant? Switch between conversations: one does the work, a new one reviews, and in the next step, do the opposite.

Rafael did step 2 with Codex and asked Claude to review it. Claude found a sending failure that the previous route had missed in this task.

Route A

Claude does the step. Codex reviews the diff.

Route B

Codex does the step. Claude reviews the diff.

Both routes are valid. Choose the one that finds more real issues in your task, not the more famous one.

2Fix what has evidence; reject with a reason

Not every finding is worth fixing. Fix what points to where, shows the failure, and matches a criterion in the briefing. Reject what is a matter of taste or outside the scope, and write down why.

Carla received two findings. One pointed out repeated cities caused by accents: she fixed it. The other suggested changing the report format, which is under “Out of scope” in the briefing: she rejected it.

Fixed

“São Paulo” and “Sao Paulo” counted as two cities. Criterion 2 of the briefing.

Rejected

“Change the report format.” Reason: it is listed under Out of scope in the briefing.

Stuck here? That's normalNot sure whether to fix or reject something? Ask: does this break a criterion in the briefing? If yes, fix it. If not, make a note and move on.

3The review record

A short file for each step keeps a record of the review: who reviewed it, what they found, what was fixed, what was rejected, and why. It stays in the pilot folder, alongside the code.

Rafael’s revisao-etapa-2.md has six lines. In module 5, another conversation will read this file to continue.

revisao-etapa-2.md · clinic form
1 Did the work: Codex · Reviewed: Claude
2 Finding 1: no warning appears when sending fails · fixed
3 Finding 2: change the button colors · rejected (Out of scope in the briefing)
4 Checked afterward: the test submission arrived at reception
  1. 1Who did the work and who reviewed it.
  2. 2Finding fixed, with what it was.
  3. 3Finding rejected, with the reason.
  4. 4The check after the fix.

4The lab: a complete change

In Git, the entire module becomes a sequence: branch, step, package, review, fix, record, and commit. The commit comes last, after the review. It’s the first change in your pilot, completed and checked by two people.

Carla completed step 2 in the app, in the order shown in the window below. From the step to the commit, it took about forty minutes, and revisao-etapa-2.md was saved in the folder.

A reviewed change
1 Test branch (lesson 11)
2 Step with the criteria pasted in and a stopping point (lesson 12)
3 Package and read-only review (lesson 13)
4 Fix, record, git add -A, and commit (this lesson)
  1. 1Separate from the version that works.
  2. 2Small change.
  3. 3Another look at the actual work.
  4. 4Record it, and only then commit.
The branch and the commit come from lessons 11 and 12. In the app, ask the assistant for each step. The complete lab, with the commands, is in the supplementary material.

Practice now 0/4

A pilot change, with the review recorded

You’re done when the pilot folder has a completed revisao-etapa-N.md, based on the review.md from lesson 13, and the step is saved in a commit. About 12 minutes, on a computer.

Everything happens on the test branch; the version that works doesn’t change. Don’t have the review.md from lesson 13? Practice with Carla’s example in the template. If the review asks for something big, write down "save it for another step" and continue.

# Stage <N> review
Done by: <model or conversation> · Reviewed by: <other>
Finding 1: <what it is> · fixed | rejected (reason)
Finding 2: <what it is> · fixed | rejected (reason)
Checked afterward: <what you tested and the result>

Carla’s example:
# Stage 2 review
Done by: Codex · Reviewed by: AI chat, new conversation
Finding 1: city with and without an accent counted as two · fixed
Finding 2: change the report format · rejected (Out of scope in the briefing)
Checked afterward: each city appears once; total matches the spreadsheets

You finished module 3 with a change that was made, reviewed, corrected, and recorded.

Lesson cheat sheet

Review recorded

  1. Routesswitch who does the work and who reviews it; compare in your task.
  2. Decidefix it with evidence; reject it with a reason.
  3. Filerevisao-etapa-N.md next to the code.

Your next step

You now deliver changes with two sets of eyes and a record another person can understand.

Reread your revisao-etapa-N.md tomorrow. If anything is unclear, write one more line.

In module 4: the review is done. But does the change really work? Time to verify and decide.

Supplementary material · The complete module labMore depth on the topic. Does not count toward lesson time.

A reviewed change, from start to finish

Set aside 30 to 45 minutes. In the next pilot stage, switch the route: whoever reviewed does the work, and whoever did the work reviews it. Have only one assistant? Use one conversation to do the work and a new one to review it.

  1. On the test branch, request the stage with the criteria pasted in and the stopping point (lesson 12).
  2. Check git status and record the result for each criterion in conferencias.txt.
  3. Generate the package and request a read-only review (lesson 13).
  4. Fix or reject each finding and record it in revisao-etapa-N.md.
  5. Check again and make the commit.
git add -A
git diff --staged --output=mudanca.diff
codex exec --sandbox read-only -o review.md "Leia briefing.md, mudanca.diff e conferencias.txt. Liste os achados. Para cada achado: onde, o que falha, impacto e a menor correção. Não altere arquivos."
git add -A
git commit -m "etapa 2 revisada"

In Claude, the read-only review is claude -p --permission-mode plan "…" > review-claude.md.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 15 · Dev with AI v6.2 · INEMA.CLUB

Module 4 · Lesson 1 of 5

Each criterion calls for one check

An operations analyst uses a pencil to draw arrows connecting each item on a paper form to a line on a printed spreadsheet, to know how to check each one.

You can write, next to each criterion in your briefing, how it will be checked and who will check it.

In module 1, you wrote the criteria. A criterion without a way to check it becomes decoration in the briefing. At delivery time, each one needs a check agreed on in advance.

In 1 minute

  1. Each criterion gets one line: how to check and who checks it.
  2. A repeated rule becomes an automatic test. Appearance and meaning are checked manually.
  3. A test that has never failed proves nothing. Ask to see it fail first.

1A criterion without a check is decoration

Every acceptance criterion needs an agreed way to check it. Without one, each person checks in a different way, or nobody checks.

Rafael took the three criteria from the clinic’s form and wrote down next to each one how he would check it.

briefing.md · how I check
1 Message reaches reception → submit it through the page and check the inbox
2 Phone field rejects letters → automatic test
3 Page doesn’t scroll sideways on a phone → open it on a phone
  1. 1Manual check, done by Rafael.
  2. 2A rule that repeats: the computer checks it.
  3. 3Appearance: manual check.

2Automatic for rules, manual checks for meaning

An automatic test checks a rule the same way, as many times as you want. It’s great for “rejects letters” or “the total matches.” It’s not good for “the text makes sense” or “the page looks good.”

Carla added a check cell to the spreadsheet: it compares the report total with the sum of the spreadsheets and shows “matches” or “doesn’t match.” She reads the list of cities with her own eyes.

Automatic test

Rules that repeat: sum, format, required field. It runs the same way every time.

Manual check

Meaning, appearance, tone, the user’s path. Someone needs to look.

A good criterion says which of the two will check it.
How Carla made the check cell

In an empty cell in the report spreadsheet, she wrote a formula that compares the total with the sum of the spreadsheets. For example: =SE(B2=SOMA(C2:C8);"confere";"não confere"), with B2 as the report total and C2 through C8 as the totals from each spreadsheet. Replace the cells with yours. Don’t know how to set it up? Ask AI for the formula and say which cells to compare.

Stuck here? That's normalNot sure whether a criterion can have an automatic test? Ask: can you write the rule as a calculation or format (sum, numbers only, filled-in field)? Then it’s automatic. Does it depend on looking, reading, or using it? Manual check.

3You don’t need to write the test

AI writes the automatic test. Your job is to request one test per criterion and look at its result: the text it shows when it runs. Don’t accept just “I created the tests; they all pass.”

Rafael asked for one test for the phone criterion and its result, before and after the fix.

Code assistant

RafaelCreate tests for the form.

AII created tests for the form. They all pass.

Which tests? Checking which criterion? There’s no way to tell.

RafaelCreate a test for criterion 2 in briefing.md: the phone field rejects letters. Test it with the text "abc". Run it before fixing the issue and show me the result. Then fix it and run it again.

AIBefore the fix: FAILED — the field accepted "abc". After the fix: PASSED — the field rejected "abc".

One criterion, one test, and the result shown both times.

Tap or click the two buttons in the frame and compare the two responses.

4A test that has never failed proves nothing

A poorly written test can always pass, whether the error is there or not. That’s why it’s worth seeing the test fail once, while the error still exists. If it has never failed, you don’t know whether it checks the right thing.

In a copy, Carla deliberately changed a value in the last spreadsheet, and the cell kept saying "checks out." The formula added C2 through C7 and left out the last spreadsheet.

Always passed

"Checks out" with the correct values and with a value deliberately changed.

What it proves: nothing.

Failed once already

With the formula adding C2 through C8: "doesn't check out" with the changed value; "checks out" with the correct ones.

What it proves: that it catches the error.

Test yourself

The AI wrote a test for "the total matches," and it passed on the first try. What do you ask next?

Practice now 0/3

How you'll check each criterion

You’re done when every criterion in your briefing has a “how I’ll check” line and a “who will check” line. About 10 minutes, on a computer.

You’ll only write in your file; nothing runs yet. Don’t have the briefing.md from lesson 5 (in the pilot folder)? Use three criteria from any task of yours.

## How I check
1. <criterion 1> → <how I check> · <automatic or manual> · who: <name>
2. <criterion 2> → <how I check> · <automatic or manual> · who: <name>
3. <criterion 3> → <how I check> · <automatic or manual> · who: <name>

Carla's example:
1. Total equals the sum of the spreadsheets → check cell · automatic · who: Carla
2. Each city appears once → read the list · manual · who: Carla
3. Source below each table → look at each table · manual · who: colleague on the leadership team

You already know how each criterion in your pilot will be checked.

Lesson cheat sheet

Criterion and verification

  1. How I checka line next to each criterion.
  2. Automatic or manualrepeated rule × meaning and appearance.
  3. See it faila test only counts after it catches the error once.

Your next step

You already connect each criterion to an agreed check.

Ask the AI for the test for the criterion you marked, with the result before and after the fix.

In the next lesson: the test passed. Does that mean the system works for the person using it?

Lesson 16 · Dev with AI v6.2 · INEMA.CLUB

Module 4 · Lesson 2 of 5

The test passed. Now use it for real.

A developer standing by the window fills out the website form on a phone as if they were a patient.

You can check a change in your pilot by following the path someone who will use it takes, from beginning to end, with test data.

A test checks one rule. People using the system go through several rules in a row, on a phone, in a hurry. Many errors only show up along that entire path.

In 1 minute

  1. A passing test doesn't mean the system works for the person using it.
  2. Follow the user's path from start to finish using test data.
  3. If AI operates the screen for you, you handle passwords, payments, and submissions.

1A passing test doesn't mean the system works

The automated test checks one rule at a time. It doesn't see the small phone screen, the keyboard covering the button, or the message landing in the spam folder.

On the clinic form, the phone number test passed. On Rafael's phone, the keyboard covered the submit button, and he couldn't scroll to it.

Test only

What Rafael saw: PASSED — the field rejected "abc".

What the patient saw: a submit button he couldn't reach.

Entire path

What Rafael did: opened it on a phone, filled it out as a patient, and tried to submit it.

What he found: the hidden button, before the clinic did.

2Follow the user's path

Think about who will use it and where. Follow the same path on the same type of device, from the first click to the final result.

The management team reads Carla's report on a phone, through email. She started sending the report to herself and opening it on her phone before sending it.

Patient path · clinic form
1 Open the site on a phone using the link the clinic shares
2 Enter a name, phone number, and message
3 Submit
4 Check the reception inbox, including spam
  1. 1The same device and the same way in as the patient.
  2. 2Type it in; don't paste it.
  3. 3Go all the way through without skipping a step.
  4. 4The result at the other end, not the "sent" screen.

3Test data, never real data

Check with made-up data and make it clear that it's for testing. That way, no one mistakes the test for a real case, and no real person's data is shared.

Rafael submits it as "Test Patient" and agrees with reception that messages with that name can be deleted.

Carla checks a copy of the spreadsheets, never the originals.

Real data

A real patient's name and phone number in the test form.

Test data

"Test Patient," an example phone number, and reception notified.

Stuck here? That's normalCan't test without doing something real, like sending an email? Send it to yourself first. Only then, send it once to the real destination, after letting the recipient know.

4When AI operates the screen for you

Optional: if your assistant can't operate the browser, skip this part. Some assistants can operate the browser or computer, such as Claude Code with the Chrome extension. This only works if the tool is available in your account: having the assistant on your computer isn't enough.

Even with the tool, three things stay with you: the password, payment, and final click to send. The AI prepares and stops. If it refuses an action for safety reasons, don't switch to another model to get around it: do it yourself the usual way.

Assistant with browser use

RafaelOpen the clinic's contact page and fill in: name "Paciente Teste", phone "11 90000-0000", message "teste do formulário". Stop before sending.

AII filled in the name "Paciente Teste", phone "11 90000-0000", and message "teste do formulário". I stopped before sending. Check the fields and tell me if I can click send.

The AI prepared it; sending is waiting for your confirmation.

Practice now 0/3

The whole process, once

You're done when you've followed the user's path from beginning to end and noted what you saw. About 10 minutes, on the device the user uses.

Use test data and a copy, never the original. If your pilot doesn't have anything working yet, follow the process using the result of any AI task from this week.

You're already checking the system the way it will be used.

Lesson cheat sheet

User path

  1. Beyond testingthe whole process catches what a loose rule doesn't.
  2. Same deviceand test data, with the recipient informed.
  3. AI on screenit prepares; the password, payment, and sending are yours.

Your next step

You're already checking the pilot the way the user will use it.

Today, agree with the person receiving your pilot results on a test name, such as "Paciente Teste".

Next lesson: you found a problem. How much should you ask the AI to change?

Supplementary material · Supervised computer useFurther reading on the topic. Doesn't count toward lesson time.

Level 4 of the Use Both kit

The kit calls this "supervised computer use": delegating an authorized action in an interface when there's no other way. Prefer a test account and made-up records. Confirm that the assistant actually has the computer-use tool. Enter the password and verification code yourself, never in the request. Ask it to prepare and stop before sending, buying, publishing, or deleting. Review the final fields and check the result afterward.

The kit also says: if a tool refuses a sensitive action, don't switch to another model to get around the refusal. Use the usual process, carried out by a person.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 17 · Dev with AI v6.2 · INEMA.CLUB

Module 4 · Lesson 3 of 5

The smallest fix, and one line to keep it from happening again

A developer writes a single line in a lined notebook beside the laptop and a cup of coffee, recording a failure they just fixed.

You can request the smallest fix for a failure in your pilot and record it in one line in the FALHAS.md file.

When something breaks, you may want to ask, "redo it properly." AI redoes much more than necessary, and you have to check everything again. And without a record, the same mistake comes back next month.

In 1 minute

  1. Ask for the smallest fix that solves it, and say what must not change.
  2. Record the failure in one line: date, what broke, the fix, prompt or infrastructure.
  3. After about ten lines, the pattern in your mistakes appears on its own.

1Fix what broke, not the whole system

The smallest fix is the shortest change that makes the criterion pass again. The smaller it is, the less there is to check again and the lower the chance of breaking something else.

The send button disappeared under the phone keyboard. Rafael almost asked, "redo the form." He only asked for the button to stay visible while the keyboard was open.

Redo everything

Request: "The form is bad on mobile, redo it."

Result: a new page, and three criteria to check again.

Smallest fix

Request: "Keep the send button visible while the keyboard is open. Don't change anything else."

Result: one change, one thing to check.

2The fix request has three parts

What failed, with evidence. What to fix, and only that. What the AI should report at the end: what changed and what stayed the same.

Carla's formula counted the deadline from the wrong date. She asked to fix only the formula, using the example spreadsheet that showed the error.

Code assistant

CarlaFailed: the shipping deadline is counted from the order date. Evidence: order 3 in the attached spreadsheet. Fix only the deadline formula so it counts from the departure date. Don't change the columns or the format. At the end, say what changed and what stayed the same.

AII changed the deadline formula so it counts from the departure date. The columns and format stayed the same. To check: calculate order 3 by hand and compare.

Failure with evidence, a clearly defined fix, and a report of what changed.

3One line in FALHAS.md

Fixed it? Before moving on, write one line in a file called FALHAS.md, in the pilot folder. Four fields: date, what broke, the smallest fix, and whether it was prompt (your request) or infrastructure (something outside the request). No long story.

Rafael creates FALHAS.md the same way he created the briefing.md in lesson 5. You can also ask the assistant: "create FALHAS.md in the pilot folder with this header". The most recent line goes at the top.

FALHAS.md · pilot folder
1 | date | what broke | smallest fix | prompt or infrastructure |
2 | 28/09 | send button under the phone keyboard | button visible while keyboard is open | prompt |
3 | 25/09 | message didn't reach reception | change the sending address | infra |
  1. 1The header, written once.
  2. 2Today's failure, at the top.
  3. 3The older ones go below.

4Prompt or infra: where the error came from

Prompt, here, means your request: mark “prompt” when the request led to the error, or the AI misunderstood. Infra is when the problem is outside the request: machine, network, service outage, account, address. If it was both, mark both.

After ten lines, Carla saw that six were “prompt” and all mentioned columns. She started pasting the column names into every request.

Prompt

“The request didn't say which date to count the time limit from.”

Infra

“The clinic's sending address was wrong in the email provider.”

The label shows where the fix belongs: in the way you ask or outside it.

Stuck here? That's normalNot sure whether it was prompt or infra? Mark “?” and move on. Over time, similar lines will show the answer.

Test yourself

The AI couldn't finish because its service went down. Which label goes on the line?

Practice now 0/3

Your first line in FALHAS.md

You're done when the pilot folder has a FALHAS.md with the header and one completed line. About 8 minutes, on a computer.

It's your text file. Hasn't there been a failure in the pilot yet? Record the last time an AI made a mistake in your work, or use Rafael's button failure. The bars and dashes in the template draw the table: copy it as is.

| date | what broke | smallest fix | prompt or infra |
|---|---|---|---|
| <dd/mm> | <what broke, in a few words> | <the smallest change that fixed it> | <prompt, infra, or both> |

You now have a record that will show over time where your requests fail.

Lesson cheat sheet

Fix and record

  1. Smallest fixthe shortest change that makes the criterion pass again.
  2. Three-part requestfailure with evidence, only what to fix, a report of what changed.
  3. FALHAS.mddate, what broke, fix, prompt or infra.

Your next step

You can already fix things without starting a major overhaul and keep what you learned.

The next time there's an AI failure, any kind, write the line before moving on to the next task.

Next lesson: what about an error that doesn't go away in one or two attempts?

Lesson 18 · Dev with AI v6.2 · INEMA.CLUB

Module 4 · Lesson 4 of 5

A difficult error needs one goal

An operations analyst examines a printed spreadsheet with a magnifying glass, looking for an error, with a desk timer beside the laptop.

You can write the goal for a difficult error: a success you can see, the limits, and when the AI should stop and report.

Some errors don't go away in one or two attempts. The temptation is to say, “fix this, whatever it takes.” Then the AI runs for a long time, spends a lot, and no one can say whether it got there.

In 1 minute

  1. Replace “fix this” with a goal that has observable success: what input, what result.
  2. Set limits in the request: files and attempts. And a reporting point. Keep the time or spending cutoff outside the AI.
  3. Once it meets the criterion, the AI stops. No extra frills.

1“Fix this” has no end

A request without a goal leaves the AI to decide when to stop. It may stop too early or never stop. A goal tells you how to know when you’ve reached it.

Carla’s report showed a different total only in some weeks. She wrote the goal using a week that failed and the expected result.

No goal

“The total is sometimes off. Fix this.”

Goal

“Using the spreadsheets for week 38, the report total matches the sum of the spreadsheets. The check cell shows ‘matches.’”

2The three parts of a goal

Observable success: the input, the expected result, and how to check it. Limits: which files, how many attempts. Reporting point: when to stop and say what you tried. And, outside the request, the time or spending cutoff, as in lesson 4.

Rafael used the same template for an error that happened only with a phone number from another state.

meta.md · total error
1 Success: week 38 → total matches the sum; cell says “matches”
2 Limits: only the file juntar.py; 3 attempts · cutoff: 30-minute alarm
3 Report: when it passes or when it reaches the limit
  1. 1You can check without asking the AI anything.
  2. 2Limits in the request; the cutoff stays outside the AI, as in lesson 4.
  3. 3What you tried, what failed, what’s left.

3Reached it: stop. Didn’t reach it: report.

Specify both outcomes in the request. If it meets the criterion, the AI stops without adding features no one asked for. If it reaches the limit, it stops and describes the block.

On the third attempt, Carla’s AI stopped and showed where the difference was.

Code assistant

CarlaGoal: using the spreadsheets for week 38, the report total matches the sum of the spreadsheets. How to check: the check cell shows “matches.” Only the file juntar.py. At most 3 attempts. If it passes, stop. If not, stop and say what you tried, what failed, and what’s left.

AII stopped after 3 attempts; the goal still isn’t met. I tried three ways to add up week 38. Failed: one spreadsheet for week 38 has duplicate requests on two rows. What’s left: you need to say whether a duplicate request counts once.

Without reaching the goal, the AI gave a clear account of what was blocking it, with the question only Carla can answer.

Stuck here? That's normalDon’t know the expected result? Then you don’t understand the error yet. First, ask the AI only to find a case that fails, without fixing anything.

4Goal mode: more time doesn’t mean better quality

Some assistants have a goal mode, where they pursue the objective on their own. According to the author of the video that comes with the Use Both kit, Codex has a goal mode (/goal), which uses 3 to 5 times more time and tokens. This is an account, not a measurement: check whether the mode exists in your version and in the app.

With or without this mode, the conditions are the same: observable success, limits in the request, and a real stop. Running for a long time proves nothing; what proves it is the criterion passing.

Rafael doesn't use the goal mode for the site. He sends the usual request, with the same three parts.

With goal mode

Same goal, same limits, and a real time or spending limit.

Without goal mode

Usual request with the same three parts. It works with any assistant.

The feature changes how it runs. The goal is still yours.

Practice now 0/3

Write the goal for one error

You're done when you have a goal with observable success, three limits, and both outcomes. About 10 minutes, in a meta.md file in the pilot folder or in your notes app.

You only need to write it; you don't need to send it now. No difficult error in your pilot? Use Carla's total error, replacing it with your spreadsheet or page.

Goal: with <the input that fails>, <the expected result>.
How to check: <the test or check>.
Limits: only <files>; at most <3> attempts.
If it passes, stop without adding anything.
If it doesn't pass, stop and say what you tried, what failed, and what's missing.

You can now turn a difficult error into a task with a beginning, an end, and a limit.

Lesson cheat sheet

Goal

  1. Observable successinput, expected result, how to check.
  2. Limits, stop, and reportfiles and attempts in the request, stop outside it, and when to stop and report.
  3. Two outcomesit passes, stop; it doesn't pass, describe the blocker.

Your next step

You now know how to set a goal for a difficult error.

Save the goal template in your notes app, ready for the next error that resists two attempts.

Next lesson: everything checked. Who decides whether the pilot goes in, and based on what?

Supplementary material · Goal with a stop conditionFurther reading on this topic. Not included in the lesson time.

Level 5 of the Use Both kit

For a well-defined problem that needs a long stretch of work: write the objective with an observable success condition (which tests, which input, which expected output); limit files, tools, review rounds, and spending, starting small; use the app's goal feature if there is one, or a usual request with the same conditions; set a real usage or time control if you need a cap; ask for a report at a defined point or after a repeated blocker. When the acceptance tests pass, stop without adding features.

Done means test evidence that meets the condition, or a reported blocker with what was tried. A long runtime is not a quality score. The video's comparisons of speed, persistence, and cost are the author's experience, not a measured guarantee.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 19 · Dev com IA v6.2 · INEMA.CLUB

Module 4 · Lesson 5 of 5

You decide what gets included

An operations analyst signs an acceptance sheet beside the checked sheets marked with checkmarks.

You can complete the acceptance review of your pilot: for each criterion, the evidence you saw and the decision to accept, accept with an open item, or send it back.

It’s time to deliver. The tests passed, the other assistant approved it, and you used it for real. All that’s left is the decision, written down with the evidence behind it. This is what you show if someone asks, “Who checked?”

In 1 minute

  1. Acceptance is a decision based on evidence, criterion by criterion.
  2. Evidence is what you saw, not what the AI said.
  3. Three outcomes: accept, accept with a documented open item, or send it back.

1Acceptance fits on one sheet

Each acceptance criterion in the briefing becomes one line: the criterion, the evidence, and the result. At the end, the decision and your name.

Rafael put together the acceptance review for the clinic form in a file called aceite.md, next to briefing.md.

aceite.md · clinic form
1 Message arrives at reception → inbox screenshot with “Test Patient” · ok
2 Phone number rejects letters → test result: failed before, passed after · ok
3 No horizontal scrolling → phone screenshot with no horizontal scrolling · ok
4 Decision: accept · Rafael · 28/09
  1. 1Evidence from a manual check.
  2. 2Evidence from an automated test.
  3. 3The user’s path, from lesson 17.
  4. 4The decision, with name and date.

2Evidence is what you saw

“The AI said it passed” isn’t evidence. Evidence is a screenshot of the inbox, the test result, the sum you calculated, or the link opened on a phone. To take a screenshot: Windows+Shift+S on Windows, Cmd+Shift+4 on Mac, and the power button + volume down on most phones. It’s something another person can look at and agree with.

In the report acceptance review, Carla replaced “reviewed by AI” with the check cell saying “checks out” and the list of cities she read herself.

Statement

“Codex reviewed it and said everything looks good.”

Evidence

“Check cell: ‘checks out’. I read the list of cities: none are repeated.”

3Three outcomes, not two

Not everything is “perfect” or “redo it.” Often the right choice is to accept with a small open item, documented with an owner and deadline. What you can’t do is leave the open item only in your head.

All of Carla’s criteria passed, but the report title had the wrong week. She accepted it with an open item: correct the title by Friday.

Acceptance decision
1 Accept: all criteria have evidence
2 Accept with an open item: a small, documented issue with an owner and deadline
3 Send back: an important criterion lacks evidence or fails
  1. 1Standard delivery.
  2. 2Deliver, and document the open item in the acceptance.md file itself: what, who, by when.
  3. 3Make the smallest correction and check again, as in lesson 18.

Stuck here? That's normalNot sure whether to "accept with an open item" or "send back"? Ask: if the client sees this issue tomorrow, can they still use it? If so, it's an open item. If not, send it back.

4No one signs for you

Cross-review and tests help you decide. The decision remains human. Not even approval from two models can sign in your place.

Rafael sends the clinic the link and a sentence: "checked by me, criterion by criterion; the acceptance is in the project".

What helps you decide

Tests, the user journey, cross-review.

Who decides

You, with your name and date on the acceptance.

Test yourself

A criterion has no evidence, but both AIs say it passes. What decision should you make?

Practice now 0/4

Your pilot acceptance

Done when the pilot folder has an acceptance.md file with evidence for each criterion and your signed decision. About 12 minutes, on a computer.

This is your file; nothing is sent. Is the pilot not ready yet? Accept what exists: the honest decision may be "send back," and that is a result too.

# Acceptance: <pilot name>

| criterion | evidence (what I saw) | result |
|---|---|---|
| <criterion 1> | <screenshot, test result, checksum, link> | <ok, failed, or not checked> |
| <criterion 2> | <...> | <...> |
| <criterion 3> | <...> | <...> |

Decision: <accept, accept with an open item, or send back>
Open item (if any): <what, who, by when>
Signed: <your name> · <date>

You completed module 4 with your pilot acceptance, signed and backed by evidence.

Lesson cheat sheet

Acceptance

  1. For each criterioneach row with its evidence and result.
  2. Evidencewhat you saw and another person can inspect.
  3. Decisionaccept, accept with pending items, or send back; with a name and date.

Your next step

You can now wrap up a delivery with a written decision and evidence for each criterion.

Show aceite.md to someone on your team and ask if they would agree with the decision just by reading the file.

In module 5: tomorrow, another conversation, or another assistant continues the pilot. How do you hand it off without explaining everything again?

Lesson 20 · Dev with AI v6.2 · INEMA.CLUB

Module 5 · Lesson 1 of 5

Handoff: the state in a file

At the end of the day, a developer writes a short page of notes and leaves it on the laptop keyboard before closing it.

You can write down in a file where the work stopped, the decisions, the pending items, and the exact next action.

A new conversation with AI starts without remembering anything. Without a record, you explain everything again, and AI sometimes reverses a decision that was already made. A short file solves this.

In 1 minute

  1. The handoff is the shift handover: a short file at the end of the session.
  2. Four parts: where things stopped, decisions, pending items, and the exact next action.
  3. It records what happened. It is not an instruction for AI to start doing things.

1The new conversation remembers nothing

Each conversation with the model has a maximum size. When it ends, or when you open another one, what you agreed on is left behind.

The handoff carries that agreement into the next conversation. It is like a shift handover in a hospital: the person arriving reads it before touching the patient.

Rafael stopped work on the clinic form on a Friday. On Monday, he opened a new conversation and said only "continue the form".

New conversation, Monday

RafaelContinue the contact form.

AII can create the form with name, email, phone, and message fields. Which email should submissions go to?

AI started from scratch and asked what had already been decided.

RafaelRead handoffs/latest.md and tell me where we stopped before changing any file.

AIWe stopped at sending it to reception. Decided: the phone field rejects letters. Pending: the test message did not arrive. Next action: send a test and check the reception inbox.

AI started with what the file says. Check the response against the file, line by line.

Tap or click the two buttons on the board and compare the same Monday.

2The four parts of a handoff

A good handoff fits on one screen. It says where the work stopped, what was decided, what is pending, and what the next action is.

Carla writes hers every Friday, when she finishes the week's report. It takes five minutes.

handoffs/latest.md · Carla's report
1 Where things stand: weekly report generated in the "relatorios" folder
2 Decisions: cities in alphabetical order; original spreadsheets stay unchanged
3 Pending: order total not yet checked against the sum
4 Next action: add up the orders column in the 5 spreadsheets and compare
  1. 1Where the work is, with the folder name.
  2. 2What doesn't need to be discussed again.
  3. 3What hasn't been checked yet.
  4. 4Just one step you can take tomorrow morning.

3The next action needs to be specific

The last part is where things most often go wrong. "Continue the form" doesn't say anything. A good next action says what to do, where, and how to tell whether it worked.

Rafael replaced the vague sentence with one any assistant can follow, even another model.

Vague

"Next action: continue the form."

Specific

"Next action: send a test message through the contact page and check the reception inbox. If it doesn't arrive, check the submission in contato.html."

Net gain: the new conversation starts with the right step, without asking anything.

Stuck here? That's normalDon't know what the next action is? Write the question that still needs an answer, such as "find out why the message isn't arriving." A clear question works too.

4Ask the AI itself for the handoff, then check it

At the end of the session, ask the assistant to write the handoff. It has the conversation details at hand, but it makes mistakes. Then read and correct it: you are responsible for what's there.

One thing to keep in mind: the handoff records what happened. It doesn't authorize the next conversation to send, publish, or delete anything on its own.

Carla asks for the handoff before closing. The first time, the AI wrote "total checked." Carla hadn't checked it, so she changed it to "pending."

End of session

CarlaBefore ending, write the handoff in handoffs/latest.md with: where things stand, decisions, pending items, and the exact next action. Mark as pending anything I haven't checked.

AIHandoff saved. Pending: order total not checked against the sum of the spreadsheets. Next action: add up the orders column in the spreadsheets and compare it with the report total.

"Mark as pending anything I haven't checked" keeps the handoff from claiming more than what happened.

Practice now 0/3

Your pilot handoff

Done when the latest.md file, inside the pilot's handoffs folder, has all four parts. About 10 minutes, on a computer: this module requires a computer.

It's just a text file of your own. Didn't do the pilot in the previous modules? Use any work you'll pick up again tomorrow. To create a new file in Windows: right-click the folder › New › Text Document, then rename it (lesson 5 shows you step by step). First, turn on View › Show › File name extensions; otherwise, the file becomes latest.md.txt. You can also ask the assistant: "create handoffs/latest.md with this template".

# Handoff — <date>

## Where things stand
<what was done and which folder or file it is in>

## Decisions
- <what is not up for discussion again>

## Pending items
- <what has not been checked yet>

## Next step
<one exact step: what, where, and how to know it worked>

You’re already leaving the work ready to be picked up in any conversation.

Lesson cheat sheet

Handoff

  1. Shift handoffa short file at the end of each session.
  2. Four partswhere things stand, decisions, pending items, next step.
  3. A record, not an orderthe person who arrives reads it; they don’t start doing things.

Your next step

You already know how to end a session without losing what was agreed on.

In your next work session with AI, ask for the handoff before closing and correct one line that says more than what happened.

Next lesson: the handoff is written. How do you get the new conversation to read and check it before acting?

Lesson 21 · Dev with AI v6.2 · INEMA.CLUB

Module 5 · Lesson 2 of 5

Prime: read and check before acting

In the morning, with a cup of coffee nearby, an operations analyst reads the page of notes left by the computer before getting started.

You can ask a new conversation for a project summary with sources and contradictions before it changes any files.

The handoff only helps if someone reads it. And reading isn’t enough: the new conversation needs to show what it understood so you can check before it acts. In this course, this first request is called a prime. It isn’t a command for the tool: it’s text you paste in.

In 1 minute

  1. The prime is the first request of the session: read the files and tell me where we are.
  2. The response answers five questions, with the file for each answer.
  3. Read only. Nothing is sent, published, or deleted at this stage.

1The first request is “read and tell me”

The prime pairs with the handoff. One is written at the end of a session; the other is read at the start of the next.

Instead of saying “continue,” you ask: read these files, don’t change anything, and tell me where we are.

Carla always starts her Monday conversation with the same request, saved in a text file. She copies, pastes, and waits for the summary.

New conversation, Monday

CarlaRead briefing.md and handoffs/latest.md. Don’t change any files. Tell me: the goal, one rule, the latest decision, the next step, and any contradictions, citing the file for each piece of information.

AIGoal: weekly report of requests by city (briefing.md). Rule: don’t change the original spreadsheets (briefing.md). Latest decision: cities in alphabetical order (latest.md). Next step: add up the requests column and compare it with the total (latest.md). Contradictions: none found.

Each piece of information comes with the file it came from. That makes it easy to check.

2The prime’s five questions

A good prime always answers the same five questions. The objective comes from the briefing; the decision and next action come from the handoff. If one of them is wrong, the problem is in the file, not the next task.

Rafael checks the summary in one minute: he compares each line with the handoff he wrote himself on Friday.

What the prime returns
1 Objective and acceptance criteria
2 One rule, with its source file
3 The last decision made
4 The exact next action
5 Contradictions between the files
  1. 1Comes from briefing.md.
  2. 2Shows whether it read the rules, not just the handoff.
  3. 3Comes from the handoff.
  4. 4Comes from the handoff.
  5. 5This is where the prime helps most.

3The prime only reads; you decide the next step

The handoff is a record of the past, not a new request. If it says "the report still needs to be sent," the new conversation won’t send it on its own. It summarizes and waits for you.

Once, Rafael’s assistant read "next action: publish the page" and started publishing right away. Since then, his prime has said "don’t change anything and wait for my instruction".

Read it and started doing it

"I saw in the handoff that the page still needs to be published. Publishing the page now."

Read it, summarized, and waited

"The handoff lists publishing the page as the next action. I’ll wait for your confirmation."

The phrase "don’t change anything" in the prime separates reading from acting.

Stuck here? That's normalDoes this seem like too much distrust of AI? It’s the same care someone brings when starting a shift: first read and check, then make changes.

4Finding a contradiction is worth its weight in gold

When two files say different things, the prime points it out. It’s better to find that in the summary than in the middle of the work.

Another week, Carla’s briefing asked for cities in alphabetical order. That week’s handoff said "sort by total requests". The prime showed the conflict, and she decided in ten seconds.

Prime response

AIContradiction: briefing.md asks for cities in alphabetical order; handoffs/latest.md says "sort by total requests". Which one should be followed?

CarlaUse alphabetical order. Correct the handoff.

The question came up before the work, not after the report was ready.

Test yourself

The prime suggested a next action that isn’t in any file. What does that indicate?

Practice now 0/3

Build the prime for your pilot

You’re done when a new conversation has returned the summary with sources and you’ve checked each line. About 10 minutes, on a computer: this module requires a computer.

The request prohibits changing files: read-only only. Don’t have the handoff from lesson 21? Remove the handoff line from the request and use only the aula 5 briefing.md, which is in the pilot folder. If your file is briefing.txt, change the name in the request. If the AI starts changing anything, click the stop button (or press Esc).

Read briefing.md and handoffs/latest.md, in this order.
Do not change any files or run anything.
Tell me, citing the file for each piece of information:
1. The goal and acceptance criteria.
2. One rule, with its source file.
3. The latest decision.
4. The exact next action.
5. Any contradiction between the files.
Then wait for my instruction.

You can start a session knowing what the AI understood before it changes anything.

Lesson cheat sheet

Prime

  1. First requestread the files, change nothing, tell me where we are.
  2. With a sourceeach piece of information in the summary points to a file.
  3. Wait for the instructionthe handoff is a record, not authorization.

Your next step

You can check what the new conversation understood before it starts working.

Save the prime request in a text file in the pilot folder, so you can always paste it the same way.

In the next lesson: briefing, handoff, rules. Where does each file live, and which one tells you to read them?

Lesson 22 · Dev com IA v6.2 · INEMA.CLUB

Module 5 · Lesson 3 of 5

What matters lives in the project, not in the tool

An operations analyst and a colleague organize folders of different colored paper in a filing cabinet, with the laptop open on the desk.

You can set up the minimum files and reading order in the pilot folder for any assistant to follow.

Instructions, decisions, and tasks stored only in a tool’s memory are tied to it. When you switch assistants, you lose everything. Stored in text files in the project, they can be read by any assistant.

In 1 minute

  1. What belongs to you lives in text files in the project folder.
  2. AGENTS.md sets the rules and reading order.
  3. Folder names are agreed on; AGENTS.md tells the assistant to read them.

1Tool memory × project files

Each assistant stores things its own way: its own memory, conversation history, or an instruction file. This works as long as you use only that assistant.

When knowledge lives in text files inside the project folder, the assistant becomes just the one that carries out the work. Switching assistants, or model, no longer costs you your work.

Rafael had the clinic website rules only in CLAUDE.md and Claude Code’s memory, which Codex doesn’t look for. When he tried Codex, he had to explain everything again and forgot two rules.

Tied to the tool

Rules in an assistant’s memory. Decisions scattered across old conversations.

When switching: explain everything again.

In the project

Rules, task, and handoff in text files in the pilot folder.

When switching: the other assistant reads the files.

2The essential files

You don't need much. Five files cover the pilot: rules, objective, task, context, and handoff.

Carla set up the report folder like this in fifteen minutes. She already had two of them: the briefing and the handoff.

Carla's pilot folder
1 AGENTS.md
2 briefing.md
3 tasks
current.md
4 context
overview.md
5 handoffs
latest.md
  1. 1Stable rules and reading order.
  2. 2Objective, what's in and out of scope, criteria, and limits (lesson 5).
  3. 3The current task and what remains.
  4. 4Verified project facts, with the source.
  5. 5The handoff (lesson 21).

3AGENTS.md tells the assistant what to read

Folder names are just an agreement. No assistant opens "tasks" on its own. AGENTS.md tells the assistant what to read, with the reading order right at the top. The assistant usually follows it; the prime in lesson 22 checks whether it did.

Codex reads AGENTS.md on its own when starting. Claude Code reads CLAUDE.md. To have it follow the same rules, put the line @AGENTS.md in CLAUDE.md.

Rafael wrote the rules just once, in AGENTS.md. His CLAUDE.md has one line, and both assistants follow the same rules.

AGENTS.md · clinic website
1 Reading order: briefing.md, tasks/current.md, handoffs/latest.md
2 Rule: don't change the home page or the colors
3 Rule: nothing goes live without Rafael's approval
4 CLAUDE.md, alongside it, with one line: @AGENTS.md
  1. 1What to read, and in what order.
  2. 2The "Out of scope" that applies to every task.
  3. 3The final decision is made by a person.
  4. 4This way, Claude Code reads the same file.

Stuck here? That's normalIs your assistant neither of these? That's okay. For your first request, write "read AGENTS.md before anything else". In a regular chat that can't open the folder, paste AGENTS.md, then each file in the reading order, one after another, with the file name above it.

4Each file has an owner

A file without an owner becomes a mess: the AI changes a rule that was yours, or you forget to update the task. Agree on who changes what.

Carla made it clear in AGENTS.md: only she changes rules and decisions; the AI writes the handoff, and she checks it.

You decide

AGENTS.md, briefing.md, and the decisions. The AI can suggest a change; you approve it.

The AI writes, you check

handoffs/latest.md at the end of the session, and progress in tasks/current.md.

The two cards work together: one holds what’s stable, and the other records the day-to-day.

Practice now 0/3

Set up your pilot’s core

You’re done when the pilot folder has AGENTS.md with the reading order, tasks/current.md, and handoffs/latest.md. About 12 minutes, on a computer: this module requires a computer. Start with the three; add context/overview.md when you have verified facts to record.

These are your text files; nothing is deleted. To create a new folder in Windows: right-click › New › Folder. To create a new file: follow the steps for briefing.md in lesson 5. Don’t have the handoff from lesson 21? Create the file with the phrase "first session".

# AGENTS.md — <pilot name>

Reading order: 1) briefing.md 2) tasks/current.md 3) handoffs/latest.md
Read before taking action. Nothing is sent or published without my approval.

## Rules
- <what applies to all tasks>

## Owners
- AGENTS.md, briefing.md, and decisions: only I make changes.
- handoffs/latest.md: the AI writes it at the end of the session, and I check it.

You now have the pilot’s knowledge in a place any assistant can read.

Lesson cheat sheet

The project’s brain

  1. Text fileswhat’s yours stays in the folder, not in the tool.
  2. AGENTS.mdrules and reading order; CLAUDE.md points to it.
  3. Ownersyou decide the rules; the AI writes the handoff.

Your next step

You’ve set up the core that makes the pilot independent of any one assistant.

Read the AGENTS.md you wrote as if you were someone else. Remove one rule that only you understand.

Next lesson: with everything in files, move the pilot from one assistant to another without starting over.

Supplementary material · Migrate or stay assistant-agnosticMore on this topic. Not included in the lesson time.

The complete core

The agente-claude-codex kit uses seven places, each with an owner: AGENTS.md (rules and reading order), context/overview.md (verified facts, with source and date), context/current-state.md (what works and what is pending), context/sources.md (where each piece of information comes from), context/decisions/ (one accepted decision per file), tasks/current.md (objective, owner, completion criteria, next action), and handoffs/latest.md (the continuation for the next session).

Separate fact, preference, hypothesis, and decision

Mixing fact with hypothesis, or decision with preference, is what makes the assistant repeat an old mistake and contradict what has already been resolved. Information only moves from loose memory to approved context with your approval.

Move everything from one assistant to another

Moving instructions, skills, and memory from one assistant to another is the topic of the Claude → Codex course: audit, adapt, prove it in a new session, and hand off.

Sources: Claude → Codex Area — INEMA Events · Claude → Codex Course · agente-claude-codex kit guide

Lesson 23 · Dev with AI v6.2 · INEMA.CLUB

Module 5 · Lesson 4 of 5

Switch assistants without starting over

A developer closes one laptop and works on the other, with the project notes folder open between them, as if handing the work over.

You can move your pilot from one assistant to another with handoff and prime, and check that the new assistant understood before it starts working.

You’re going to switch assistants. You reach the usage limit, another assistant does one part better, or you want an outside perspective. Without a method, switching costs an hour of explanation. With the files from lesson 23, it takes ten minutes.

In 1 minute

  1. For the assistant leaving: handoff. For the one coming in: prime.
  2. Before letting it work, check the five prime questions (lesson 22).
  3. A wrong answer in prime usually means the file is vague. Fix the file.

1The switch is coming

No one uses one model forever. The account’s usage runs out in the middle of the week, the price changes, or another assistant does one part of the work better.

A good switch doesn’t depend on memory. It goes through the files: the handoff from the one leaving and the prime for the one coming in.

On Wednesday, the notice appeared that Rafael’s Claude usage limit was almost reached. He asked for the handoff right away and continued in Codex. If the limit had already run out, he would write the four handoff lines by hand.

Start over by explaining

Tell the new assistant about the project from memory. One hour, and two forgotten rules.

Handoff and prime

The new assistant reads the files, summarizes them with sources, and Rafael checks. Ten minutes.

2The switch, in four steps

The order matters. First the record, then the reading, then the check. Only then the work.

Carla has Codex and an AI chat. She switches the other way around: Codex creates the report and, for the review, she switches to the chat in a new conversation, pasting in the files.

Switch assistants
1 In the assistant you’re leaving: request a handoff and review it
2 In the assistant you’re switching to: run prime without changing anything
3 Review the summary with the five questions
4 Only then: "you can continue with the next action"
  1. 1Lesson 21.
  2. 2Lesson 22.
  3. 3The same five questions from lesson 22.
  4. 4You decide the order of the work.

3Prime also runs in the terminal in read-only mode

Anyone who uses the terminal can run Codex prime with an extra safeguard: read-only sandbox mode. Even if it wanted to, it couldn’t change files.

Rafael runs prime in the website folder. The response is saved in prime.md, and he reads it carefully before allowing the work to proceed.

Terminal · pilot folder
$ codex exec --sandbox read-only -o prime.md "Leia AGENTS.md e siga a ordem de leitura. Não altere nada. Responda as cinco perguntas do prime, com o arquivo de cada resposta."
...
$ cat prime.md

-o saves the last response to a file. cat displays that file on the screen. If the folder doesn’t use Git, add --skip-git-repo-check: without it, codex exec refuses to run.

In the app or extension, the same prime request from lesson 22 serves as this command.
And what about Claude Code in the terminal?

The equivalent is claude -p --permission-mode plan "<the same request>" > prime.md. -p responds once and exits. plan mode is the planning mode: it reads but doesn’t change files.

4The five questions to ask before letting it work

The assistant you switched to answers the five prime questions, the same ones from lesson 22. If an answer is wrong, check the files first. The problem is almost always vague wording in a text like the briefing or handoff, not the model.

Rafael’s Codex got the next action wrong: it said "create the form." The handoff only said "continue." Rafael rewrote the next action and ran prime again.

The five prime questions
1 What is the objective and acceptance criteria?
2 Name one rule and the file it comes from.
3 What was the last decision?
4 What is the exact next action?
5 Are there any conflicts between the files?
  1. 1Check it against briefing.md.
  2. 2Check AGENTS.md.
  3. 3Check the handoff.
  4. 4Check the handoff.
  5. 5If there is one, you decide which one takes precedence.

Stuck here? That's normalFive questions seem like a lot for each handoff? Start with 4. If the next action is correct, the others usually are too.

Practice now 0/3

Hand off the pilot to another assistant

You’re done when the assistant that joined has answered all five questions and you’ve marked each one correct or incorrect. About 10 minutes, on a computer: this module requires a computer.

Everything here is reading. Have only one assistant? Switch to a new conversation with it: the method is the same. Don’t have the lesson 23 AGENTS.md? Replace the first line with "read briefing.md and handoffs/latest.md".

Read AGENTS.md and follow its reading order.
Do not change any files or run anything.
Answer, citing the file for each answer:
1. What is the objective and acceptance criterion?
2. Cite one rule and the file it comes from.
3. What was the latest decision?
4. What is the exact next action?
5. Are there conflicts between the files?
Then wait for my instruction.

You can now switch assistants without losing what the pilot knows.

Lesson cheat sheet

Switch without starting over

  1. Leave and joinhandoff from the one leaving, prime for the one joining.
  2. Five questionsobjective, rule, decision, next action, conflict.
  3. Incorrect answerfix the file, not the assistant.

Your next step

You can now hand off one assistant’s work to another and check that it understood.

The next time the limit warning appears, switch using the four steps instead of waiting for the limit to reset.

Next lesson: a second opinion within the same assistant, and the lab that completes the cycle.

Lesson 24 · Dev with AI v6.2 · INEMA.CLUB

Module 5 · Lesson 5 of 5

Second opinion and the handoff → prime cycle

A developer and an operations analyst sit side by side: she reads a page of notes aloud while he checks on the laptop.

You can complete the full cycle in your pilot: handoff in one session, prime in a new session, review, and a second opinion.

In the previous four lessons, you saw the pieces: handoff, prime, project files, and switching assistants. Now they become a routine. And the same cycle gives you a second opinion as a bonus, without needing another assistant.

In 1 minute

  1. New conversation + prime + "point out flaws" = a second opinion from the same assistant.
  2. Claude Code has /advisor, which costs extra. Check before using it.
  3. The daily cycle: session, handoff, prime, review.

1Second opinion within the same assistant

A new conversation doesn’t carry over the previous one’s attachment to its work. With prime, it reads the briefing, the handoff, and the actual work. Then you ask for a critique: flaws, with the passage that shows each one.

It’s the cross-review from module 1, done with what you already have. Another model reviews it with even less attachment when you have access.

Carla has Codex and an AI chat. This time, she reviewed without leaving Codex: she opened a new conversation, ran prime, and asked for a critique.

New conversation, after prime

CarlaNow read relatorios/semana.md and point out flaws against the acceptance criteria in briefing.md, with the passage that shows each one. Don’t change anything.

AIFlaw 1: criterion 2 in briefing.md says "Each city appears once, with its total." In the totals table in relatorios/semana.md, the city of Pelotas appears in two rows.

The criterion comes from Carla’s briefing, since lesson 3. She opened the report and checked: there really were two rows.

2Claude Code’s /advisor: know what it does first

In Claude Code, there’s the /advisor command, followed by a model name. It connects a stronger model, which the main model consults during the session. The same command switches or turns off the advisor.

Two things to keep in mind. Claude Code itself warns that the advisor charges separate usage credits. And Codex doesn’t have this command. A new conversation with prime works in any assistant, at no extra cost.

Rafael typed /advisor in Claude Code and read the billing notice. For the pilot, he stuck with the new conversation; the advisor is for a difficult error.

Works in any assistant

New conversation, prime with the files, request for flaws with evidence.

Only in Claude Code, with a cost

/advisor: a stronger model consulted during the session, charged separately. Check the cost first.

Start with the "Works in any assistant" card. Use the other when the cost is worth it.

3The daily cycle

Put it all together, and a workday with AI has four stages. It’s the same in any assistant because it depends only on text files.

Rafael ends every session with the handoff and starts every session with the prime. He says he’s no longer afraid to close the conversation.

The daily cycle
1 Session: work on the next action
2 Handoff: save and check handoffs/latest.md
3 Prime: a new conversation reads and summarizes without changing anything
4 Check: five questions, then get to work
  1. 1The work itself.
  2. 2Always before you close.
  3. 3Always when you open.
  4. 4Go back to 1.

Stuck here? That's normalDoes this feel like too much ritual for a half-hour task? For short tasks, a one-line handoff is enough: the exact next action.

4The lab: the cycle in your pilot

The module lab brings all five lessons together in one run. The result is a ready-to-resume pilot that any assistant can pick up, any day.

Carla did the lab on a Friday and resumed on Monday. The prime got all five questions right on the first try.

Without the cycle

Every Monday starts with “where was I again?”, and the AI repeats questions.

With the cycle

Every Monday starts with a checked summary and the exact next action.

Net gain: ten minutes for the handoff and prime instead of half an hour of reconstruction.

Practice now 0/4

Lab: close the cycle in the pilot

You’re done when a new conversation has answered all five questions correctly and checked each acceptance criterion against the passage, marking it “fail” or “ok.” No failures is also a result. About 12 minutes, on a computer: this module requires a computer.

Everything is reading, except the handoff, which is your own file. Haven’t done the earlier lessons? Use briefing.md from lesson 5 and a one-line handoff. If the AI starts changing anything, click the stop button.

Read AGENTS.md and follow its reading order.
Do not change any files or run anything.
Answer, citing the file for each answer:
1. What is the objective and the acceptance criterion?
2. Name one rule and the file it comes from.
3. What was the last decision?
4. What is the exact next action?
5. Is there a conflict between the files?
Then, for each acceptance criterion in briefing.md,
say “ok” or “fail,” with the passage from the work that shows this.

You’ve finished module 5 with a pilot that any assistant can resume and review.

Lesson cheat sheet

Cycle and second opinion

  1. Second opiniona new conversation, prime, and request for failures with passages.
  2. New commandcheck with your assistant before using it.
  3. Daily cyclesession, handoff, prime, check.

Your next step

You already work with AI without depending on the memory of a conversation or a tool.

Tomorrow, start the day with the pilot’s prime and time how long it takes to get to the first action.

In module 6: how much did all this cost? Model, effort, and budget by stage.

Supplementary material · Handoff and prime in the kitsFurther reading on this topic. Not included in the lesson time.

Use Both Level 6

The Use Both kit calls this “handoff and prime”: the handoff saves the state to a portable file; prime has the next assistant read and verify that state before continuing. In the agente-claude-codex kit, each handoff becomes a dated snapshot in handoffs/history/, which is never overwritten, and handoffs/latest.md is the copy of the most recent one.

The /advisor in Claude Code

The /advisor command, followed by a model name, sets an advisor model that the main model consults during the session. The advisor needs to be at least as capable as the main model, and its use is billed separately in credits. The INEMA summary presents using it to investigate bottlenecks as guidance to verify: test it on one of your tasks and compare it with a new conversation using prime.

Sources: Codex + Claude Area — Eventos INEMA · Claude → Codex Area — Eventos INEMA · Use Both Guide

Lesson 25 · Dev with AI v6.2 · INEMA.CLUB

Module 6 · Lesson 1 of 5

Choose based on the task, not the ranking

A developer compares two printed answers on the desk, with the criteria sheet beside them, marking with a pen what passed in each one.

You can compare two models on one of your tasks, using your own acceptance criteria, and say which one did a better job on that task.

Every week, a new list appears of the “best AI right now.” It answers a general question. Your question is different: which one does this task better, on your account, today.

In 1 minute

  1. A ranking is an average across many tasks. Yours is just one.
  2. Compare using the same task, the same briefing, and the same standard.
  3. Count how many criteria passed, not the tone of the response.

1A ranking doesn’t answer your question

A model that performs well on average can perform poorly on your task. And names change quickly. In September 2026, the most frequently mentioned models in this course’s sources were Claude Opus 5.5 and GPT-6 Astra.

A few months from now, they’ll be different. The way you choose stays the same.

Rafael read that one model “was the best for code.” On the clinic form, that model accepted letters in the phone number field.

The question you ask

RafaelWhat’s the best model for programming?

AIIt depends on how you use it. Several models stand out at programming, each with different strengths.

A general question gets a general answer.

RafaelSame briefing for both models. Which one passed all three acceptance criteria in the form?

Rafael’s worksheetModel A: 3 out of 3. Model B: 2 out of 3; it failed to reject letters in the phone number field.

The answer comes from your review, not someone else’s opinion.

Tap or click both buttons in the panel and compare the two questions.

2Same task, same briefing, same standard

A fair comparison changes just one thing: the model. The request, the material, and the criteria stay the same. If you change two things, you won't know which one made the difference.

The standard is what you already have: the acceptance criterion in your briefing. Didn’t make the lesson 5 briefing? Write three yes-or-no criteria for the task now.

Carla ran both rounds in Codex, using the same spreadsheet folder and the same briefing. Between rounds, she changed only the model in the selector near its name.

Carla's comparison
1 Same request: briefing.md for the weekly report
2 Same material: a copy of this week's spreadsheets
3 Same yardstick: the three acceptance criteria
4 Only change: the model, in the same tool
  1. 1The same file for both.
  2. 2A copy for each, so one doesn't change the other's.
  3. 3Written before seeing the responses.
  4. 4The only difference between the two rounds.

Stuck here? That's normalNot sure where to change the model? In Claude Code, the /model command lists the models on your account; in the app, look for the selector near the model name. Do you only have access to one? Compare the same model in two new conversations with the same request: you'll still learn to use the yardstick.

3Count the criteria, not the tone

A polished response, with headings and confident wording, seems better. The yardstick asks something else: did it pass each criterion or not?

The first model's report had a chart and polished text, but listed one city twice. The second was plain and passed all three criteria.

Looks better

Appearance: chart, headings, elegant summary.

Yardstick: 2 out of 3. One city appears twice.

Passed the yardstick

Appearance: a simple table.

Yardstick: 3 out of 3. Total matches, cities are unique, source is at the bottom.

Net gain: the choice comes from the record, not the first impression.

4The choice applies to that task, on that date

The comparison result applies to that type of task, in that month. Note the task, model, criteria that passed, and the date. When the model, price, or task changes, compare again.

Rafael keeps one line per comparison in a file named escolhas.md, in the pilot folder. In three months, he'll know which model to use for each type of request.

What to note

28/09/2026 · contact form · model A 3 out of 3, model B 2 out of 3 · keep A.

When to compare again

New model, new price, or a different kind of task.

One line per comparison already becomes a history.

Test yourself

Carla wants to know which model to use for the weekly report. What does she do first?

Practice now 0/3

Which model is right for this task?

You’re done when you’ve answered the three questions about the case. About 10 minutes, on paper or in a notes app.

This is just a case to read. If you’re unsure, go back to the criteria: count what passed, not what seems better.

The case. Rafael gave the same briefing to two models for a client’s pricing page. Criteria: all three plans appear; the button leads to payment; the page doesn’t scroll sideways on a phone. Model A delivered a beautiful page: all three plans appear and it doesn’t scroll sideways, but the button didn’t open anything. Model B delivered a simple page that passed all three. He used a slightly different request for B.

Show answer key

1. A: 2 of 3. B: 3 of 3. 2. It wasn’t fair: the request for B was different. 3. Repeat A’s test using the same request as B, compare again, and only then write down the dated entry.

You’re already comparing models using the criteria for your task, not their reputation.

Lesson cheat sheet

Choosing a model

  1. Your taskthe ranking is an average; your question is your own.
  2. One differencesame briefing and materials, only the model changes.
  3. With a daterecord your choice and repeat the comparison when something changes.

Your next step

You’re already choosing a model based on evidence from your own task.

Create the file escolhas.md in the pilot folder and write the first line, even if it’s only for the model you’ve used so far.

In the next lesson: there’s another control within the same model: effort. When is it worth increasing, and when is it a waste?

Supplementary material · Starting route, not a rankingFurther reading on the topic. It doesn’t count toward lesson time.

What the sources say

The Codex + Claude area of Eventos INEMA notes that Mark Kashef’s video mentions Claude Opus 5.5 and GPT-6 Astra. INEMA’s summary from September 27, 2026 mentions the same names, but points out that models, access, and billing vary. The recommendation is to compare the quality of the result on the specific task instead of setting a universal rule for which one is better.

The route card in the Use Both kit (Claude/Opus plans, Codex critiques; Claude builds, Codex reviews what changed) is an initial preference recorded in Mark Kashef’s video. The Codex + Claude area of Eventos INEMA cautions that these are editorial choices, not a measured ranking. Try them and stick with the route that produces better evidence for your task.

Sources: Codex + Claude area — Eventos INEMA · Use Both guide

Lesson 26 · Dev with AI v6.2 · INEMA.CLUB

Module 6 · Lesson 2 of 5

High effort where it matters, low effort for routine tasks

An operations analyst sticks three colored notes onto a sheet with five drawn stages, marking how much effort each stage deserves.

You can mark the effort as high, medium, or low for each step in your pilot cycle, and say why.

Many people set effort to the maximum for everything, to be safe, or to the minimum, to save money. Both choices have a cost: one in time and account usage, the other in errors in the difficult parts.

In 1 minute

  1. Effort is a different control from the model: it determines how much the model thinks before responding.
  2. Planning and solving a difficult problem call for medium or high effort. Routine tasks call for less.
  3. More effort won’t bring back the missing file.

1Model and effort are two controls

The model is which AI does the work. Reasoning effort is how much it analyzes before responding. You can change one without changing the other.

Rafael uses the same model all day. What he changes is the effort: high when planning the access area, low when changing button text.

Model

Which AI does the work. Switch when the comparison in lesson 26 shows that another one delivers better results.

Effort

How much it analyzes. Increase or decrease it at each step while keeping the same model.

Two different controls. Today’s lesson is about the second one.

2When to increase effort: planning, difficult problems, thorough reviews

Increase effort when a mistake is costly or when there are many parts to consider together. Planning, tracking down an error no one can find, and reviewing a major change all call for more effort.

The diagram walks through the whole cycle, from the briefing to the handoff.

Carla asked for high effort only when planning how to combine spreadsheets with different columns. The rest of the work was done at medium effort.

Pilot cycle · effort at each step
1 Define the briefing: you write it; the AI only reviews it, at medium effort
2 Plan and critique: medium or high effort
3 Build: medium effort; repetitive parts, low effort
4 Review and verify: medium effort; major change, high effort
5 Record the handoff: low effort
  1. 1The criteria are yours; the AI only points out gaps.
  2. 2A bad plan ruins everything that follows.
  3. 3Follow the plan; increase effort only if you get stuck.
  4. 4The bigger the change, the more effort the review needs.
  5. 5Summarizing what already exists is routine.

3When to lower effort, and when effort won’t help

Routine tasks, such as renaming, formatting, or changing text, usually turn out the same with less effort, take less time, and use less. Check against the criteria. But “the minimum for everything” isn’t a safe rule: the difficult parts come out worse.

And there’s a limit: if the spreadsheet or the criteria are missing, no amount of effort will solve it. First the materials, then the effort.

Carla raised the effort to the maximum when the report came out wrong. It was still wrong: Thursday’s spreadsheet was missing from the folder.

Maximum effort, without the spreadsheet

Request: "Think carefully and redo the report."

Result: it took three times as long, and the total was still wrong.

Medium effort, with the spreadsheet

Request: the same one, with Thursday’s spreadsheet in the folder.

Result: total equal to the sum of the spreadsheets.

Net gain: what was missing was material, not reasoning.

Stuck here? That's normalNot sure if the task is difficult? Start at medium. Increase it only if the answer fails a criterion because it needs more analysis, not because material is missing.

4Where to choose the effort

Each tool shows this control differently, and some don’t show it at all. In the app or extension, when available, it’s near the model name. In chat, it may appear as "reasoning," "think more," or "extended thinking"; sometimes it’s just on or off. In the terminal, it goes in the command itself.

Rafael asks for a critique of the plan with high effort and the list of button labels with low effort, each in its own command.

Where to find the control
1 App or extension: near the model name, when available
2 Chat: "reasoning," "think more," or "extended thinking"
3 Terminal: in the command itself, as in the box below
  1. 1Sometimes there are levels; sometimes, a menu.
  2. 2Often, it’s just on or off.
  3. 3The level is written alongside the request.
Can’t find the control? The tool decides on its own. The rest of the lesson still applies: materials first, criteria second.
Terminal
$ codex exec -c model_reasoning_effort=high "Leia plan-v1.md e aponte falhas com evidência."
$ claude -p --effort low "Liste os textos de botão que aparecem em contato.html."

In Codex, the effort is model_reasoning_effort (from minimal to high; some models accept more), written in Codex’s config.toml or in the command with -c. In Claude Code, it’s --effort (low, medium, high, xhigh, or max).

The same model, two effort levels, each at the step that calls for it.

Test yourself

The AI got the report total wrong. Before increasing the effort, what do you check?

Practice now 0/3

The effort level for each step of your pilot

You’re done when each of the five steps in your pilot has a level and a reason. About 8 minutes, on paper or in a notes app.

You only write things down; nothing runs. Don’t have a pilot? Use Rafael’s form or Carla’s report as an example.

Step | Effort | Why
1. Define the briefing | <low/medium/high> | <reason>
2. Plan and critique | < > | < >
3. Build | < > | < >
4. Review and verify | < > | < >
5. Record the handoff | < > | < >
Show answer key

A reasonable answer: medium for the briefing (review only); high for planning and critique (a bad plan spoils everything that follows); medium for building; medium for review, high if the change is large; low for the handoff. Step 4 is a good time to check the materials: is a file missing? Fix that before increasing the effort.

You already distribute effort across the stages instead of using one fixed setting for everything.

Lesson cheat sheet

Effort by stage

  1. Two controlsModel is which AI; effort is how much it analyzes.
  2. Increase itfor the plan, the difficult problem, and the major review.
  3. Materials firsteffort doesn't replace a missing file.

Your next step

You already choose the effort based on the stage, not out of habit.

For your next routine request, lower the effort by one level and check the criteria to see whether the result stayed the same.

In the next lesson: how much did each stage of your pilot cost, in time and account usage?

Supplementary material · What the summary says about effortFurther reading on the topic. Doesn't count toward lesson time.

Effort, according to the sources

The INEMA summary from September 27, 2026: planning and difficult problems may justify medium or high effort; routine tasks can use less; a complex review may also need more. Using the minimum level for everything else is not a safe rule.

The message that inspired the course sums up the reason: adjusting the model and effort for each stage avoids spending the resources of a complex task on simple actions.

Sources: executive summary and message from 09/27/2026, in the course source files, folder docs/.

Lesson 27 · Dev with AI v6.2 · INEMA.CLUB

Module 6 · Lesson 3 of 5

Cost without results means nothing

A developer records the time and usage for one stage in a small notebook, with a desk timer beside the laptop.

For each stage of your pilot, you can record the time, account usage, and problems found.

"It was expensive" and "it was worth it" are usually impressions. Without a simple measure, the decision in lesson 30 becomes a matter of opinion. Measuring means recording three numbers for each stage, as you go.

In 1 minute

  1. For each stage, record: your time, account usage, and problems found.
  2. Your time counts, including the time you spend checking.
  3. A problem found before delivery is the result the cost paid for.

1One row per stage

The measure fits in a spreadsheet with eight columns. One row for each stage of the cycle, filled in as soon as the stage ends. After a week, no one remembers how long it took.

Rafael opened the piloto-custos spreadsheet in the pilot folder and filled in the plan row as soon as the critique ended.

piloto-custos · columns
1 Stage
2 Model and effort
3 Your time (min)
4 Old method (min)
5 Usage before
6 Usage after
7 Problems found (by whom)
8 Criteria that passed
  1. 1Which step in the cycle.
  2. 2With which configuration.
  3. 3Request, wait, read, and check.
  4. 4How long it would take without AI; an estimate is fine, labeled "estimated".
  5. 5What the usage page shows before.
  6. 6And after the step.
  7. 7Review, test, or you.
  8. 8How many out of how many.

2Your time counts

AI responds in two minutes, but you spent twenty reading and checking. The time that counts is yours, from the request to the check. That is what your work hour pays for.

Carla used to record only the wait for the response: three minutes. When she included checking the total, the step took twenty-five minutes.

Just the wait

Recorded: 3 minutes.

Conclusion: "AI prepares the report in three minutes."

From the request to the check

Recorded: 25 minutes, including 15 minutes of checking.

Conclusion: still less than the two hours by hand.

Net gain: the honest comparison is with the old way, including the check. That’s why the spreadsheet has the "Old way" column.

3Account usage: the usage page before and after

AI usage is measured in tokens. You don’t need to count them: just open the account’s usage page before and after the step and record the difference, in the unit it shows.

The usage page is in the account settings on the service’s website (in English, "Usage"). Lesson 4 showed you where to look. Company account? Ask the person who manages it for the usage page.

Rafael recorded the percentage of his subscription’s weekly limit before and after the plan review. The review at high effort used much more than the text changes.

Rafael’s entry · plan review

BeforeWeekly usage: 18%

AfterWeekly usage: 23%

In the spreadsheetPlan review · high effort · 5 points of the weekly limit · 2 issues found

The exact number matters less than comparing steps with each other.

Stuck here? That's normalDoes your usage page not show a number for each task? Record what it shows, before and after. If nothing shows up, just track the time: that’s enough to compare. If other people use the account at the same time, that can distort the before and after: write "shared account" next to it.

4Issues found are the result

Cost alone doesn’t tell you whether it was worth it. The other side of the equation is the issues found before delivery, and by whom: through cross-review, through testing, or by you.

During the pilot week, the review by another model found the city listed twice, and the manual test found the wrong total. Both would have gone to the board.

Cost

Your time and account usage at each step.

Result

Problems found before delivery, criteria that passed.

An expensive step that found two errors may be worth more than a cheap one that found nothing.

Test yourself

The review step used the most of your account. What do you look at before cutting it?

Practice now 0/3

Your pilot cost spreadsheet

Done when the spreadsheet has the columns and at least two completed rows. About 10 minutes, on a computer.

It's your spreadsheet; nothing is sent. Don't remember the time for an earlier step? Write an estimate and mark it "estimated": from now on, note it down right away.

Columns (one per cell, in the first row):
Step | Model and effort | Your time (min) | Old way (min) | Usage before | Usage after | Problems found (by whom) | Criteria that passed

Carla's example:
Weekly report | Codex, medium | 25 | 120 | 40% | 42% | city listed twice (review), wrong total (test) | 3 out of 3

You now have cost and result side by side, step by step.

Lesson cheat sheet

Measure the pilot

  1. Right awayone row per step, as soon as it ends.
  2. Your timefrom the request until the review.
  3. Resultproblems found before delivery, and by whom.

Your next step

You now measure the pilot with numbers, not impressions.

For your next task with AI, note the start and end times of your review. It takes ten seconds.

Next lesson: your account price and limit change. Where can you check, and what changes between a subscription and pay-per-use?

Lesson 28 · Dev with AI v6.2 · INEMA.CLUB

Module 6 · Lesson 4 of 5

Prices and limits change: check your account

An operations analyst, looking thoughtful, checks the AI account settings on the laptop, with the stack of this month's invoices beside it.

You can look at your account and tell whether you pay by subscription or by usage, and what the limit is. You also note where you checked and the date.

Prices, limits, and included tools change from one month to the next and from one account to another. What you saw in a video may not apply to yours. The only reliable source is your account, on that day.

In 1 minute

  1. Subscription: fixed amount, with a usage limit per period. Pay-per-use: you pay for what you use.
  2. Note the type, the limit, where you saw it, and the date. Never the password or key.
  3. OpenRouter and open models are options to check, not a recipe.

1Subscription and pay-per-use

With a subscription, you pay a fixed amount and have a usage limit for each period. When you reach the limit, you wait for the next period or, if the plan offers it, buy extra usage. With pay-per-use, each consumed token is added to the bill, and you set the spending cap, if there is one.

Pay-per-use usually comes through an API, with a secret key.

Rafael uses Claude Code through a subscription. A client asked for an automation that calls the AI on its own every night: that one is pay-per-use, with a key and spending cap in the client’s account. That way, the cost stays with the person using it, and his personal subscription isn’t used for the client’s work.

Subscription

Pays: a fixed amount per month.

Limit: usage per period; when you reach it, you wait or buy extra, if the plan offers it.

Pay-per-use

Pays: for what you use.

Limit: the spending cap, if you set one.

Neither is generally better. It depends on the volume and who uses it.

2An account record, dated

Write down what your account shows today: type, limit, where you saw it, and the date. When something changes, you’ll notice by comparing it with the record. The password and key never go in the record or in a project file.

Carla made a record for her company’s Codex account. She found out that her colleague used one account and she used another, with different limits.

conta-ia.md · Carla’s record
1 Service: Codex, through the company account
2 Type: subscription
3 Limit: "40% of the week used, renews Monday"
4 Where I saw it: account settings › Usage · 28/09/2026
5 Password and key: never here
  1. 1Which service and whose account it is.
  2. 2Subscription or pay-per-use.
  3. 3Copy what it says there, without interpreting it. If it only shows a percentage, write it down that way.
  4. 4Where and when you checked.
  5. 5Keep credentials in the password manager.

3OpenRouter and open models: check first

OpenRouter is a service that gives you access to several models with a single account, paying per use. Open models can run on a service like this or on your own computer.

They can be good options for certain tasks. But in this course, treat them as options to check: the current price, quality for your task using the lesson 26 yardstick, and where the data goes.

Carla considered a cheaper open model for the report. She stopped first: the spreadsheets contain client names, and she didn’t know where the service would store the data.

Switch based on price

“This one is cheaper, so I’ll use it.” No testing, and no idea where the data goes.

Check first

Today’s price, test it against the rubric on a task with no customer data, and read the data policy.

Stuck here? That's normalNot sure if you can send customer data to another service? When in doubt, don’t send it. Ask the person who handles this at your company, and test first with made-up data.

4“AI will get more expensive” is a hypothesis

Some people say AI use is subsidized today and will get more expensive. Maybe. But that’s a market prediction, not a proven fact.

This course’s method works without it: measure cost and results for each task, and keep project knowledge in files any assistant can read. If the price goes up, you already know where to cut. If it goes down, you know where to expand.

Rafael didn’t change anything because of the prediction. He changed things when the spreadsheet from lesson 28 showed two things. Critique at high effort was worth it. Switching texts wasn’t.

Decide based on the prediction

“It’s going to get more expensive, so I’ll use the minimum for everything.”

Decide based on the measurement

“The spreadsheet shows where the effort pays off. I’ll keep it there.”

Test yourself

A video says the subscription limit is a certain number of hours per week. What do you do?

Practice now 0/3

Your account record

You’re done when the file conta-ia.md has the type, limit, where you checked, and the date. About 8 minutes, on a computer.

You only check your account and write it down; nothing is purchased or changed. Never copy a password or key into the record. Can’t find the limit? Write “not shown” and the date.

# AI Account
Service: <which one, and whose account it is>
Type: <subscription or pay per use>
Limit: <what the page shows; if it only shows a percentage, write down the percentage>
Spending cap: <set? how much?>
Where I saw it: <path on the site> · <date>
Password and key: never here

You now know what your account charges and what its limits are, with the date and source.

Lesson cheat sheet

Account and price

  1. Typesubscription with a limit for a set period, or pay per use with a spending cap.
  2. Dated recordwhat the account shows today; never include credentials.
  3. To checkOpenRouter, open models, and price predictions.

Your next step

You now check the price and limit at the source, not based on hearsay.

Set a reminder for one month from now: reopen the usage page and update the record.

Next lesson: with the cost, results, and account details in hand, decide whether the pilot becomes part of your routine.

Supplementary material · What the summary marks as “to check”Further reading on the topic. Not included in the lesson time.

Additional guidance

The INEMA summary from September 27, 2026, separates what was verified from what was suggested in the conversation. OpenRouter and open models can be part of the strategy when they fit the project, but they are not recommendations verified by the page that was analyzed. Confirm them in your environment and with the tools available before making them standard practice.

About future pricing

The idea that AI is subsidized today and may cost more per use in the future is treated as a market hypothesis. The practical recommendation already stands without that prediction: measure cost and results per task, and keep the project knowledge portable.

Sources: executive summary dated 27/09/2026, in the course source files, folder docs/.

Lesson 29 · Dev with AI v6.2 · INEMA.CLUB

Module 6 · Lesson 5 of 5

Expand, adjust, or stop: the pilot decision

A developer and an operations analyst talk at a small meeting table in front of the pilot's one-page report, deciding on the next step.

You can deliver the one-page report for your pilot and write the decision: expand, adjust, or stop, with the reason.

A pilot without a decision at the end is just an experiment. The report brings together what you measured throughout the course and answers one question: is it worth doing it this way again, for more tasks?

In 1 minute

  1. One-page report: task, criteria, problems found, cost, and decision.
  2. Three possible outcomes: expand, adjust, or stop. All are results.
  3. The decision is yours, with a date and the next review scheduled.

1The report fits on one page

Everything you produced in the course becomes six lines. They cover the task from the briefing, the acceptance criteria that passed, the problems found, and who found them. Then the cost and comparison from the spreadsheet in lesson 28, and the decision.

Rafael put together the report for the clinic form in fifteen minutes, copying from the files in the pilot folder.

relatorio-piloto.md · Rafael
1 Task: clinic contact form
2 Criteria: 3 of 3 passed the test
3 Problems found: 2 in the plan review, 1 in the test
4 Cost: time and account usage, from the piloto-custos spreadsheet
5 Comparison: the same type of task done the old way
6 Decision and date of the next review
  1. 1From briefing.md.
  2. 2From the acceptance criteria in module 4.
  3. 3From the spreadsheet, problems column.
  4. 4From the spreadsheet in lesson 28.
  5. 5From the spreadsheet's "Old way (min)" column; an estimate is fine, marked "estimated".
  6. 6What this lesson teaches you to write.

Stuck here? That's normalDid you skip any step in the pilot? Write "not done" on that line. An honest report with gaps is worth more than a complete one you made up.

2Three outcomes, all legitimate

Expand means using the method for more tasks of the same type. Adjust means repeating the pilot while changing one thing. Stop means concluding that, for this type of task, the old way is better. Stopping is a result too.

Carla expanded it to the monthly report. Rafael adjusted it: the cross-review of the plan found too much, but the one for swapping texts found nothing, so he’s going to remove it.

When to choose each outcome
1 Expand: the criteria passed and the cost fits the budget
2 Adjust: it passed, but one step cost time without finding anything
3 Stop: it didn’t pass, or it cost more than the old way
  1. 1More tasks of the same kind.
  2. 2Change one thing and repeat.
  3. 3Write down why and move on without guilt.

3The decision is yours, with a date

The model can summarize the report and point out numbers that don’t add up. The person responsible for the work decides whether to expand. Write down the decision, the reason, and when you’ll review it.

Carla wrote: "Expand it to the monthly report. Reason: 3 of 3 criteria and 25 minutes compared with two hours. Review in 30 days."

No decision

"The pilot was interesting. We’ll see."

After one month: no one remembers what was measured.

Decision in writing

"Expand. Reason: 3 of 3 criteria, 25 minutes compared with 120. Review in 30 days."

After one month: she compares it with the new month.

4Where to go next

The course cycle is the beginning. Three open INEMA paths continue from here.

The Use Both guide covers the six levels of using both together and includes ready-to-use prompts. claudex automates a debate between Claude and Codex about a large plan. And the Claude → Codex course teaches you to store project context in files that any assistant can read.

claudex runs in the terminal, with Claude Code and Codex ready to use. The other two are for people who use the app.

Rafael chose Use Both. He wants to give the assistant a goal with a success test and a stop limit, as in lesson 19. Carla chose the Claude → Codex course: her team uses different assistants.

Next steps
1 Use Both: six levels and ready-to-use prompts
2 claudex: automatic debate about a large plan (terminal)
3 Claude → Codex: portable context in files
  1. 1For anyone who wants to deepen their use of both together.
  2. 2For people who often plan big things.
  3. 3For people who switch assistants or work on a team.

Practice now 0/4

Your pilot report and decision

Ready when relatorio-piloto.md has all six lines and the decision with a review date. About 12 minutes, on a computer.

This is your file; nothing is sent. Missed a step? Write "not done." If you want, ask an assistant to check whether the report's numbers match the spreadsheet. The decision is still yours.

# Pilot report
Task: <from briefing.md>
Criteria: <how many passed out of how many>
Issues found: <how many, and by whom: review, test, you>
Cost: <your time and account usage, from the spreadsheet>
Comparison: <the "Old way (min)" column in the spreadsheet; if it's an estimate, write "estimated">
Decision: <expand, adjust, or stop> · Reason: < > · Review on: <date>

You closed out the pilot with a decision another person can understand and verify.

Lesson cheat sheet

Close out the pilot

  1. One pagetask, criteria, issues, cost, comparison, and decision.
  2. Three outcomesexpand, adjust, or stop, each with a reason.
  3. Review scheduledthe decision has a date for review.

Your next step

You finished the course with a method tested on your own task: one does the work, the other checks it, and you decide.

Show the pilot report to someone on your team or to a client. It takes five minutes and tests whether the report explains itself.

Continue with Use Both, claudex, and the Claude → Codex course.

Lesson 30 · Dev com IA v6.2 · INEMA.CLUB