PTENES
Skip to content
MODULE 3.1

🎚️ Right model, make your quota go further

Always using the strongest model is like asking the company director to check a stamp. It works, but it uses up your quota. The ROTEAMENTO.md from the kit says which model to call for each type of task.

6
Topics
~35
Minutes
Base
Level
Foundation
Type
0 of 60%
1

Understand why quota isn’t dollars

The kit runs through subscription for Claude Code and Codex. You pay for the monthly plan, period. You don’t get billed per call.

But the subscription has a usage limit per period: the quota. Each request to the agent uses up a bit of it. When it runs out, you wait for the period to roll over.

The first line of the runtime/ROTEAMENTO.md summarizes the module: "Don't always use the most powerful model. Each call uses up some of your subscription quota."

🆕 New here? Three words from this module

  • Quota — the usage limit on your subscription for a given period. It’s not money: it’s how much more you can ask for before you have to wait.
  • Model — the AI's "head" that responds. Claude Code has several (opus, sonnet, haiku…), Codex too. The more powerful ones think better and use more of your quota.
  • API equivalent — how much the same task would cost if billed by API usage. It’s a benchmark for comparison, even if you don’t pay that amount.
Period quota (your subscription) m m executor m super for a simple task super again… m = smaller model: short block · executor: medium block · super: long block quota ran out: wait for the period to roll over

How to read the diagram: the purple bar is the quota. The short blue blocks are requests to the smaller model; the long red ones are the stronger model used unnecessarily. Two unnecessary calls to the top tier use more than several correctly routed tasks. The sizes are illustrative, not measurements.

✓ Quota (what you have)

  • ✓ Comes with the subscription you already pay for
  • ✓ It ends, but comes back in the next period
  • ✓ Gets better results when each task uses the right model
  • ✓ It’s what the entire kit uses

✗ API dollars (what you don’t need)

  • ✗ Separate account, billed per use
  • ✗ No refunds: what you spent, you paid
  • ✗ No kit recipe depends on it
  • ✗ In the course, it appears only as a benchmark

💡 A real number to give you an idea

The kit's CHANGELOG records that the three-role team (module 3.2), creating a one-sentence file, spent the equivalent of ~US$ 0,73 of API costs. With a subscription, this is quota, not a charge. But it shows that even a tiny task has a cost when it calls large models.

💳
Subscription

monthly plan

⏳
Quota

limit per period

📏
Equivalent

benchmark, not a bill

🎚️
Routing

model by task

2

Read the four model levels

The kit doesn’t use standalone model names. It uses model level: super, top, executor, and smaller. Each level says when to use and point to the model in Claude Code, Codex, and on your computer.

This is the table from runtime/ROTEAMENTO.md, copied from the kit:

LevelWhenClaude CodeCodexLocal (free, Ollama)
superrare: difficult problem, major decisionfable · high effortgpt-6-astra · high—
topplan, decide, review the hard partsopus · mediumgpt-6-astra · mediumqwen3.8:27b
executordo everyday workopus · low / sonnetgpt-6-sol · lowqwen3.8:27b
smallertriage, format, checkhaiku / sonnetgpt-6-lunallama3.2

What to look for in the table: the "When" column is the one you use day to day. The three on the right only tell you which name to type in each tool. Note that gpt-6-luna, the bridge’s default in module 2.1, is the lower level.

🆕 New here? "Effort" and "Local"

Effort is how much the model thinks before responding (low, medium, high). The same model with high effort thinks more and uses more quota. Local is a model that runs on your own computer through Ollama: it uses no quota, but depends on your machine being powerful enough. The doctor shows whether you have Ollama (line ollama, marked as optional).

superrare topplan, decide, review the hard parts executorthe day-to-day work smallertriage, format, check — start here more quota per call ↑ most used ↓

How to read the diagram: the width of a step represents how often it should appear in your day. The smallest one (purple, glowing) is the broad foundation. The super model, red and narrow, is an exception: if it appears all the time, something is wrong with the routing.

🔺
super

big decision

🧠
top

plan and review

🛠️
executor

do

🔎
smaller

triage and check

3

Start with the smallest thing that solves it

Routing rule 1: "Start with the smallest option that works; move up only if it fails." O AGENTS.md from the kit repeats this for both agents: "use the smallest model that gets the job done".

In practice, you move up one level at a time, and only for a reason: the lower level got something wrong, let something through, or couldn’t handle the size.

1

Smallest fix: format and check

Clara wants the list of available times formatted as a message. That’s formatting: a lower level.

2

Did it fail? Executor

Sônia asks for the total per client from the ERP CSV. If the smaller model gets the sum wrong, move up to the executor. She checks the correct total by hand (module 2.3).

3

Decision? Top

"How do I connect the ERP without an API to the agent?" involves planning and decision-making. That’s where the top tier pays off.

4

Very: only if the top gets stuck

A truly difficult problem, a big decision. Rare, as the table says.

The Codex bridge is set up this way from the start: without you saying anything, it uses gpt-6-luna, the lower level. To move up, you change the variable CODEX_MODELO only for that call.

🎯 Objective: call Codex at the lower level and learn how to step up

In the terminal, from inside the kit folder (module 2.1 bridge, already with chmod +x):

runtime/pontes/codex-exec.sh "Responda apenas PONG"

Result confirmed in CHANGELOG 0.1.0:

PONG

To raise Codex to executor level for this call only, put the variable in front:

CODEX_MODELO=gpt-6-sol runtime/pontes/codex-exec.sh "Responda apenas PONG"
How to verify: the first call returns PONG using the standard gpt-6-luna. The variable applies only to the line where it appears; the next call goes back to the lowest level.

💡 Move up with a written reason

When you move up a level, write down why in one line ("the kid got the total for Padaria Lua wrong"). After a few weeks, these lines show which tasks should start at a higher level.

⬇️
Start at the bottom

rule 1

⬆️
Move up if it goes wrong

one step at a time

🌙
gpt-6-luna

bridge pattern

🔧
CODEX_MODELO

switches to a call

4

Ask another model to review it

Rule 3: "Review by another model (e.g., Claude does it, Codex reviews — recipe R1) catches errors the same model doesn’t see."

It’s the same idea as asking a colleague to reread your writing. The person who wrote it tends to read what they meant to write, not what’s actually there. A model from another company has different blind spots.

Claude does the work codex-exec.sh Codex reviews read-only (N4) the review comes back you receive the two opinions

How to read the diagram: the blue arrow is the bridge from module 2.1 carrying the request to Codex. The dashed purple curve is the review coming back. Codex only reads: Claude decides what stays, and you have the final say.

🎯 Objective: get a second opinion from another model without leaving Claude

Open claude in the kit folder and paste (R1 recipe prompt):

Use runtime/pontes/codex-exec.sh to ask Codex to review the README.md file: what might confuse a beginner? Then compare its feedback with your opinion.
How to verify: Claude asks to run the bridge, shows what Codex replied, and then separates where the two agree and disagree.

✓ Review by another model

  • ✓ Different blind spots cancel each other out
  • ✓ The reviewer can be a lower-tier model
  • ✓ Disagreement becomes a question for you
  • ✓ Already included in the kit’s bridge

✗ The same model reviewing itself

  • ✗ Tends to agree with what it just did
  • ✗ Repeats the same reading error
  • ✗ Gives a false sense of being “verified”
  • ✗ Uses up quota without adding a fresh perspective
👀
Another perspective

rule 3

🤝
R1 recipe

Claude does it, Codex reviews

📖
Read-only

bridge pattern

⚖️
You decide

in disagreements

5

Build a team only when it’s worth it

Rule 2: "Use a team of agents only when the work is worth more than the cost of assembling the team." A team (module 3.2) calls several models, and each one reads the task and the project from scratch.

The kit's own example shows the cost: the team created a one-sentence file and spent the equivalent of ~US$ 0,73 in API costs. It was great as a test. For writing one sentence in day-to-day work, a single agent would be enough.

TaskIs one agent enough?Is a team worth it?
Schedule Clara’s free timeYes, lower levelNo
Answer a question about the POLICYYesNo
Sônia’s monthly ERP summary, with verified totalsRisky: no one checksYes: plan, do, and check
Build a client system bridgeRiskyYes

What to look for in the table: A team makes sense when mistakes are costly and someone needs to review the work. If the task fits in one sentence and mistakes are immediately visible, one agent is enough.

⚠️ An agent inside an agent costs more than it seems

Rule 4, copied from the kit: "An agent call from within another agent carries the project's instructions and tools. For simple bridges, use a lightweight folder or --bare from Claude." In other words, each internal call rereads the folder’s rules. In a crowded folder, that multiplies the cost.

💡 Quick question

Before assembling a team, ask: "if the agent gets this wrong, will I notice right away?" If you will, use a single agent. If you won’t, the team’s reviewer is worth the quota.

👤
One agent

short task

👥
Team

work greater than the cost

📦
Lean folder

less to reread

💸
~US$ 0,73

R2 equivalent

6

Keep the table up to date

The last line of the ROTEAMENTO.md warns: "Model names change. Check with claude --help e codex --help and update this table."

The levels stay. What changes is the name inside each cell. That's why the recipes say "lower level" rather than a fixed name.

🎯 Objective: see which models your tools support today

In the terminal, one command at a time:

claude --help
codex --help

Or let the agent compare. Open claude in the kit folder and paste:

Rode claude --help e codex --help, compare com a tabela de runtime/ROTEAMENTO.md e me diga quais nomes de modelo mudaram. Não edite o arquivo: proponha a mudança numa linha da tabela Aprendizado da runtime/POLITICA.md.
How to verify: in the help, look for the model option (the one that chooses the model for the call). If the agent proposed something, the row appears in the Learning table with the status "proposed", and the ROTEAMENTO.md stays the same until you approve.
1 · check--help for both CLIs 2 · proposerow in Learning 3 · you approveyes or no 4 · updateROTEAMENTO.md the four levels don't change; only the name inside each cell

How to read the diagram: the glowing purple box is you. The agent checks and proposes, but doesn't change the rules file on its own. It's the "propose → approve → incorporate" of the POLITICA.md, which you’ll examine in depth in module 4.2.

Quick test (optional): Sônia only wants to format the total by customer that she has already checked as a table. Which level should she use?

❓
--help

source of the current name

🏷️
Fixed level

name changes

📝
Learning

agent proposes

✅
You approve

and only then changes

🎓 Module summary

✓
Quota isn’t dollars — it's the subscription limit, and each call uses some of it.
✓
Four levels — super, top, executor, and lowest, with the name on each tool.
✓
Start with the smallest — and push with CODEX_MODELO only if it makes a mistake.
✓
Review by another model — Claude does it, Codex reviews.
✓
Team only when it’s worth it — and the table is checked against --help.

Next module:

3.2 — Three-role team