Claude · GPT-6 · Grok — September 2026

The right model for each task

Many models shipped at the same time. The practical answer is simple: use the best model for each task, with no loyalty to any platform.

Banner: which AI model to use for each task
Quick view

Each model in one line

Click a model for the full, plain-language explanation.

⚖️ Main rule: don't trust benchmarks alone

Test the models on your own tasks.

Beware of the honeymoon effect: new models look amazing in their first days. Wait a week and watch stability, quality and cost before changing your whole stack.

The stack today

Who does what

Used through subscriptions (Claude plan + Codex plan). “Cost” here means how much of the quota each model uses up.

Planning and hard problems
GPT-6 Astra
General use and complex projects
Opus 5.5
Cost-effective execution
GPT-6 Sol
Repetitive tasks
GPT-6 Luna
1

Planning and hard problems

GPT-6 Astra · medium–high effort. Architecture, scenario planning, demanding cases.

2

General use and complex projects

Claude Opus 5.5 · low effort. Fast, direct, economical, needs few instructions.

3

Creative, design, video, browser

Claude Opus 5.5 · low–medium. Gets it right the first time where taste matters.

4

Executing a ready plan

GPT-6 Sol · medium–high. Good value when the task is already well defined.

5

Reviewing and fixing code

GPT-6 Sol · medium–high. Strong when there is a clear criterion (tests that pass or fail).

6

Grunt work

GPT-6 Luna · Low for simple and fast; High / X-High for heavy work. Unambiguous tasks only.

7

Many parallel subtasks

Opus 5.5 orchestrates, Sol or Luna execute. Judgment once, cheap execution many times.

8

Ambiguous case

Astra proposes, Opus 5.5 reviews. Two cheap views beat one expensive one.

✕

Out of the stack

Grok 4.7 (worse than 4.6) · Fable 5.1 (expensive, surpassed by Opus 5.5) · Sonnet 5 and Haiku 4.5 (weak value).

How to decide

From request to model in four questions

Look at the hardest step of the task, not the task as a whole.

What is the hardest step?→Smallest model that handles it→Lowest effort that covers the risk→Check the result→Log it

🚫 No loyalty

Go wherever you are best served — team Claude or team Codex, it doesn't matter.

📈 Don't pay for the top

Above “high” effort there was no real gain. Raise it only with evidence.

🍯 Honeymoon

Every new model looks great in its first days. Wait a week before changing your stack.

🎯 Quota per result

A model that uses more quota but nails it the first time can beat four attempts on the cheaper one.

📁 One agent, one folder

Two agents in the same repository trip over each other. Isolate each one and tell it where it may write.

🧪 Test it yourself

Three real tasks, two models, same prompt. Thirty minutes beat any benchmark.

Requirements

What you need

Nothing to install beyond the tools you already use.

Claude subscription

Claude Code or Claude Desktop, for Opus 5.5.

# check quota usage
/usage

Codex subscription

Codex CLI or Codex Desktop, for GPT-6 Astra, Sol and Luna.

codex  # opens a session

The repository

Model cards, rules and prompts in Markdown (written in Portuguese).

git clone https://github.com/inematds/modelos
How to use · step by step

Day-to-day use

From a quick lookup to turning the catalog into a skill.

1

Check the stack

The README table is the quick reference. When in doubt, open the model's card.

cat README.md               # stack per task + rules
cat modelos/opus-5-5.md     # card: best for, avoid for, effort
2

Use a ready-made pattern

For flows with more than one model, copy the matching prompt.

prompts/planejar-executar.md     # Astra plans → Sol executes
prompts/segunda-opiniao.md       # Astra proposes → Opus 5.5 reviews
prompts/orquestrador-workers.md  # Opus 5.5 splits work → Sol/Luna do it
prompts/tarefa-bracal.md         # Luna in batch, no ambiguity
3

Log what you observe

One line per observation, newest on top. This is what makes the honeymoon rule practical.

# log.md
| date | model | task | result | source |
4

Run the quick test

Three real tasks (easy, medium, hard), same prompt on the current model and the candidate. Repeat the hard one after a week.

regras.md             # 30-minute protocol
avaliacao/bateria.md  # 8 real cases, 0–3 score, quota used
5

Turn it into a skill once the stack settles

There is a ready draft that answers “which model should I use for this?”.

cp -r skills/escolher-modelo ~/.claude/skills/
Usage patterns

Three combinations that work

The expensive part (judgment) happens once; the long part (execution) runs cheap.

🗺️ Plan → execute

Astra writes a plan with verifiable steps; Sol executes without changing it and stops if anything goes off-script.

⚖️ Second opinion

Astra proposes; Opus 5.5 reviews as a skeptic. If they agree, go ahead. If not, the disagreement is what you decide.

🎛️ Orchestrator + workers

Opus 5.5 splits the work and reviews it; Sol or Luna handle each part in its own folder.

📊 Effort in practice →

Opus 5.5 and GPT-6 Astra data across every effort level, with charts. Nei's take: medium by default, high when you need more reasoning, xhigh only in extreme cases.

Next steps

What comes next

The stack is a starting point and will be revisited with real use.

Now
Catalog publishedModels explained, stack, model cards, rules, prompts and test battery.
Sep 30
Post-honeymoon reviewConfirm or adjust the stack after a week of use.
Later
Model-picking skillInstall the escolher-modelo skill or fold the stack into the existing model router.