PTENES
Claude Code skill · triage

Model and effort don’t convert

Before dispatching a subagent or workflow, the orchestrator-router decides separately which model (haiku/sonnet/opus/fable) and how much effort (low→max) the task calls for — instead of collapsing everything into "sonnet+medium".

The skill guides, the orchestrator decides, the agent executes — the skill doesn't act on its own; the agent executes
What it is

Two axes, two questions, never a compromise

maestro-roteador is a Claude Code skill: it takes a raw problem and decides, before dispatching, which model and effort each part of the work deserves.

🎯 Model = repertoire

Answers just one question: what level of rigor does the worst step Does the task require it? Choose the smallest model that handles this step well — never "just in case."

⚙️ Effort = deliberation

Add up ambiguity + cost of error. One clear solution and a cheap error → low. Several plausible paths or a costly error (production, lost data) → high or higher.

⚖️ Errors aren't negotiable; efficiency is

When the risk is correctness, there’s no cheap option. When it’s just polish/taste, the skill offers the cost × quality choice, and the user decides.

📊 Reference: effort in practice →

Third-party data on Opus 5.5 and GPT-6 Astra at every effort level. For both, the maximum level was not preferred, which reinforces the ladder: start low and increase only with evidence. The page also includes Nei's opinion: medium by default, high when more reasoning is needed, xhigh only in extreme cases. It is a reference and does not change the skill's procedure. (EN · ES)

How it works

Procedure, always in this order

The skill runs inline in the main turn — never in a subagent — and follows four fixed steps before dispatching any work.

1. Isolate the riskiest step→ 2. Model = the smallest that gets the job done→ 3. Effort = ambiguity + cost of error→ 4. Volume: test 1 item before the batch

🪤 Anti-overhead

A task cheaper than triage itself doesn't get triaged. "Fix this typo" goes straight to the turn's default.

🚨 Named risk = high

If you named a production/data-loss risk yourself in the justification, the minimum effort is high. “Medium with care” doesn’t exist.

🪜 An evidence-based ladder

Is the error cheap and reversible? Start at the lowest plausible effort and increase only with evidence of insufficient results — never "high just in case." The ladder does NOT apply to irreversible risks.

🎭 A matter of taste

When the structure is already guaranteed by a template/skill and only finishing touches remain, the triage presents the cost×quality tradeoff and lets the user choose.

Router flow: your request comes in, the routing agent analyzes it, chooses the right agent (research, code, content, analysis, review), and delivers the result
Savings

Cache: the hidden cost of switching models

The prompt cache is model-specific and depends on the context prefix. Losing the cache for a 100k-token context makes the next input ~12,5× more expensive (1,25× for writing vs. 0,1× for reading). The triage matrix needs to account for this.

🧱 Switch only at a boundary

Switching the /model of the main turn discards the accumulated cache. Switch only at a work boundary (end of phase, handoff) — never mid-block. With a large context, switching may cost more than the savings from the smaller model.

🤖 Subagents are free in cache

Each Agent/Workflow has its own context: routing a part to a smaller model via a subagent does NOT touch the main cache. This is the preferred way to apply the matrix without switching costs.

🚫 No artificial keepalive

Never send “still there?” just to keep the cache alive — Claude Code manages the cache on its own, and the ping costs more than it saves. Long pause? Make a lean handoff and let it expire.

📋 A cache-preserving routine

Work in continuous blocks (plan → execute → test → document), keep the start of the context stable (don't turn tools and MCPs on or off mid-block), and group related tasks in the same session.

⏸️ Pauses

Short pause: continue as usual. Long pause or task switch: brief handoff (objective, status, decisions, next step) + fresh context. With the direct API, calls spaced 10–50 min apart pay the 1h TTL cache rate (2× for writing).

Perspective

Harness > effort: where the result really comes from

In a real test with the same task at 12 effort levels and 2 providers, the functional results were nearly identical — the high levels added a favicon and shadows, for 2–5× the tokens. What changes the result isn’t the effort slider.

📝 A clear spec beats high effort

A prompt that says exactly what "done" means delivers at low what max effort tries to guess. Effort can't make up for a vague specification.

🛠️ The model is a brain in a jar

Tools, files, terminal, skills, and instructions — the harness — are the model's arms. Investing in the harness pays off more than increasing effort.

🧠 Too much effort gets in the way

Overthinking is too much of the effort axis itself: excessive deliberation on a simple task re-explores paths already decided and can worsen the result, not just make it more expensive. The right effort is the lowest level that covers the risk — beyond that, you’re buying noise, not safety.

The right agent for the right task, every time: the router analyzes type, complexity, risk, tools, and cost before dispatching specialized agents
Prerequisites

What needs to be live

That's all: Claude Code installed and the skill available at ~/.claude/skills/ (user) or .claude/skills/ (project).

Claude Code

CLI with support for skills and Agent/Workflow calls with parameters model/effort.

# confirms installation
claude --version

Git

To clone the repo and create the installation symlink.

# confirms git
git --version

No extra dependencies

No external service, no API key, no build. It’s pure Markdown read by Claude Code.

# skill content
skills/maestro-roteador/SKILL.md
User guide · step by step

Install and use the triage

Install via symlink (to receive repo updates automatically) and three ways to trigger the triage.

1

Clone the repo

Download the skill.

git clone https://github.com/inematds/maestro-roteador.git
2

Install via symlink (user)

It’s available in any project and receives updates from the repo without reinstalling.

ln -sfn "$(pwd)/maestro-roteador/skills/maestro-roteador" ~/.claude/skills/maestro-roteador  # user symlink
3

Or copy it to just one project

Without a symlink, directly in the project’s skills folder (won’t receive automatic updates).

cp -r maestro-roteador/skills/maestro-roteador .claude/skills/  # project-only skill
4

Ask for triage directly

Trigger phrases: "which model", "how much effort", "triage this".

"Faz a triagem disso: renomear 80 arquivos em lote seguindo um padrão."  # -> haiku + low
5

Or just ask for the work as usual

When dispatching subagents/workflows, the skill applies the matrix on its own and assigns each part with model e effort correctly — without you needing to request triage separately.

"Migra esses dados de produção pro schema novo."  # -> sonnet + high (named risk)
6

Read the dispatch plan

The output always uses the same YAML format, with the riskiest step and the risk stated for each part.

tarefa: <resumo>
partes:
  - o_que: <subtarefa>
    pior_passo: <qual e por quê>
    modelo: haiku|sonnet|opus|fable
    esforco: low|medium|high|xhigh|max
    risco: <custo do erro em 1 linha>
turno_principal: <recomendação de /model, se valer trocar>
Validation

Skill TDD: baseline vs. with the skill

The skill was written against a measured baseline: without it, agents collapse every task into "sonnet + medium". With it, all three test scenarios landed in the right part of the matrix.

ScenarioWithout the skillWith skill
Rename 80 files in bulksonnet + mediumhaiku + low
Script with brand voicesonnet + mediumfable + low
Data migration with production at risksonnet + mediumsonnet + high

💡 Why this matters

The generic “sonnet+medium for everything” middle ground overpays for mechanical tasks and underpays for tasks with real risk — the three scenarios show both mistakes at once.

🧪 How to test it yourself

Run “without the skill” (disable it) and “with the skill” on the same raw request and compare the chosen model/effort — this is exactly the test that validated the three scenarios above.

Roadmap

Current state and known limitation

INEMA research/education project — no formal feature roadmap; what exists today and the skill’s documented limit.

Today
Functional skill, validated against the baselineDecision rules, quick-reference table, common pitfalls, and output format defined in SKILL.md, tested in the three scenarios above.
Automatic
Subagents and workflowsFor Agent/Workflow calls, the model/effort decision is automatic — the skill applies the matrix without needing to be called explicitly.
Limit
Main conversation only recommendsFor the main turn’s model, the skill only makes a suggestion — the user always makes the actual switch via /model.