Before dispatching a subagent or workflow, the orchestrator-router decides separately which model (haiku/sonnet/opus/fable) and how much effort (low→max) the task calls for — instead of collapsing everything into "sonnet+medium".
maestro-roteador is a Claude Code skill: it takes a raw problem and decides, before dispatching, which model and effort each part of the work deserves.
Answers just one question: what level of rigor does the worst step Does the task require it? Choose the smallest model that handles this step well — never "just in case."
Add up ambiguity + cost of error. One clear solution and a cheap error → low. Several plausible paths or a costly error (production, lost data) → high or higher.
When the risk is correctness, there’s no cheap option. When it’s just polish/taste, the skill offers the cost × quality choice, and the user decides.
Third-party data on Opus 5.5 and GPT-6 Astra at every effort level. For both, the maximum level was not preferred, which reinforces the ladder: start low and increase only with evidence. The page also includes Nei's opinion: medium by default, high when more reasoning is needed, xhigh only in extreme cases. It is a reference and does not change the skill's procedure. (EN · ES)
The skill runs inline in the main turn — never in a subagent — and follows four fixed steps before dispatching any work.
A task cheaper than triage itself doesn't get triaged. "Fix this typo" goes straight to the turn's default.
If you named a production/data-loss risk yourself in the justification, the minimum effort is high. “Medium with care” doesn’t exist.
Is the error cheap and reversible? Start at the lowest plausible effort and increase only with evidence of insufficient results — never "high just in case." The ladder does NOT apply to irreversible risks.
When the structure is already guaranteed by a template/skill and only finishing touches remain, the triage presents the cost×quality tradeoff and lets the user choose.
The prompt cache is model-specific and depends on the context prefix. Losing the cache for a 100k-token context makes the next input ~12,5× more expensive (1,25× for writing vs. 0,1× for reading). The triage matrix needs to account for this.
Switching the /model of the main turn discards the accumulated cache. Switch only at a work boundary (end of phase, handoff) — never mid-block. With a large context, switching may cost more than the savings from the smaller model.
Each Agent/Workflow has its own context: routing a part to a smaller model via a subagent does NOT touch the main cache. This is the preferred way to apply the matrix without switching costs.
Never send “still there?” just to keep the cache alive — Claude Code manages the cache on its own, and the ping costs more than it saves. Long pause? Make a lean handoff and let it expire.
Work in continuous blocks (plan → execute → test → document), keep the start of the context stable (don't turn tools and MCPs on or off mid-block), and group related tasks in the same session.
Short pause: continue as usual. Long pause or task switch: brief handoff (objective, status, decisions, next step) + fresh context. With the direct API, calls spaced 10–50 min apart pay the 1h TTL cache rate (2× for writing).
In a real test with the same task at 12 effort levels and 2 providers, the functional results were nearly identical — the high levels added a favicon and shadows, for 2–5× the tokens. What changes the result isn’t the effort slider.
A prompt that says exactly what "done" means delivers at low what max effort tries to guess. Effort can't make up for a vague specification.
Tools, files, terminal, skills, and instructions — the harness — are the model's arms. Investing in the harness pays off more than increasing effort.
Overthinking is too much of the effort axis itself: excessive deliberation on a simple task re-explores paths already decided and can worsen the result, not just make it more expensive. The right effort is the lowest level that covers the risk — beyond that, you’re buying noise, not safety.
That's all: Claude Code installed and the skill available at ~/.claude/skills/ (user) or .claude/skills/ (project).
CLI with support for skills and Agent/Workflow calls with parameters model/effort.
# confirms installation
claude --versionTo clone the repo and create the installation symlink.
# confirms git
git --versionNo external service, no API key, no build. It’s pure Markdown read by Claude Code.
# skill content
skills/maestro-roteador/SKILL.mdInstall via symlink (to receive repo updates automatically) and three ways to trigger the triage.
Download the skill.
git clone https://github.com/inematds/maestro-roteador.git
It’s available in any project and receives updates from the repo without reinstalling.
ln -sfn "$(pwd)/maestro-roteador/skills/maestro-roteador" ~/.claude/skills/maestro-roteador # user symlink
Without a symlink, directly in the project’s skills folder (won’t receive automatic updates).
cp -r maestro-roteador/skills/maestro-roteador .claude/skills/ # project-only skill
Trigger phrases: "which model", "how much effort", "triage this".
"Faz a triagem disso: renomear 80 arquivos em lote seguindo um padrão." # -> haiku + low
When dispatching subagents/workflows, the skill applies the matrix on its own and assigns each part with model e effort correctly — without you needing to request triage separately.
"Migra esses dados de produção pro schema novo." # -> sonnet + high (named risk)
The output always uses the same YAML format, with the riskiest step and the risk stated for each part.
tarefa: <resumo> partes: - o_que: <subtarefa> pior_passo: <qual e por quê> modelo: haiku|sonnet|opus|fable esforco: low|medium|high|xhigh|max risco: <custo do erro em 1 linha> turno_principal: <recomendação de /model, se valer trocar>
The skill was written against a measured baseline: without it, agents collapse every task into "sonnet + medium". With it, all three test scenarios landed in the right part of the matrix.
| Scenario | Without the skill | With skill |
|---|---|---|
| Rename 80 files in bulk | sonnet + medium | haiku + low |
| Script with brand voice | sonnet + medium | fable + low |
| Data migration with production at risk | sonnet + medium | sonnet + high |
The generic “sonnet+medium for everything” middle ground overpays for mechanical tasks and underpays for tasks with real risk — the three scenarios show both mistakes at once.
Run “without the skill” (disable it) and “with the skill” on the same raw request and compare the chosen model/effort — this is exactly the test that validated the three scenarios above.
INEMA research/education project — no formal feature roadmap; what exists today and the skill’s documented limit.