✅ When a skill IS worth it
A skill costs context and maintenance. It only pays off when it delivers real, repeated gains. There are four green flags—the more of them you see, the clearer the decision to package it. A practical rule: if two or more the signals line up; it's worth creating the skill.
🔁 Repeatable workflow
You run the same sequence of steps every week — generating a component, configuring a deploy, writing a type of test. Repetition is the number one signal.
Example: PR pattern, commit convention, feature scaffold.
🧠 The team's contextual knowledge
Something only your team knows: internal conventions, legacy gotchas, architecture decisions. The model can't guess this on its own.
Example: "in this repo, always use wrapper X instead of direct fetch."
🔍 Verifiable Output
You can check whether it worked—tests pass, lint is clean, schema validates. Verifiable output lets you iterate on and measure the skill.
Example: code that compiles, JSON that validates against a schema.
🪜 Multi-step
The task has several linked steps with order and dependencies. The skill encapsulates the entire process, not just one isolated tip.
Example: read config → generate files → run build → validate.
💡 Calibration tip
If you catch yourself explaining the same thing to the agent for the third time in a week, that’s a repeatable workflow + contextual knowledge. Two green signals: stop and write the skill now.
🚫 When It’s NOT Worth It
Knowing when to create is just as important as knowing when to no create. Most useless skills come from the impulse to "I'll organize this into a skill" for something that didn't need one. Three red flags indicate over-engineering.
✗ Do NOT create a skill when...
- ✗One-off task: you'll only do this once. A Skill is pure overhead.
- ✗The model already does this on its own: format markdown, translate, summarize — native capability, no added value.
- ✗Simple one-step query: "what's the capital of France?" doesn't become a skill.
- ✗You still don’t have the process clear, even in your own mind.
✓ Do something else
- ✓One-off → just include it in the prompt at hand.
- ✓Native capability → trust the model; don't duplicate it.
- ✓1 step → one sentence in the prompt solves it.
- ✓Diffuse process → do it manually 3x, then extract the pattern.
The "already does it on its own" test
Run the task before creating without skill (baseline). If the result is already good, the skill adds nothing — it only takes up context. A skill is worthwhile only if the with-skill vs. baseline delta is visible.
🌳 Skill vs. CLAUDE.md vs. MCP vs. subagent
Four tools solve different problems. Choosing the wrong one is the most common design mistake. The decision tree below leads you to the right choice with two or three questions.
CLAUDE.md — always active
Rules that apply in every session, with no trigger. Response style, language, general prohibitions. They weigh on context all the time—use them only for what’s universal.
MCP—connection
A server that gives the agent tools to communicate with external services (database, API, browser). MCP is plumbing; a skill is knowledge. Don’t confuse them.
Subagent — isolation
Heavy, parallel work with its own context (e.g., searching the entire repo). Keeps the main context clean. It can even use skills within it.
Skill—on demand
Knowledge that triggers only when the description matches. It doesn’t add weight when irrelevant. It’s the default choice for a repeatable workflow + context that doesn’t always fit in CLAUDE.md.
🐘 Anti-pattern: oversized scope
The first anti-pattern that kills skills: the temptation to create a skill "everything-about-X". An 800-line skill that covers all of React, from useState to deploy. It seems efficient. It’s the opposite.
✗ Huge scope
- ✗
tudo-sobre-react.mdwith 800 lines - ✗Description becomes generic: "helps with React"
- ✗Triggers on everything or nothing
- ✗Agent absorbs it poorly — instruction diluted
- ✗Impossible to test and version
✓ Atomic skills
- ✓
react-component-scaffold - ✓
react-hooks-conventions - ✓
react-a11y-guidelines - ✓Each one has a precise trigger and <500 lines
- ✓They compose together when the context calls for it
The division rule—one trigger per skill:
# ✗ ANTES (uma skill, vários gatilhos misturados) skills/ tudo-sobre-react/SKILL.md # 800 linhas, dispara em "react" # ✓ DEPOIS (uma responsabilidade cada) skills/ react-component-scaffold/SKILL.md # dispara em "novo componente" react-hooks-conventions/SKILL.md # dispara em "usar hook / estado" react-a11y-guidelines/SKILL.md # dispara em "acessibilidade"
💡 Splitting heuristic
If you can’t write a one-sentence description that says exactly when the skill triggers, it’s too big. Each distinct trigger = one skill. Atomicity isn’t about aesthetics: it’s what makes the trigger work.
📢 Anti-pattern: uppercase MUSTs and overfitting
The second anti-pattern: stuffing the skill with "YOU MUST ALWAYS" in ALL CAPS and paste overly specific examples. The model memorizes the exact cases and fails on any variation. The skill passes your tests and breaks in real life—that’s overfitting.
overfit vs. generalization:
# ✗ OVERFIT (decora o caso, não o princípio) VOCÊ DEVE SEMPRE nomear o arquivo de "Button.tsx". VOCÊ DEVE SEMPRE usar a cor #14b8a6. # ✓ GENERALIZA (ensina o porquê, vale fora dos exemplos) Nomeie componentes em PascalCase, igual ao nome exportado, porque o resolver de imports e a navegação do editor dependem dessa correspondência. Ex.: Button → Button.tsx, UserCard → UserCard.tsx.
✗ MUSTs and overfitting
- ✗ALL CAPS and shouted imperatives
- ✗Examples pasted in as fixed rules
- ✗No explanation of why
- ✗Breaks on any new case
✓ Principle + why
- ✓Calm, natural imperative
- ✓Examples illustrate; they don't constrain
- ✓Explain the reasoning behind
- ✓Generalizes to new cases
Why the why matters
The model is good at reasoning from principles. When you explain the reason (“because the import resolver depends on this”), it applies the rule correctly in situations you never anticipated. Uppercase MUSTs only signal anxiety — they don’t improve adherence and even encourage memorization instead of understanding.
🎯 Anti-pattern: vague description
The third anti-pattern — and the deadliest one. The description is the trigger: the only skill level that’s always in context. If it’s vague, the skill never triggers (under-triggering) or triggers in the wrong context. A brilliant skill with a poor description is invisible.
bad trigger vs. good trigger:
# ✗ VAGA (nunca dispara ou dispara errado) description: Ajuda com código. # ✓ ESPECÍFICA (diz O QUE faz E QUANDO usar) description: Gera componentes React com a convenção do time (PascalCase, hooks, props tipadas). Use quando o usuário pedir um novo componente, refatorar JSX, ou criar UI em React/Next. Acione também ao mencionar "componente", "tela" ou "página" num projeto React.
Say WHAT it does
The concrete capability, not the field. "Generates React components using the team's convention" — not "helps with frontend."
Say WHEN to use it
The explicit triggers: verbs and nouns that appear in the user’s request. “Use when asking for a new component, refactoring JSX...”
Be a little pushy
The model tends to under-trigger. Include “also trigger when mentioning X” for borderline cases. It’s better to trigger too often than to disappear when it should show up.
💡 The 5-second test
Read only the description (not the body). Can you tell exactly which request should trigger it? If you hesitated, the model will too. The description is the only part that determines whether the skill exists in practice.
✅ Module Summary
Next Module:
5.2 — 🚀 Publish, Version, and Measure — you decided it’s worthwhile and avoided the anti-patterns; now take the skill to the world: repo, git, skills.sh, trigger evals, and the lifecycle