Detailed content
📦 What is a skill in Claude Code
One skill is a package of instructions and tools that changes Claude Code's behavior for a specific task. Instead of explaining the same process every day ("before coding, write a test"), you package that logic into a skill, and the agent loads it when needed.
Technically, a skill is a directory with a file SKILL.md (YAML frontmatter + Markdown body) that describes when to activate e how to act. It can include slash commands, sub-agents, hooks, and references to other skills.
⚙️ Three ways to invoke a skill
-
•
Explicit slash command: user types
/grill-meor/tdd— deliberate invocation. -
•
Automatic trigger: the agent reads the skill description and activates it when the request matches (e.g., "let’s refactor this" ->
/improve-codebase-architecture). -
•
Composition: one skill calls another.
/make-plancan call/grill-meinternally.
# Minimum SKILL.md structure --- name: grill-me description: Stress-tests a plan before coding. Use when the user says "I'm going to start implementing", "I have an idea", "does it make sense to do it this way?". --- # Grill Me You're a skeptical senior. Before accepting any plan: 1. List 5 assumptions implicit in the user's request. 2. For each one, ask "what if this is wrong?". 3. Suggest the cheapest experiment to validate each one. 4. Only then return the refined plan.
💡 Instruction priority
When there's a conflict, the order of precedence is: system prompt > active skill > CLAUDE.md > user request. A well-crafted skill “anchors” behavior even if the user tries a shortcut.
Example: if /tdd if this is active and the user says "skip the test this time," the agent refuses—the skill wins.
🎯 Problem #1 — Misalignment
The most costly problem isn’t incorrect code—it’s code right about the wrong thing. The user asks for X but wants Y, the agent implements X perfectly, and two weeks later the feature gets thrown away.
📚 Citation—David Thomas (The Pragmatic Programmer)
"No-one, not even users, knows exactly what they want. Even when they do, they cannot articulate it. And even when they articulate it, they will change their minds tomorrow."
Conclusion: accepting the user's literal request is almost always a trap. Skills like /grill-me exist to force a discovery conversation before of the code.
✗ BEFORE — Without grilling
User:
Agent:
- ✗ Code a CSV button in 2 hours.
- ✗ User tests: "but I wanted Excel with formatting."
- ✗ Rework. The CSV goes in the trash.
✓ AFTER — With /grill-me
The agent asks:
- ✓ "Who will open this file? Excel, Sheets, a BI tool?"
- ✓ "Do you need formatting (colors, totals), or just raw data?"
- ✓ "How many rows? CSV files freeze above 1M in Excel."
User responds -> agent codes an XLSX with formulas. Zero rework.
💡 Warning sign
If the agent has never corrected you, asked for clarification, or disagreed with you, it's serving you, not helping you. A good discovery skill should create productive friction in the first few exchanges.
🗣️ Problem #2 — Verbosity
When agent and human don't share vocabulary, every conversation has to be explained from scratch. Eric Evans (author of Domain-Driven Design) calls this the absence of ubiquitous language: terms the whole team—humans and agents—uses consistently.
📖 Eric Evans — Domain-Driven Design
Evans’s core idea: the code should speak the language of the business. If you have to translate "PaymentBatch" to "batch of charges" every time you open a file, the model is wrong.
With agents, the loss is amplified: each new conversation starts from scratch because the agent didn't "live through" the previous discussions. Skills with a fixed glossary (e.g., /grill-with-docs) inject the ubiquitous language into every prompt.
A concrete example. Imagine a course where “lesson” can be a placeholder OR a real instance materialized with a student and progress. Without ubiquitous language, you write it like this:
✗ Verbose (without ubiquitous language)
62 words. Quotation marks around "real" three times. Rebuilds the concept from scratch.
✓ Concise (with ubiquitous language)
21 words. The terms "materialization cascade" and "materialized" carry all the context.
💡 The 3 benefits of a ubiquitous language with agents
- 1. Brevity: one word carries a paragraph of context.
- 2. Accuracy: "materialize" means exactly one thing, with no ambiguity.
- 3. Audit: commits, PRs, and tickets share the same dictionary.
🧪 Problem #3 — Code that doesn't work
LLMs are trained to produce code that seems right. Types match, imports exist, syntax is valid — and yet the logic is wrong. Without an external signal (tests), the agent believes its own output.
The classic solution is TDD (Test-Driven Development): write the test first, watch it fail (red), implement the minimum to make it pass (green), then refactor. Skills like /tdd imposes this cycle on the agent.
RED — Write the failing test
Before writing a line of implementation, describe the expected behavior in a test.
Why it works: it forces you to define what before the how. If you can’t write the test, you don’t understand the requirement. A failing test confirms that you’re testing something real—not an empty test that always passes.
GREEN — Implement the minimum
The rule: the shortest, simplest code that makes the test pass. It can be ugly. It can be duplicated. It works.
Why it works: it blocks the temptation to "go ahead and get it ready for the next case." Each test buys one line of code, nothing more.
REFACTOR — Clean up with the safety net in place
With the tests green, you can now reorganize names, extract functions, remove duplication—without worrying about breaking anything.
Why it works: the refactoring happens afterward that behavior is locked in by tests. If a refactor breaks the logic, the test alerts you in seconds.
⚠️ Attention — the TDD anti-pattern with agents
The most common mistake: asking the agent to “implement X and then write the tests.” The agent will write tests that are in its code, not tests that would validate the requirement. A test afterward isn’t TDD—it’s theater.
🏚️ Problem #4 — Ball of mud
Code rots. Foote and Yoder documented this in 1997 with the term "Big Ball of Mud": systems that grow without coherent architecture because each feature was rushed in, without refactoring. With agents, the pace increases—and so does the decay.
The wrong intuition: "I’ll ask the agent to refactor everything now." The right intuition: refactor as an opportunity, not an event. You don’t rewrite the system; you identify a duplication and unify it when you already need to work in that area.
✗ Destructive refactor
- ✗ "Refactor the entire payment module" — 2 weeks, 50 files, impossible-to-review PR.
- ✗ Change a class name and break 14 places no one noticed.
- ✗ Feature flag stays on for 3 months. New and old code coexist.
- ✗ Result: a ball of mud with another layer on top.
✓ Deepening opportunities
-
✓
Before touching a file, run
/improve-codebase-architecturein it. - ✓ Skill identifies specific duplication: “this function exists in 3 places as variations.”
- ✓ Unifies only that case. An 80-line PR, reviewable in 10 minutes.
- ✓ Result: every feature makes the code a little better (the scout rule).
💡 Practical example — identifying duplication
Are you going to work on the UserController. First, run /improve-codebase-architecture UserController. The skill answers:
- UserController.create (line 42)
- AdminController.invite (line 78)
- WebhookController.signup (line 31)
They differ only in how they handle a blocked domain. I suggest extracting
EmailValidator.validate(email, options) and refactor the 3 callers in a PR before adding your new code."
🔗 How skills compose
Individual skills already help. But the real gains appear when they compose into workflows — each one solves a problem, and one’s output feeds the next.
Typical workflow for a feature, from scratch to production code:
/grill-with-docs Stress-testa a ideia contra docs e ADRs │ (elimina desalinhamento, problema #1) ▼ /to-prd Vira documento de produto curto e cravado │ (fixa linguagem ubiqua, problema #2) ▼ /to-issues Quebra o PRD em issues pequenas e testaveis │ (escopo pequeno = menos ball of mud) ▼ /tdd Para cada issue: teste vermelho -> verde -> refactor │ (resolve problema #3) ▼ /triage Revisa diff, sugere refator pontual antes do merge (combate problema #4, deepening opportunity)
Each arrow is a deliberate transition. You don’t have to hold the entire context in your head—the skills handle the handoff. The PRD becomes input for the /to-issues; the issue becomes input for the /tdd; the diff becomes input for the /triage.
🧱 Principle: skills are building blocks, not monoliths
A skill that tries to do “the whole feature cycle” is worse than five small chained skills. Reasons:
-
•
Reuse:
/grill-meworks for any plan, not just features. -
•
Replacement: replace
/tddby/bddin one project doesn’t break the rest. - • Maintenance: a short skill fits in your head and is easier to audit.
-
•
Free composition: same skill across different flows (e.g.,
/grill-mebefore the PRD, before the ADR, before migration).
💡 Practical rule
If you're writing a skill and the description needs more than 2 sentences to explain when to activate it, it's doing too much. Break it up.
📌 Module Summary
Next Module:
1.2 — Anatomy of a SKILL.md: frontmatter, description triggers, and body