📚 Where the rules come from
The figures in this lesson were verified in October 2026 against the official guide Skill authoring best practices and in the Claude Code docs on skills, hooks, and context window. The basic rules have been stable for about a year; what changed were the models. The adjustments for the 5.5 models at the end come from another source and are marked as such.
Detailed content
📏 Size and depth
O SKILL.md is a summary, not a manual. The official guide calls for a body
under 500 lines; the rest goes into linked reference files
directly from SKILL.md. A reference you can only reach through another reference is
nested: the agent may read only its preview (the first lines, something like a head -100)
and the rules at the end of the file disappear without warning.
lines is the SKILL.md body limit. If it exceeds that, split it by topic into references/.
deep: each reference is linked from SKILL.md itself, never only from another reference.
lines or more in a reference call for a summary at the top, so even a partial read shows everything the file covers.
# Pricing reference
## Contents
- Base prices per plan
- Regional taxes
- Discounts and coupons <- at the end, but listed here
- Refund rules
## Base prices per plan
...
💡 Practical tip
Open SKILL.md and count how many references it cites. Any reference that appears only inside another reference is a candidate to move up one level, with a line in SKILL.md saying “read X when Y.”
🎚️ Degrees of freedom
Not every step deserves the same level of control. The guide classifies them as degrees of freedom: the more fragile and costly the error, the less room the agent should have. A single skill often mixes all three. The test for each step is simple: “what if the agent does this step differently?”
Text instruction
Several answers work, and context determines the choice. E.g., brainstorming titles, reviewing code using good judgment.
Model with room to vary
There’s a preferred pattern, but it can be adapted. E.g., a weekly report with a template and parameters.
Exact script, no loose parameters
An error costs money, deletes data, or publishes something. E.g., invoice, tax filing, database migration: “run exactly this script.”
💡 Practical tip
The guide’s image: a narrow bridge with a chasm on either side calls for a handrail (low freedom); an open field needs only a direction (high freedom). Mark each step in your skill A, M, or H before rewriting it.
🏷️ Description: third person, “when to use” and limits
The description is included in the system prompt alongside the descriptions of all the other skills. That’s why the guide asks for third person (“Processes spreadsheets…”, never “I can…” or “You can…”), what the skill does and when to use it, with the words the person types. And there are size limits:
characters is the field maximum description in the frontmatter.
characters is where Claude Code cuts it off description + when_to_use combined in the skills list. Anything beyond that the model won’t see.
✓ Third person + when
"Generates invoices and sends payment reminders. Use when the user asks to bill a client, issue an invoice or chase a late payment."
✗ First person, no trigger
"I can help you with invoices."
💡 Practical tip
The same applies to the body: include only what the model doesn't know. Don't explain what an invoice is; keep your prices, terms, and internal rules. Context is shared with the conversation, other skills, and history.
✅ Checklist in the response and verification loop
A multi-step task needs a checklist the agent copies into its response and keeps checking, with a loop-back line: “if the total doesn't match, go back to step 2.” And every important output goes through a run the loop → fix → repeat until the check passes. The check doesn't have to be code: comparing the draft with the style guide and listing every deviation also counts.
Copy this checklist into your answer and tick each step:
- [ ] 1. Read the client data
- [ ] 2. Compute totals with scripts/total.py
- [ ] 3. Validate: python3 scripts/check_invoice.py out.json
- [ ] 4. Generate the PDF
If step 3 fails, fix the data and go back to step 2.
Only continue when validation passes.
⚠️ Attention
“Review carefully before delivering” isn’t a loop. A loop has a criterion that passes or fails and an instruction for what to do when it fails. For long tasks, a verifier with a clean context that didn’t do the work often finds more than self-criticism.
🧪 Test with every model
The same skill behaves differently with each model. The official guide asks you to test it with all the models that will use it and cite three questions, one for each family: Haiku, Sonnet, and Opus. The page doesn’t mention Fable.
Does the skill provide enough guidance? A smaller model needs more detail.
Is it clear and concise?
Does it avoid overexplaining? Redundant instructions are hard on large models.
💡 Practical tip
Test the skills you depend on, with three real requests each. For the others, the cost of testing isn't worth it.
📦 Packages, hooks, and compaction
Three practical rules to help the skill work beyond your machine and in a long session:
Packages with installation commands
Don’t assume the library is installed on your colleague’s machine: list the exact packages with the installation command next to the script that needs them. In the Claude API, the skill doesn’t install anything; only what’s already in the environment is available.
A rule that must never be broken becomes a hook
Uppercase is followed almost always; a hook is always followed. In Claude Code, a hook can be declared in the skill’s own frontmatter: after the skill is loaded, it runs before every command and stays active for the rest of the session.
Compaction keeps only the beginning
When a conversation is compacted, Claude Code keeps only the first 5,000 tokens of each loaded skill. A critical rule at the end of a long SKILL.md may disappear after compaction: put the most important rules at the top.
---
name: invoices
description: Generates invoices and sends payment reminders. Use when
the user asks to bill a client, issue an invoice or chase a late
payment. Not for accounting reports.
hooks:
PreToolUse:
- matcher: "Bash"
hooks:
- type: command
command: "./scripts/check-limit.sh"
---
# Invoices
Critical rules first: never send an invoice above the approval
limit without a human OK. (the hook enforces it)
🆕 Adjustments for models 5.5
According to the prompting guide for Claude Fable 5 and Opus 5.5 cited by skill-creator-plus (RoboNuggets, MIT), skills written for older models tend to be too prescriptive and may make the result worse. This isn't text from the skill best practices page; treat it as practical guidance and confirm it in your own use.
Step-by-step instructions for what the model already knows and repeated common sense tend to be dead weight. Only a run with and without the line can prove it.
An instruction to “write out your reasoning step by step” may be refused by 5.5 models. Ask for the answer, a short explanation, or a summary of the actions.
“Below 1,536 characters, because Claude Code cuts off there” lets the model handle the case the rule didn’t anticipate.
When everything shouts, nothing stands out. One guiding sentence is worth more than a list of prohibitions.
📋 Copyable checklist
Paste it into a conversation with the agent along with your skill, or use it yourself before publishing.
Check this skill against the rules below and mark each item: - [ ] The SKILL.md body is under 500 lines - [ ] Each reference is linked directly from SKILL.md, with no nested references - [ ] Every reference over 100 lines starts with a table of contents - [ ] Each step has the right degree of freedom: plain text, template, or exact script - [ ] The description is in third person and says what it does and when to use it - [ ] The description is no more than 1,024 characters and, together with when_to_use, fits within 1,536 - [ ] The body includes only what the model doesn’t know - [ ] Long tasks have a checklist the agent copies into its response, with a line to return to - [ ] There is a run, fix, and repeat loop with a pass-or-fail criterion - [ ] The skill has been tested with every model that will use it - [ ] Every package used has its installation command next to it - [ ] Any rule that must not be broken has been made a hook in the skill’s frontmatter - [ ] Critical rules are at the top, within the first 5,000 tokens - [ ] No instruction asks for reasoning to be written in the response For each item that fails, say where the problem is and propose a fix before editing.
🔎 Audit your skills with the validator
The mechanical parts of these rules can be measured. The auditar-skills (INEMA project, mirror of robonuggets/skill-creator-plus, MIT licensed) includes a pure Python validator, with nothing to install and no API calls, that reads the skills folder and lists errors and warnings by rule. Guide: inematds.github.io/auditar-skills/guia.
git clone https://github.com/inematds/auditar-skills cp -r auditar-skills/skill-creator-plus ~/.claude/skills/ python3 ~/.claude/skills/skill-creator-plus/scripts/validate_skill.py --all ~/.claude/skills
📊 Real example
On an INEMA production machine, with the folder ~/.claude/skills full of its own and third-party skills, the validator ran in about one second:
- ST5 · missing summary in a reference over 100 lines: the biggest offender, with 268 occurrences.
- DS3 · description without “when to use”: 75 skills.
- ST4 · nested reference: 60 occurrences.
💡 Practical tip
Measure before rewriting. Start with the skills you use most and the cheapest errors to fix: a summary at the top and “when to use” in the description solve most of the list.
✏️ Practical exercises
1. Run the validator
Run the command above in your skills folder and note the three rules that appear most often. Compare them with the real example in this lesson.
2. Mark the degrees of freedom
Choose one of your skills with more than five steps and mark each step high, medium, or low. Is any low-level step written as a loose instruction? Replace it with a script.
3. Run the checklist on a real skill ⭐
Paste the copyable checklist and your skill into a conversation with Claude. Fix any items that fail, run the validator again, and confirm that the number of errors has gone down.
🎯 Module summary
End of Track 1:
You already understand the anatomy, structure, trigger, and current rules. In Track 2, you'll dissect a real skill from start to finish and build your own.