🌱 From intent to first draft
Capture what the person wants, do enough research, and turn it into a well-written SKILL.md from the start — imperative, with the why, and without shouting MUSTs.
Before writing a single line, skill-creator asks 4 questions: what the skill enables Claude to do, when it should trigger, what output format is expected, and whether it's worth putting together test cases.
Most bad skills come from poorly defined intent. If the conversation already contains the workflow ("turn this into a skill"), extract it from the messages first and only ask for what’s missing.
what it enables · when it triggers · output format · need tests? · confirm before proceeding
Proactively ask about edge cases, input and output formats, sample files, success criteria, and dependencies. Research in parallel with subagents when useful MCPs are available.
Arriving with context ready reduces friction for the user. Write test prompts only after completing this part — a poor interview creates a skill that covers the happy path and breaks everywhere else.
edge cases · input/output formats · example files · success criteria · dependencies · parallel research
Based on the interview, fill in name (identifier), description (trigger: what it does AND when to use it), and the Markdown body with the instructions. Every "when to use" detail belongs in the description, not the body.
The description is the primary trigger mechanism. Since Claude tends to under-trigger, it should be a little "pushy" — list concrete contexts where the skill should kick in even when the user doesn't explicitly ask for it.
name · pushy description · what + when · body <500 lines · compatibility (rare)
Write in the imperative, use theory of mind, explain why each instruction exists instead of piling on uppercase MUSTs, and keep the skill general instead of tailoring it to the examples.
Today’s LLMs are smart: when you explain why, they go beyond rote instructions and solve the real case. ALL CAPS ALWAYS/NEVER is a yellow flag—rephrase it to explain the reason.
imperative · theory of mind · explain why · avoid MUSTs · general skill, not narrow
Write a first draft without getting stuck, then reread it with fresh eyes and improve it. Draft → review → improve is a cycle within the writing itself.
The first draft is almost never the best. A fresh reread reveals redundant or ambiguous instructions, or ones that tell the model to waste time.
quick draft · reread with fresh eyes · cut redundancy · clarity · iterate on the writing
Decide when to move resources out of SKILL.md: if 3 runs repeat the same script, turn it into scripts/; large docs go in references/ loaded on demand.
Progressive disclosure: metadata always in context, body when triggered, resources only when needed. Bundling a repeated script saves every future invocation from reinventing the wheel.
scripts/ · references/ · assets/ · progressive disclosure · rule of 3 repetitions · TOC in docs >300 lines
🔁 Test, evaluate, and iterate
The heart of the loop: test realistic prompts, run with-skill vs. baseline, evaluate qualitatively and quantitatively, generalize from feedback, and optimize the description using the test metric.
After the draft, create 2-3 realistic test prompts — the kind of thing a real user would type — and show them to the user for validation before running them.
Artificial test prompts produce misleading evaluations. The prompts are in evals/evals.json without assertions yet — the assertions come later, while the runs are running.
2-3 prompts · real user language · validate with the user · evals.json · no assertions yet
For each test case, launch two subagents in the same turn: one with the skill, one without (baseline). Launch them all at once so they finish together.
Without a baseline, you don’t know whether the skill added anything. For a new skill, the baseline = no skill at all. For an existing skill, the baseline = the old version (snapshot before editing).
with_skill · without_skill · same turn · snapshot of the old version · workspace per iteration
Evaluate on two fronts: qualitative (review the outputs in the viewer) and quantitative (verifiable assertions that produce a pass_rate). Subjective skills are evaluated only qualitatively.
Good assertions are objectively verifiable and have descriptive names. Don't force assertions on things that call for human judgment (writing style, design).
qualitative in the viewer · verifiable assertions · pass_rate · descriptive names · don’t overfit to subjective criteria
Generalize from feedback instead of gradually overfitting to a few examples, keep the prompt lean, and read the transcripts (not just the final outputs).
The skill will be used a million times in different prompts. If it only works for the test examples, it’s useless. Avoid fiddly changes and oppressive MUSTs.
generalize · don't overfit · keep it concise · read transcripts · explain why · repeated script becomes a bundle
Apply the improvement → rerun all test cases in a new iteration → review with the user → read the feedback → repeat until it meets the requirements.
The loop stops when the user is happy, all feedback is empty, or you’re no longer making meaningful progress. Each iteration goes in its own directory.
apply → rerun → review → repeat · iteration-N · previous-workspace · stopping criteria
Generate ~20 trigger queries (a mix of should-trigger and should-not-trigger), focus on near-misses, run the optimization loop, and choose the description based on the test set metric.
The description determines whether the skill triggers. Obvious negatives don't test anything — the valuable ones are near-misses that share keywords but need something else.
20 queries · should-trigger / should-not · near-misses · 60% train / 40% test · best_description based on the test set
⭐ The Best Meta-Skills (for creating skills)
The tools for people who CREATE skills: skill-creator (246k), find-skills (1,8M), and the scaffolding pattern. What each one does, when to use it, and where it fits in the workflow.
A skill whose job is to help you work with other skills—discovering, creating, testing, optimizing, and packaging them. It operates one level above an ordinary skill.
Many people create without checking whether something better already exists and without an eval loop. Meta-skills solve this: discover before creating, create methodically, validate with data.
meta-skill · discover · create · test · optimize · package
Anthropic’s meta-skill orchestrates the entire cycle—draft → eval → iterate—and includes a separate description optimizer.
It’s the central axis of the creation workflow. Everything in modules 4.1 and 4.2 comes from it; don’t improvise a parallel process.
draft → eval → iterate · scripts · aggregate_benchmark · run_loop · package_skill
The most installed skill in the catalog (1.802.925, vercel-labs). It finds relevant skills for a task before you create one from scratch.
Prevents the most costly mistake creators make: spending hours writing something that already exists in a better form. It's step zero in the workflow.
discovery · use / extend / create · step zero · catalog of 39.366 skills
A standard way to generate the initial scaffold—SKILL.md with frontmatter and scripts/, references/, and assets/ folders—instead of typing everything by hand.
It speeds up the start and removes the blank page. But create only the folders you’ll use—empty folders confuse the model and violate the "keep it lean" principle.
skeleton · pre-filled frontmatter · don’t create everything upfront · trim boilerplate
The sequence that prevents rework: find-skills (discover) → scaffolding (generate a base) → skill-creator (create and iterate) → description optimizer (refine the trigger).
The three don’t compete; they connect in sequence. Using them out of order is how skills with fewer than 100 installs come about.
discover → foundation → create/iterate → optimize → package · golden rule
A situation → tool cheat sheet: "I need a skill for X" → find-skills; "I’m going to create one" → skill-creator; "it doesn’t trigger when it should" → description optimizer.
Make the right decision at the right time. Install all three and, since Claude tends to under-trigger them, mention them explicitly the first few times until it becomes automatic.
situation → meta-skill · install all three · pro tip for triggering · package_skill
🛠️ How to Create: Complete Walkthrough with Evals
A complete worked example: intent (4 questions) → draft → test prompts → evals.json with verifiable assertions → run with-skill vs. baseline → iterate. Ready-to-use JSON templates.
The 4 questions applied to the “margem-xlsx” case: what does it enable, when does it trigger, what’s the output format, and is it worth testing? Verifiable output → worth testing.
Question 4 determines the rest of the walkthrough. File transformations and code generation are worth evaluating; style/art aren’t.
what it enables · when it triggers · format · worth testing · extract from conversation · confirm
Write the complete draft: name, pushy description (what it does AND when to use it), and an imperative body that explains why.
The description is the trigger. "Works with spreadsheets" under-triggers; listing verbs + contexts ("margin, profit... even without asking for a column") covers near-triggers.
name · pushy description · imperative body · explain why · don’t guess columns
2-3 prompts written as a real user would write them — with backstory, paths, and details — saved in evals/evals.json with no assertions yet. JSON template ready.
Artificial prompts produce misleading evaluations. Validate with the user before running—it’s cheap and avoids running everything for nothing.
2-3 prompts · natural language · paths and backstory · evals.json · validate first
While the runs are in progress, write objective assertions with descriptive names in eval_metadata.json—preferably checked by a script. Ready-to-use template.
"the spreadsheet looks good" is subjective; "values match (C-D)/C" is verifiable and distinguishes with_skill from baseline.
verifiable assertions · descriptive name · checked by a script · discriminates · not subjective
Two subagents in the same turn per test case (with skill / without skill = baseline), output organized by iteration and eval, with total_tokens and duration_ms in timing.json.
Launching them together avoids bias. Timing can only be captured when the notification arrives—process each one right away.
with_skill / without_skill · same turn · workspace per iteration · timing.json
Numbered timeline: grade each run → aggregate the benchmark → open the viewer before you judge → read the feedback and generalize → rerun in iteration-2.
In this case, all 3 runs wrote nearly identical calc_margem.py files — a sign they should be bundled. In iteration 2, the pass_rate went up and token usage went down.
grading.json · aggregate_benchmark · generate_review · feedback · repeated script bundle
🚀 Advanced Tips: Description Optimization and Benchmarking
The optimization loop (60/40 split, near-misses, best_description based on the test score), how triggering really works, blind comparison, and how to read the benchmark without fooling yourself.
run_loop.py splits the eval set into 60% train / 40% test, evaluates the description (3 runs per query), proposes improvements based on what failed, and reevaluates, up to 5x.
Activation is stochastic; 3 runs provide a stable trigger rate. Use the session’s model ID so the test matches what the user experiences.
run_loop · 60/40 · 3 runs/query · propose → reevaluate · session model ID · background
~20 realistic queries, half should-trigger and half should-not. The gold is in the near-misses — phrases that share words but need something else.
"Margin of error in a survey" uses "margin" but isn't spreadsheet profit — it forces the description to distinguish. Obvious negatives teach nothing.
should-trigger 8-10 · should-not 8-10 · near-misses · avoid obvious ones · trigger-eval.json
Skills appear in available_skills with a name + description, and Claude decides whether to consult them—but only for tasks it can’t easily solve on its own.
"Read this PDF" may not trigger even with a perfect description. Your test queries need to be substantive (multi-step, specialized).
available_skills · non-trivial tasks only · substantive queries · under-triggering · pushy
Give two outputs to an independent agent without saying which is which, let it judge their quality, and then analyze why the winner won.
Because it’s blind, it doesn’t favor “the new” just because it’s new. It’s optional and requires subagents—save it for when the question is costly.
anonymize · judge quality · why it won · optional · subagents · human review is enough
The benchmark provides pass_rate, tokens, and time by config, with mean ± standard deviation and delta. The analyst’s pass reveals what the average hides.
An assertion that passes with and without the skill measures nothing and inflates the pass_rate; high deviation may be flaky; a high pass_rate may hide a token explosion.
pass_rate · delta · non-discriminating · variance/flaky · token/time tradeoff
Take the best_description (chosen by the test score, not the train score), apply it to the frontmatter, show the before and after, and package it with package_skill.
Choosing based on the train set would reward a description that memorized the examples. The held-out test simulates the real world—it's what separates generalizing from shining only in the lab.
best_description · test score > train · optimize only after it's ready · review the eval set · package_skill
Learning path overview
Ask, research, write. A draft that starts with imperative instructions and explains why.
Run, measure, generalize. The loop that turns a draft into a skill that works a million times.
find-skills discovers, skill-creator creates, scaffolding generates the foundation. The tools for people who create skills.
From intent to evals.json with assertions. A worked example from start to finish, JSON ready to copy.
60/40 split, near misses, best_description based on the test. Fine-tune the trigger and read the benchmark without fooling yourself.