🧪 Write 2-3 realistic test prompts
After the draft, create 2-3 realistic test prompts — the kind of thing a real user would type. Show the user before running: "Here are some test cases I want to try. Are they right, or do you want to add more?". Save the prompts in evals/evals.json still no assertions—they come later, while the runs are underway.
✗ Artificial prompt
"Format this data"
"Create a chart"
Too generic. Doesn't test anything or trigger skills.
✓ Realistic prompt
"okay, my boss sent me an xlsx (it's in Downloads, something like 'Q4 sales final FINAL v2.xlsx') and wants a profit margin column in %. Revenue is in column C, costs in D, I think"
Concrete, with backstory, paths, and real details.
evals/evals.json — prompts only, no assertions yet:
{
"skill_name": "example-skill",
"evals": [
{
"id": 1,
"prompt": "User's task prompt",
"expected_output": "Description of expected result",
"files": []
}
]
}
⚖️ Run with-skill vs. baseline in the same turn
For each test case, launch two subagents in the same turn: one with the skill, one without (baseline). This matters — don't run the with-skill ones first and come back for the baselines later. Launch everything at once so they finish around the same time.
With-skill run
Points to the skill’s path. Save it in iteration-N/eval-ID/with_skill/outputs/.
Baseline — new skill
Same prompt, without any skill. Save to without_skill/outputs/.
Baseline — existing skill
The old version. Before editing, create a snapshot (cp -r skill snapshot/) and point the baseline to it. Save it in old_skill/outputs/.
💡 Why the baseline?
Without comparing against the baseline, you don’t know whether the skill added anything. Maybe the model would have handled the case on its own. The baseline is what separates “the skill helped” from “this would have happened anyway.”
📊 Evaluate: qualitative + quantitative evals
The evaluation has two fronts. Qualitative: review outputs in the viewer, click through each case, and leave feedback. Quantitative: objectively verifiable assertions that produce a pass_rate. Subjective skills (style, design) are better evaluated qualitatively—don’t force assertions where human judgment is needed.
✗ Bad assertion
- ✗Subjective: "the output looks nice"
- ✗Opaque name: "check_1", "assert_x"
- ✗Always passes, with or without the skill (doesn’t discriminate)
- ✗Forced into a creative writing skill
✓ Good assertion
- ✓Verifiable: "margin column exists and is %"
- ✓Descriptive name, readable in the viewer
- ✓Checked by script, not by eye
- ✓Distinguishes with_skill from baseline
eval_metadata.json — with assertions added:
{
"eval_id": 0,
"eval_name": "margem-lucro-xlsx",
"prompt": "The user's task prompt",
"assertions": [
"Arquivo .xlsx de saída existe",
"Coluna 'margem' presente e formatada como %",
"Valores batem com (C - D) / C"
]
}
🎯 Show before judging
Generate the eval viewer with generate_review.py and put the outputs in front of the human before before you try to fix it yourself. The "Outputs" tab shows one case at a time; the "Benchmark" tab shows pass_rate, time, and tokens per configuration, with mean ± standard deviation and the delta.
🧠 How to think about improvement
This is the heart of the loop. The key insight: the skill will be used a million times with different prompts. You and the user iterate on a few examples because it's quick, but if the skill only works for those examples, it is useless. Generalize from feedback instead of overfitting.
Four ways to think
Generalize from feedback
Avoid fiddly overfit changes and oppressive MUSTs. If a problem is stubborn, try other metaphors or work patterns — it's cheap to test and could yield something great.
Keep it concise
Remove what isn’t pulling its weight. Read the transcripts, not just the final outputs: if the skill makes the model waste time, cut the part causing that and see what happens.
Explain why
Even if the user’s feedback is terse or frustrated, understand the task and convey that understanding. ALL CAPS ALWAYS/NEVER is a yellow flag — rephrase it with the reason.
Look for repeated work
If the 3 test cases wrote a build_chart.py similar, that’s a strong sign you should bundle this script. Write it once, put it in scripts/, save every future invocation.
💡 Thinking time isn't the bottleneck
The skill-creator’s advice is literal: write a draft review, look at it again with fresh eyes, and improve it. Get into the user’s mindset and understand what they truly want and need. It’s worth mulling over.
🔁 The iteration loop
With the improvement in mind, run the cycle: apply → rerun → review → repeat. Each iteration goes in its own directory (iteration-2/, iteration-3/...), including the baselines.
Apply the improvements
Edit the SKILL.md based on the feedback and what you generalized from it.
Rerun all test cases
In a new iteration-N+1/, including baselines. New skill → baseline always without_skill.
Review with the user
Launch the reviewer with --previous-workspace pointing to the previous iteration for comparison.
Read the feedback and repeat
Empty feedback = the user thought it was okay. Focus on cases with specific complaints.
When to stop
- •The user says they’re happy
- •All the feedback is blank (everything OK)
- •You’re not making meaningful progress
🎚️ Description optimization
The description is the primary mechanism that determines whether Claude invokes the skill. After creating or improving one, optimize it for better trigger accuracy. Generate ~20 trigger queries — a mix of should-trigger e should-not-trigger — and run the loop, choosing the description based on the test set metric.
✓ Should-trigger (8-10)
- ✓Same intent, different phrasing (formal/casual)
- ✓Cases where the user doesn't name the skill but needs it
- ✓Uncommon use cases and conflicts with another skill where this one should win
✗ Should-not-trigger (8-10)
- →Near-misses: they share keywords but need something else
- →Adjacent domains and ambiguous phrases
- →Avoid obvious negatives — “write a fibonacci” tests nothing
trigger-eval.json — labeled queries:
[
{"query": "the user prompt", "should_trigger": true},
{"query": "near-miss tricky prompt", "should_trigger": false}
]
The automatic loop
O run_loop.py splits the eval set into 60% train and 40% held-out test, evaluates the current description (running each query 3x for a reliable trigger rate), asks Claude to suggest improvements based on what failed, and reevaluates — up to 5 iterations. Use the model ID running in the session, so the trigger test matches what the user experiences.
At the end, return best_description — chosen by the test score, not by train, to avoid overfitting. Apply it to the frontmatter and show the before and after.
💡 How triggering works
Skills appear in available_skills with name + description, and Claude decides whether to consult it based on the description. But it only consults it for tasks it can't easily handle on its own—"read this PDF" might not trigger it even with a perfect description. That's why test queries need to be substantive enough for Claude to benefit from the skill.
✅ Module Summary
Next:
Module 4.3 — ⭐ Best Meta-Skills (for creating skills) — skill-creator, find-skills, and scaffolding: the tools skill creators use and where each fits into the workflow.