🎯 Triggering precision: the invisible signal
Beginners check whether the skill triggers. Experts check whether it triggers in the right cases and not in near-misses. A skill that triggers for a similar-but-wrong task injects irrelevant context and degrades the response — worse than not triggering.
✓ Precise triggering
- ✓Triggers on should-trigger prompts (high coverage)
- ✓Stays quiet on near-misses (high precision)
- ✓postgres-skill does NOT trigger on "mongo query"
✗ Imprecise triggering
- ✗Triggers whenever "database" is mentioned
- ✗False positive wastes context and causes confusion
- ✗False negative: never appears when needed
💡 Pro tip
Build a set of should-trigger and should-NOT-trigger examples (the near-misses). A description is only good when it gets both right. Optimize against the near-misses, not just the obvious matches.
🪙 Saving tokens and context
The description stays always in context (level 1). The body is loaded when triggered (level 2). Bundled resources are loaded on demand only (level 3). A token-expensive skill takes space away from all the others — experts measure this.
| Level | What carries over | When |
|---|---|---|
| 1 | name + description (~100 words) | always in the context |
| 2 | SKILL.md body (<500 lines) | when it triggers |
| 3 | scripts / references / assets | on demand |
The hidden cost of always-active skills
Each installed description consumes context all the time. 30 skills with bloated descriptions = thousands of tokens spent before any task. That's why a dense, short description isn't an aesthetic choice: it's a budget.
💡 Pro tip
Move everything that isn't a trigger to level 3 (bundled). If a detail isn't read on every invocation, it doesn't belong in the main body—it becomes a file loaded on demand.
📜 Read transcripts to find wasted work
The signal no one sees on skills.sh is in the your own transcript. Reread a real session: where did the agent redo the same thing, ask for unnecessary confirmation, or ignore the skill it should have used? That’s where the wasted work is.
Look for repetition
The agent rewrote the same boilerplate three times? Turn it into a bundled skill script.
Look for failure to trigger
The right skill existed but didn’t activate? The description is weak on WHEN—refine the trigger.
Look for incorrect triggering
Did a skill show up when it wasn’t needed? That’s a near miss—tighten the description to rule it out.
💡 Pro tip
Transcript is the best free eval there is. Before installing more skills, read a session and ask: What came up repeatedly? What failed silently? Each pattern is an opportunity to adjust a skill.
📊 Eval pass rates and non-discriminating assertions
An eval that always passes doesn’t measure anything. The expert signal is the discriminating assertion: one that failure without the skill and passes with it. If the baseline already passes, the assertion isn't measuring the skill's contribution.
✗ Non-discriminating assertion
assert output.contains("function")
# passa com OU sem a skill → inútilA 100% pass rate that doesn’t distinguish the baseline from the with-skill results is false reassurance.
✓ Discriminating assertion
assert migration.has_down_step() # falha sem a skill, passa com ela ✓
Measures exactly the behavior the skill is supposed to add.
the honest eval test:
baseline (sem skill) → pass-rate 30% com a skill → pass-rate 90% delta = 60pp → a skill comprovadamente ajuda baseline 95% / com-skill 96% → assertion não discrimina
💡 Pro tip
Always run the baseline. A high pass rate matters only if the baseline is low. The number that counts is the delta with-skill vs. without-skill.
🎲 Variance and flakiness
One eval run isn’t enough. The model is stochastic: the same skill can pass in one run and fail in the next. Experts run the eval several times and look at the variance—a 90% pass rate with high flakiness is less reliable than a stable 80%.
same skill, 5 executions:
skill A: 90 88 91 89 90 → média 89,6 σ baixo → confiável skill B: 100 60 95 55 90 → média 80,0 σ alto → flaky
✓ How to reduce flakiness
- ✓Imperative, unambiguous instructions
- ✓Explain why (the agent generalizes more consistently)
- ✓Run N times and report the mean + standard deviation
✗ Signs of flakiness
- ✗Result varies a lot between runs
- ✗Ambiguous or contradictory instructions
- ✗Evaluate in a single round
💡 Pro tip
Always report the mean ± standard deviation, never a standalone number. Stability is quality: prefer the consistent skill over one that sometimes shines and sometimes plummets.
⚠️ When a popular skill is still bad
High install counts aren't proof. Remember the power law: the top 100 = 43.7% of all installs and only 0.3% (131 skills) exceed 100k — there’s early-mover bias and a brand effect. A popular skill can be bad for you.
✗ Popular but problematic
- ✗100k installs but no commit in a year (abandoned)
- ✗Bloated scope that triggers in the wrong workflow
- ✗For another stack—popular in React, useless in your Vue project
- ✗Early-mover hype, not sustained quality
✓ What to check beyond the install
- ✓Recent last commit, answered issues
- ✓Fit with YOUR stack and workflow
- ✓Precise triggering for your real-world use case
- ✓Eval delta in your transcript, not the hype
The expert’s lightbulb moment
Beginners install based on the big number. Experts install based on evidence in their own workflow: precise triggering, low context cost, a real eval delta, active maintenance, and fit with the stack. Install count is where you starts to investigate—never where it ends.
✅ Module Summary
Next:
Track 3 — 🧬 Anatomy of a Skill. You already know how to judge, create, and spot the subtle signals; now let’s open up SKILL.md: frontmatter, body, and progressive disclosure.