PTENES
MODULE 2.5

🚀 Advanced Tips: Signals Only Experts See

The signals you won’t see on the storefront: triggering precision, context efficiency, reading transcripts, eval pass rate, variance—and why a skill with 100k+ installs can still be bad for you.

6
Topics
43
Minutes
Adv.
Level
Pro
Type
area of correct trigger should-trigger should-trigger near-miss Does NOT trigger near-miss
1

🎯 Triggering precision: the invisible signal

Beginners check whether the skill triggers. Experts check whether it triggers in the right cases and not in near-misses. A skill that triggers for a similar-but-wrong task injects irrelevant context and degrades the response — worse than not triggering.

✓ Precise triggering

  • ✓Triggers on should-trigger prompts (high coverage)
  • ✓Stays quiet on near-misses (high precision)
  • ✓postgres-skill does NOT trigger on "mongo query"

✗ Imprecise triggering

  • ✗Triggers whenever "database" is mentioned
  • ✗False positive wastes context and causes confusion
  • ✗False negative: never appears when needed

💡 Pro tip

Build a set of should-trigger and should-NOT-trigger examples (the near-misses). A description is only good when it gets both right. Optimize against the near-misses, not just the obvious matches.

2

🪙 Saving tokens and context

The description stays always in context (level 1). The body is loaded when triggered (level 2). Bundled resources are loaded on demand only (level 3). A token-expensive skill takes space away from all the others — experts measure this.

LevelWhat carries overWhen
1name + description (~100 words)always in the context
2SKILL.md body (<500 lines)when it triggers
3scripts / references / assetson demand

The hidden cost of always-active skills

Each installed description consumes context all the time. 30 skills with bloated descriptions = thousands of tokens spent before any task. That's why a dense, short description isn't an aesthetic choice: it's a budget.

💡 Pro tip

Move everything that isn't a trigger to level 3 (bundled). If a detail isn't read on every invocation, it doesn't belong in the main body—it becomes a file loaded on demand.

3

📜 Read transcripts to find wasted work

The signal no one sees on skills.sh is in the your own transcript. Reread a real session: where did the agent redo the same thing, ask for unnecessary confirmation, or ignore the skill it should have used? That’s where the wasted work is.

1

Look for repetition

The agent rewrote the same boilerplate three times? Turn it into a bundled skill script.

2

Look for failure to trigger

The right skill existed but didn’t activate? The description is weak on WHEN—refine the trigger.

3

Look for incorrect triggering

Did a skill show up when it wasn’t needed? That’s a near miss—tighten the description to rule it out.

💡 Pro tip

Transcript is the best free eval there is. Before installing more skills, read a session and ask: What came up repeatedly? What failed silently? Each pattern is an opportunity to adjust a skill.

4

📊 Eval pass rates and non-discriminating assertions

An eval that always passes doesn’t measure anything. The expert signal is the discriminating assertion: one that failure without the skill and passes with it. If the baseline already passes, the assertion isn't measuring the skill's contribution.

✗ Non-discriminating assertion

assert output.contains("function")
# passa com OU sem a skill → inútil

A 100% pass rate that doesn’t distinguish the baseline from the with-skill results is false reassurance.

✓ Discriminating assertion

assert migration.has_down_step()
# falha sem a skill, passa com ela ✓

Measures exactly the behavior the skill is supposed to add.

the honest eval test:

baseline (sem skill)  → pass-rate 30%
com a skill           → pass-rate 90%
delta = 60pp  → a skill comprovadamente ajuda

baseline 95% / com-skill 96%  → assertion não discrimina

💡 Pro tip

Always run the baseline. A high pass rate matters only if the baseline is low. The number that counts is the delta with-skill vs. without-skill.

5

🎲 Variance and flakiness

One eval run isn’t enough. The model is stochastic: the same skill can pass in one run and fail in the next. Experts run the eval several times and look at the variance—a 90% pass rate with high flakiness is less reliable than a stable 80%.

same skill, 5 executions:

skill A: 90 88 91 89 90  → média 89,6  σ baixo  → confiável
skill B: 100 60 95 55 90  → média 80,0  σ alto   → flaky

✓ How to reduce flakiness

  • ✓Imperative, unambiguous instructions
  • ✓Explain why (the agent generalizes more consistently)
  • ✓Run N times and report the mean + standard deviation

✗ Signs of flakiness

  • ✗Result varies a lot between runs
  • ✗Ambiguous or contradictory instructions
  • ✗Evaluate in a single round

💡 Pro tip

Always report the mean ± standard deviation, never a standalone number. Stability is quality: prefer the consistent skill over one that sometimes shines and sometimes plummets.

6

⚠️ When a popular skill is still bad

High install counts aren't proof. Remember the power law: the top 100 = 43.7% of all installs and only 0.3% (131 skills) exceed 100k — there’s early-mover bias and a brand effect. A popular skill can be bad for you.

✗ Popular but problematic

  • ✗100k installs but no commit in a year (abandoned)
  • ✗Bloated scope that triggers in the wrong workflow
  • ✗For another stack—popular in React, useless in your Vue project
  • ✗Early-mover hype, not sustained quality

✓ What to check beyond the install

  • ✓Recent last commit, answered issues
  • ✓Fit with YOUR stack and workflow
  • ✓Precise triggering for your real-world use case
  • ✓Eval delta in your transcript, not the hype

The expert’s lightbulb moment

Beginners install based on the big number. Experts install based on evidence in their own workflow: precise triggering, low context cost, a real eval delta, active maintenance, and fit with the stack. Install count is where you starts to investigate—never where it ends.

✅ Module Summary

✓
Triggering precision — triggers on should-trigger cases and stays quiet on near-misses; optimize for both.
✓
Context efficiency — a dense, concise description is a budget; move details to bundled (level 3).
✓
Reading transcripts — repetition, failure to trigger, and incorrect triggering reveal wasted work.
✓
Honest evals — discriminating assertion + delta vs baseline; a high pass rate alone is misleading.
✓
Variance and popularity — run N times (mean ± σ); high installs aren't proof — check fit, maintenance, and actual delta.

Next:

Track 3 — 🧬 Anatomy of a Skill. You already know how to judge, create, and spot the subtle signals; now let’s open up SKILL.md: frontmatter, body, and progressive disclosure.