🎚️ The description optimization loop
The description is the primary trigger mechanism. The run_loop.py automates tuning: splits the eval set into 60% train / 40% held-out test, evaluate the current description (running each query 3 times for a reliable trigger rate), calls Claude to propose improvements based on what failed, and reevaluates — up to 5 iterations.
run the loop in the background:
python -m scripts.run_loop \ --eval-set trigger-eval.json \ --skill-path ./margem-xlsx \ --model <model-id-da-sessão> \ --max-iterations 5 \ --verbose
Split 60/40
Use train to guide proposals; use held-out test to decide. The test never trains.
3 runs per query
Triggering is stochastic; running it 3x gives you a stable trigger rate instead of a noisy yes/no.
Propose → reevaluate → repeat
Claude sees what failed and proposes a new description; the loop reevaluates it on train and test, up to 5x.
💡 Use the session’s model ID
Pass the model running the current session so the trigger test matches what the user actually experiences. Run it in the background and provide periodic updates by tailing the output—the loop takes a while.
🎯 Queries: should-trigger / should-not and near-misses
Generate ~20 realistic queries—what a real user would type, with paths, column names, backstory, sometimes lowercase or with typos. Half should-trigger, half should-not. The gold is in the near-misses.
✓ Should-trigger (8-10)
- ✓Same intent, different phrasing (formal/casual)
- ✓The user doesn’t say "margin" but needs it
- ✓Uncommon use cases; conflicts with another skill where this one should win
✗ Should-not-trigger (8-10)
- →Near-miss: "clean up the duplicates in this sales spreadsheet" (same file, different task)
- →Adjacent domain, ambiguous phrase
- →Avoid obvious negatives: “write a fibonacci” tests nothing
trigger-eval.json — labeled queries:
[
{"query": "boss mandou Q4 sales.xlsx, quer margem de lucro em %",
"should_trigger": true},
{"query": "remove as linhas duplicadas dessa planilha de vendas",
"should_trigger": false},
{"query": "qual a margem de erro dessa pesquisa de satisfação?",
"should_trigger": false}
]
Why near-misses matter
"Margin of error in a survey" shares the word "margin" but has nothing to do with profit in a spreadsheet. Negatives like this force the description to truly distinguish. Obvious negatives always pass and teach the loop nothing. Bad eval queries lead to bad descriptions.
🧲 How triggering actually works
Skills appear in the list available_skills with name + description, and Claude decides whether to consult it based on the description. The detail that changes everything: Claude only consults skills for tasks it can't easily handle on its own.
✗ Weak query for testing
"read this PDF"
"open this file"
Too simple. Claude can handle it directly with basic tools—it won’t trigger a skill, no matter how good the description is.
✓ Substantive query
"take this sales xlsx, calculate the margin by SKU, and send it back formatted as %"
Multi-step and specialized—the Claude benefits from consulting the skill.
💡 Your test queries need to be substantive
If you test triggering with "read file X," you’ll conclude the description is bad when in fact the query wouldn’t trigger any skill. Simple, single-step queries are poor test cases. Add to that the fact that Claude tends to under-trigger: that’s why the description should be a little pushy.
⚖️ Blind comparison: is the new one really better?
When the user asks, "Is the new version actually better?", there’s blind comparison. The idea: give two outputs to an independent agent without saying which is which, let it judge quality, and then analyze why the winner won.
Anonymize the outputs
Output A and Output B, with no version labels. An independent comparator evaluates them without bias.
Assess quality
The agent chooses the best option. Because it is blind, it doesn’t favor “the new” just because it’s new.
Analyze why it won
Understanding the reason is more useful than the score—it tells you what to keep in the next iteration.
Optional, and that’s fine
Blind comparison requires subagents, and most cases don't need it—the human review loop is usually enough. Save it for when the question “did it really improve?” is costly enough to justify the extra rigor.
📈 Read the benchmark without fooling yourself
The benchmark provides pass_rate, tokens e time by configuration, with the mean ± standard deviation and the delta. But the number alone can be misleading. Do an analyst’s pass before celebrating.
benchmark.json (summary) — what to look at:
with_skill: pass_rate 0.92 ± 0.05 | 84.8k tok | 23.3s without_skill: pass_rate 0.41 ± 0.18 | 61.2k tok | 18.1s delta: +0.51 pass | +23.6k tok | +5.2s # leia também: assertions que passam em AMBOS (não discriminam) # e evals com desvio alto (possivelmente flaky)
✗ Reading traps
- ✗Assertion passes with and without the skill → doesn’t measure the skill
- ✗High variance in an eval → may be flaky, not a real improvement
- ✗Celebrate pass_rate while ignoring a spike in tokens/time
✓ Healthy reading
- ✓Look at the delta between with_skill and baseline
- ✓Cut non-discriminating assertions from the set
- ✓Weigh the tradeoff: is the quality gain worth the cost?
💡 A non-discriminating assertion is noise
If “file exists” passes with_skill and without_skill, it inflates the pass_rate for both equally and hides the real difference. The signal is in the assertions that only the skill makes pass. The analyst pass exists to reveal exactly these patterns that averages hide.
🏆 Pro tips: apply best_description and package it
At the end of the loop, get the best_description — chosen by the test score, not the train score — apply it to the frontmatter and show the user the before and after with the scores. Then, package it.
loop output + apply + package:
# run_loop retorna:
{ "best_description": "...nova description afinada...",
"train_score": 0.94, "test_score": 0.89, "iterations": 4 }
# aplique no SKILL.md (frontmatter) e mostre antes/depois
# depois empacote a skill final:
python -m scripts.package_skill ./margem-xlsx
# → margem-xlsx.skill (pronto para instalar)
Pro tips checklist
Always choose based on test score — train inflates due to overfitting; held-out is honest.
Optimize the description only after that the skill's content is good; don't fine-tune the trigger for something that's still changing.
Review the eval set with the user before to run — poor queries produce poor descriptions.
Keep the description pushy but honest—cover near-triggers, yes; promise what the skill doesn’t do, no.
🎯 Why the test score
If you chose the description based on the train set, you’d reward the one that memorized the training examples. The held-out test simulates queries the loop has never seen—it’s the best proxy for the real world. Choosing based on it is what separates a description that generalizes from one that only shines in the lab.
✅ Module Summary
Next:
Track 5 — 🧠 What to Think About — principles and pitfalls when deciding what becomes a skill, how to keep them healthy, and what to avoid in the long run.