Let’s build a real skill from start to finish: "margem-xlsx" — receives a sales spreadsheet and adds a profit margin column in %. Track every stage.
🎯 Stage 1 — Capture the intent (4 questions)
Before writing a single line, skill-creator asks 4 questions. If the conversation already has the workflow, extract it from the messages and ask only what’s missing. Confirm before proceeding.
the 4 questions → answers for our case:
1. O que habilita? → adicionar coluna de margem (%) num xlsx de vendas 2. Quando dispara? → usuário menciona planilha, margem, lucro, vendas 3. Formato saída? → mesmo xlsx + nova coluna formatada como % 4. Vale testar? → SIM (saída objetivamente verificável)
💡 Verifiable output → worth testing
Question 4 determines the rest of the walkthrough. File transformations, data extraction, and code generation have objective outputs—they’re worth evaluating. Style/art skills are evaluated qualitatively. Since our case is a spreadsheet transformation, we’ll do full evals.
📝 Stage 2 — SKILL.md draft
With the intent finalized, write the draft: name, description (the trigger — assertive, saying what it does AND when to use it) and the body in the imperative, explaining why.
SKILL.md — initial draft:
--- name: margem-xlsx description: Adiciona uma coluna de margem de lucro (%) a planilhas de vendas .xlsx. Use sempre que o usuário enviar uma planilha e mencionar margem, lucro, rentabilidade ou comparar receita e custos — mesmo sem pedir "coluna" explicitamente. --- # Margem XLSX Calcule a margem como (receita - custo) / receita e grave numa nova coluna formatada como porcentagem. ## Passos 1. Abra o .xlsx e identifique as colunas de receita e custo (pergunte se ambíguo — não chute). 2. Crie a coluna "Margem" à direita, formatada como %. 3. Preserve abas, fórmulas e formatação existentes. 4. Salve mantendo o nome original com sufixo "-margem".
✗ Weak description
"Works with sales spreadsheets."
Vague; doesn’t say when to trigger, so it will trigger too rarely.
✓ Pushy description
List verbs + contexts: “margin, profit, profitability... even when no column is requested.”
What it does and when to use it, with near-triggers covered.
🧪 Stage 3 — 2-3 realistic test prompts
Create 2 to 3 prompts the way a real user would write them — with backstory, paths, and details. Show them to the user, save them in evals/evals.json without assertions yet.
evals/evals.json — ready-to-use TEMPLATE (prompts only):
{
"skill_name": "margem-xlsx",
"evals": [
{
"id": 1,
"prompt": "minha chefe mandou 'Q4 sales final FINAL v2.xlsx' (tá em Downloads) e quer margem de lucro em %. receita na col C, custo na D acho",
"expected_output": "xlsx com coluna Margem em % à direita",
"files": ["Q4 sales final FINAL v2.xlsx"]
},
{
"id": 2,
"prompt": "tenho esse relatorio de vendas mensal, da pra ver quanto a gente lucra de verdade em cada produto? planilha anexa",
"expected_output": "coluna de rentabilidade por linha em %",
"files": ["vendas_mensal.xlsx"]
},
{
"id": 3,
"prompt": "preciso comparar receita vs custo por SKU nessa planilha e ver a margem",
"expected_output": "coluna Margem = (receita-custo)/receita",
"files": ["skus.xlsx"]
}
]
}
🎯 Validate before running
Tell the user: "Here are the test cases I want to try. Are these right, or do you want to add more?" Bad prompts lead to misleading evaluations — this confirmation is cheap and avoids running everything for nothing.
📊 Stage 4 — Verifiable assertions
While the runs are in progress (don't just sit idle), write the assertions: objectively verifiable, with descriptive names, preferably checked by script. Each test case gets a eval_metadata.json.
eval_metadata.json — ready-to-use TEMPLATE (with assertions):
{
"eval_id": 0,
"eval_name": "margem-q4-receita-custo",
"prompt": "...margem de lucro em %. receita col C, custo D...",
"assertions": [
"Arquivo .xlsx de saída existe com sufixo -margem",
"Coluna 'Margem' presente à direita das demais",
"Coluna Margem formatada como porcentagem (%)",
"Valores batem com (C - D) / C por linha",
"Abas e formatação originais preservadas"
]
}
✗ Bad assertion
- ✗"the spreadsheet looks good" (subjective)
- ✗"check_1" (opaque name in the viewer)
- ✗"file exists" (passes with or without the skill)
✓ Good assertion
- ✓"Values match (C-D)/C" (script checks)
- ✓Readable name: "column formatted as %"
- ✓Distinguishes with_skill from baseline
⚖️ Stage 5 — Run with-skill vs. baseline
For each test case, launch two subagents in the same turn: one with the skill, one without (baseline = without the skill, since it's a new skill). Organize the output by iteration and by eval. Capture total_tokens e duration_ms of the notification as soon as each run finishes.
workspace structure + timing.json:
margem-xlsx-workspace/
└── iteration-1/
├── margem-q4-receita-custo/
│ ├── with_skill/outputs/ ← saída com a skill
│ ├── without_skill/outputs/ ← baseline
│ └── timing.json
└── benchmark.json ← gerado pela agregação
# timing.json
{ "total_tokens": 84852, "duration_ms": 23332,
"total_duration_seconds": 23.3 }
Why in the same turn
Launch with_skill and without_skill together so they finish at about the same time. Don’t run the with-skill cases first and come back to the baselines later—that introduces bias and makes comparison harder. timing.json can only be captured when the notification arrives; process each one right away.
🔁 Stage 6 — Evaluate and iterate (timeline)
With everything run, close the loop. The numbered timeline below is the skill-creator cycle, from grading to the next iteration.
Grade every run
A grader evaluates each assertion → grading.json with fields text, passed, evidence. For programmatic checks, run a script instead of eyeballing it.
Add the benchmark
python -m scripts.aggregate_benchmark iteration-1 --skill-name margem-xlsx → pass_rate, time, and tokens per config, with mean ± standard deviation and delta.
Open the viewer BEFORE you judge
generate_review.py with --benchmark. Put the outputs in front of the human first. Don’t write your own HTML.
Read the feedback and generalize
Empty feedback = okay. Where there are complaints, generalize the fix (don't overfit), keep it concise, explain why. Script repeated across all 3 runs → bundle in scripts/.
Rerun on iteration-2 and repeat
New iteration with baselines, reviewer with --previous-workspace. Stop when the user is happy, the feedback is empty, or there’s no progress.
💡 What happened in our case
The 3 runs with the skill wrote a calc_margem.py almost identical — a strong bundling signal. In iteration 2, the script moved to scripts/ and the SKILL.md started pointing to it. Result: pass_rate went up and tokens per run went down, because each invocation stopped reinventing the wheel.
✅ Module Summary
Next:
Module 4.5 — 🚀 Advanced Tips: Description Optimization and Benchmarking — 60/40 split, near-misses, blind comparison, and how to read the benchmark without fooling yourself.