PTENES
MODULE 4.4

🛠️ How to Create: Complete Walkthrough with Evals

A complete worked example: from intent (4 questions) to SKILL.md, realistic test prompts, evals.json with verifiable assertions, running with-skill vs. baseline, and iterating. Ready-to-copy JSON templates.

6
Topics
50
Minutes
Practical
Level
Hands-on
Type
intent draft evals run iterate repeat until satisfied

Let’s build a real skill from start to finish: "margem-xlsx" — receives a sales spreadsheet and adds a profit margin column in %. Track every stage.

1

🎯 Stage 1 — Capture the intent (4 questions)

Before writing a single line, skill-creator asks 4 questions. If the conversation already has the workflow, extract it from the messages and ask only what’s missing. Confirm before proceeding.

the 4 questions → answers for our case:

1. O que habilita?  → adicionar coluna de margem (%) num xlsx de vendas
2. Quando dispara?  → usuário menciona planilha, margem, lucro, vendas
3. Formato saída?   → mesmo xlsx + nova coluna formatada como %
4. Vale testar?     → SIM (saída objetivamente verificável)

💡 Verifiable output → worth testing

Question 4 determines the rest of the walkthrough. File transformations, data extraction, and code generation have objective outputs—they’re worth evaluating. Style/art skills are evaluated qualitatively. Since our case is a spreadsheet transformation, we’ll do full evals.

2

📝 Stage 2 — SKILL.md draft

With the intent finalized, write the draft: name, description (the trigger — assertive, saying what it does AND when to use it) and the body in the imperative, explaining why.

SKILL.md — initial draft:

---
name: margem-xlsx
description: Adiciona uma coluna de margem de lucro (%) a planilhas
  de vendas .xlsx. Use sempre que o usuário enviar uma planilha e
  mencionar margem, lucro, rentabilidade ou comparar receita e
  custos — mesmo sem pedir "coluna" explicitamente.
---

# Margem XLSX

Calcule a margem como (receita - custo) / receita e grave numa
nova coluna formatada como porcentagem.

## Passos
1. Abra o .xlsx e identifique as colunas de receita e custo
   (pergunte se ambíguo — não chute).
2. Crie a coluna "Margem" à direita, formatada como %.
3. Preserve abas, fórmulas e formatação existentes.
4. Salve mantendo o nome original com sufixo "-margem".

✗ Weak description

"Works with sales spreadsheets."

Vague; doesn’t say when to trigger, so it will trigger too rarely.

✓ Pushy description

List verbs + contexts: “margin, profit, profitability... even when no column is requested.”

What it does and when to use it, with near-triggers covered.

3

🧪 Stage 3 — 2-3 realistic test prompts

Create 2 to 3 prompts the way a real user would write them — with backstory, paths, and details. Show them to the user, save them in evals/evals.json without assertions yet.

evals/evals.json — ready-to-use TEMPLATE (prompts only):

{
  "skill_name": "margem-xlsx",
  "evals": [
    {
      "id": 1,
      "prompt": "minha chefe mandou 'Q4 sales final FINAL v2.xlsx' (tá em Downloads) e quer margem de lucro em %. receita na col C, custo na D acho",
      "expected_output": "xlsx com coluna Margem em % à direita",
      "files": ["Q4 sales final FINAL v2.xlsx"]
    },
    {
      "id": 2,
      "prompt": "tenho esse relatorio de vendas mensal, da pra ver quanto a gente lucra de verdade em cada produto? planilha anexa",
      "expected_output": "coluna de rentabilidade por linha em %",
      "files": ["vendas_mensal.xlsx"]
    },
    {
      "id": 3,
      "prompt": "preciso comparar receita vs custo por SKU nessa planilha e ver a margem",
      "expected_output": "coluna Margem = (receita-custo)/receita",
      "files": ["skus.xlsx"]
    }
  ]
}

🎯 Validate before running

Tell the user: "Here are the test cases I want to try. Are these right, or do you want to add more?" Bad prompts lead to misleading evaluations — this confirmation is cheap and avoids running everything for nothing.

4

📊 Stage 4 — Verifiable assertions

While the runs are in progress (don't just sit idle), write the assertions: objectively verifiable, with descriptive names, preferably checked by script. Each test case gets a eval_metadata.json.

eval_metadata.json — ready-to-use TEMPLATE (with assertions):

{
  "eval_id": 0,
  "eval_name": "margem-q4-receita-custo",
  "prompt": "...margem de lucro em %. receita col C, custo D...",
  "assertions": [
    "Arquivo .xlsx de saída existe com sufixo -margem",
    "Coluna 'Margem' presente à direita das demais",
    "Coluna Margem formatada como porcentagem (%)",
    "Valores batem com (C - D) / C por linha",
    "Abas e formatação originais preservadas"
  ]
}

✗ Bad assertion

  • ✗"the spreadsheet looks good" (subjective)
  • ✗"check_1" (opaque name in the viewer)
  • ✗"file exists" (passes with or without the skill)

✓ Good assertion

  • ✓"Values match (C-D)/C" (script checks)
  • ✓Readable name: "column formatted as %"
  • ✓Distinguishes with_skill from baseline
5

⚖️ Stage 5 — Run with-skill vs. baseline

For each test case, launch two subagents in the same turn: one with the skill, one without (baseline = without the skill, since it's a new skill). Organize the output by iteration and by eval. Capture total_tokens e duration_ms of the notification as soon as each run finishes.

workspace structure + timing.json:

margem-xlsx-workspace/
└── iteration-1/
    ├── margem-q4-receita-custo/
    │   ├── with_skill/outputs/   ← saída com a skill
    │   ├── without_skill/outputs/ ← baseline
    │   └── timing.json
    └── benchmark.json            ← gerado pela agregação

# timing.json
{ "total_tokens": 84852, "duration_ms": 23332,
  "total_duration_seconds": 23.3 }

Why in the same turn

Launch with_skill and without_skill together so they finish at about the same time. Don’t run the with-skill cases first and come back to the baselines later—that introduces bias and makes comparison harder. timing.json can only be captured when the notification arrives; process each one right away.

6

🔁 Stage 6 — Evaluate and iterate (timeline)

With everything run, close the loop. The numbered timeline below is the skill-creator cycle, from grading to the next iteration.

1

Grade every run

A grader evaluates each assertion → grading.json with fields text, passed, evidence. For programmatic checks, run a script instead of eyeballing it.

2

Add the benchmark

python -m scripts.aggregate_benchmark iteration-1 --skill-name margem-xlsx → pass_rate, time, and tokens per config, with mean ± standard deviation and delta.

3

Open the viewer BEFORE you judge

generate_review.py with --benchmark. Put the outputs in front of the human first. Don’t write your own HTML.

4

Read the feedback and generalize

Empty feedback = okay. Where there are complaints, generalize the fix (don't overfit), keep it concise, explain why. Script repeated across all 3 runs → bundle in scripts/.

5

Rerun on iteration-2 and repeat

New iteration with baselines, reviewer with --previous-workspace. Stop when the user is happy, the feedback is empty, or there’s no progress.

💡 What happened in our case

The 3 runs with the skill wrote a calc_margem.py almost identical — a strong bundling signal. In iteration 2, the script moved to scripts/ and the SKILL.md started pointing to it. Result: pass_rate went up and tokens per run went down, because each invocation stopped reinventing the wheel.

✅ Module Summary

✓
Intent in 4 questions — what it enables, when it triggers, format, worth testing (verifiable output → yes)
✓
Draft with pushy description — what it does AND when to use it, with an imperative body that explains why
✓
evals.json without assertions, then eval_metadata.json with — JSON templates ready to copy
✓
Run with-skill vs. baseline in the same turn — workspace per iteration and eval, capture timing.json immediately
✓
Grade → aggregate → viewer → feedback → iterate — a repeated script becomes a bundle, and the loop closes

Next:

Module 4.5 — 🚀 Advanced Tips: Description Optimization and Benchmarking — 60/40 split, near-misses, blind comparison, and how to read the benchmark without fooling yourself.