PTENES
MODULE 4.2

🔁 Test, evaluate, and iterate

The heart of the loop: test realistic prompts, run with-skill vs. baseline, evaluate qualitatively and quantitatively, generalize from feedback, and optimize the description using the test metric.

6
Topics
50
Minutes
Practical
Level
Loop
Type
Draft Test Evaluate Iterate repeat until satisfy
1

🧪 Write 2-3 realistic test prompts

After the draft, create 2-3 realistic test prompts — the kind of thing a real user would type. Show the user before running: "Here are some test cases I want to try. Are they right, or do you want to add more?". Save the prompts in evals/evals.json still no assertions—they come later, while the runs are underway.

✗ Artificial prompt

"Format this data"

"Create a chart"

Too generic. Doesn't test anything or trigger skills.

✓ Realistic prompt

"okay, my boss sent me an xlsx (it's in Downloads, something like 'Q4 sales final FINAL v2.xlsx') and wants a profit margin column in %. Revenue is in column C, costs in D, I think"

Concrete, with backstory, paths, and real details.

evals/evals.json — prompts only, no assertions yet:

{
  "skill_name": "example-skill",
  "evals": [
    {
      "id": 1,
      "prompt": "User's task prompt",
      "expected_output": "Description of expected result",
      "files": []
    }
  ]
}
2

⚖️ Run with-skill vs. baseline in the same turn

For each test case, launch two subagents in the same turn: one with the skill, one without (baseline). This matters — don't run the with-skill ones first and come back for the baselines later. Launch everything at once so they finish around the same time.

1

With-skill run

Points to the skill’s path. Save it in iteration-N/eval-ID/with_skill/outputs/.

2

Baseline — new skill

Same prompt, without any skill. Save to without_skill/outputs/.

3

Baseline — existing skill

The old version. Before editing, create a snapshot (cp -r skill snapshot/) and point the baseline to it. Save it in old_skill/outputs/.

💡 Why the baseline?

Without comparing against the baseline, you don’t know whether the skill added anything. Maybe the model would have handled the case on its own. The baseline is what separates “the skill helped” from “this would have happened anyway.”

3

📊 Evaluate: qualitative + quantitative evals

The evaluation has two fronts. Qualitative: review outputs in the viewer, click through each case, and leave feedback. Quantitative: objectively verifiable assertions that produce a pass_rate. Subjective skills (style, design) are better evaluated qualitatively—don’t force assertions where human judgment is needed.

✗ Bad assertion

  • ✗Subjective: "the output looks nice"
  • ✗Opaque name: "check_1", "assert_x"
  • ✗Always passes, with or without the skill (doesn’t discriminate)
  • ✗Forced into a creative writing skill

✓ Good assertion

  • ✓Verifiable: "margin column exists and is %"
  • ✓Descriptive name, readable in the viewer
  • ✓Checked by script, not by eye
  • ✓Distinguishes with_skill from baseline

eval_metadata.json — with assertions added:

{
  "eval_id": 0,
  "eval_name": "margem-lucro-xlsx",
  "prompt": "The user's task prompt",
  "assertions": [
    "Arquivo .xlsx de saída existe",
    "Coluna 'margem' presente e formatada como %",
    "Valores batem com (C - D) / C"
  ]
}

🎯 Show before judging

Generate the eval viewer with generate_review.py and put the outputs in front of the human before before you try to fix it yourself. The "Outputs" tab shows one case at a time; the "Benchmark" tab shows pass_rate, time, and tokens per configuration, with mean ± standard deviation and the delta.

4

🧠 How to think about improvement

This is the heart of the loop. The key insight: the skill will be used a million times with different prompts. You and the user iterate on a few examples because it's quick, but if the skill only works for those examples, it is useless. Generalize from feedback instead of overfitting.

Four ways to think

🌐

Generalize from feedback

Avoid fiddly overfit changes and oppressive MUSTs. If a problem is stubborn, try other metaphors or work patterns — it's cheap to test and could yield something great.

✂️

Keep it concise

Remove what isn’t pulling its weight. Read the transcripts, not just the final outputs: if the skill makes the model waste time, cut the part causing that and see what happens.

💬

Explain why

Even if the user’s feedback is terse or frustrated, understand the task and convey that understanding. ALL CAPS ALWAYS/NEVER is a yellow flag — rephrase it with the reason.

🔧

Look for repeated work

If the 3 test cases wrote a build_chart.py similar, that’s a strong sign you should bundle this script. Write it once, put it in scripts/, save every future invocation.

💡 Thinking time isn't the bottleneck

The skill-creator’s advice is literal: write a draft review, look at it again with fresh eyes, and improve it. Get into the user’s mindset and understand what they truly want and need. It’s worth mulling over.

5

🔁 The iteration loop

With the improvement in mind, run the cycle: apply → rerun → review → repeat. Each iteration goes in its own directory (iteration-2/, iteration-3/...), including the baselines.

1

Apply the improvements

Edit the SKILL.md based on the feedback and what you generalized from it.

2

Rerun all test cases

In a new iteration-N+1/, including baselines. New skill → baseline always without_skill.

3

Review with the user

Launch the reviewer with --previous-workspace pointing to the previous iteration for comparison.

4

Read the feedback and repeat

Empty feedback = the user thought it was okay. Focus on cases with specific complaints.

When to stop

  • •The user says they’re happy
  • •All the feedback is blank (everything OK)
  • •You’re not making meaningful progress
6

🎚️ Description optimization

The description is the primary mechanism that determines whether Claude invokes the skill. After creating or improving one, optimize it for better trigger accuracy. Generate ~20 trigger queries — a mix of should-trigger e should-not-trigger — and run the loop, choosing the description based on the test set metric.

✓ Should-trigger (8-10)

  • ✓Same intent, different phrasing (formal/casual)
  • ✓Cases where the user doesn't name the skill but needs it
  • ✓Uncommon use cases and conflicts with another skill where this one should win

✗ Should-not-trigger (8-10)

  • →Near-misses: they share keywords but need something else
  • →Adjacent domains and ambiguous phrases
  • →Avoid obvious negatives — “write a fibonacci” tests nothing

trigger-eval.json — labeled queries:

[
  {"query": "the user prompt", "should_trigger": true},
  {"query": "near-miss tricky prompt", "should_trigger": false}
]

The automatic loop

O run_loop.py splits the eval set into 60% train and 40% held-out test, evaluates the current description (running each query 3x for a reliable trigger rate), asks Claude to suggest improvements based on what failed, and reevaluates — up to 5 iterations. Use the model ID running in the session, so the trigger test matches what the user experiences.

At the end, return best_description — chosen by the test score, not by train, to avoid overfitting. Apply it to the frontmatter and show the before and after.

💡 How triggering works

Skills appear in available_skills with name + description, and Claude decides whether to consult it based on the description. But it only consults it for tasks it can't easily handle on its own—"read this PDF" might not trigger it even with a perfect description. That's why test queries need to be substantive enough for Claude to benefit from the skill.

✅ Module Summary

✓
2-3 realistic test prompts — what a real user would type, validated before running, in evals.json
✓
With-skill vs. baseline in the same turn — without a baseline, you don't know whether the skill added anything
✓
Evaluate on both fronts — qualitative review in the viewer + verifiable assertions with descriptive names
✓
Generalize, don't overfit — keep it lean, read transcripts, explain why, bundle repeated scripts
✓
The loop: apply → rerun → review → repeat — until the user is happy, feedback is empty, or there is no progress
✓
Optimize the description — should-trigger / should-not, near-misses, best_description by test metric

Next:

Module 4.3 — ⭐ Best Meta-Skills (for creating skills) — skill-creator, find-skills, and scaffolding: the tools skill creators use and where each fits into the workflow.