PTENES
TRACK 4

🛠️ How to Create (the Loop)

From intent to first draft, and the test → evaluate → iterate loop with evals and description optimization. The Anthropic skill-creator method.

Draft Test Evaluate Iterate
5
Modules
30
Topics
~3h30
Duration
Practical
Level
4.1~45 min

🌱 From intent to first draft

Capture what the person wants, do enough research, and turn it into a well-written SKILL.md from the start — imperative, with the why, and without shouting MUSTs.

What it is:

Before writing a single line, skill-creator asks 4 questions: what the skill enables Claude to do, when it should trigger, what output format is expected, and whether it's worth putting together test cases.

Why learn:

Most bad skills come from poorly defined intent. If the conversation already contains the workflow ("turn this into a skill"), extract it from the messages first and only ask for what’s missing.

Key concepts:

what it enables · when it triggers · output format · need tests? · confirm before proceeding

What it is:

Proactively ask about edge cases, input and output formats, sample files, success criteria, and dependencies. Research in parallel with subagents when useful MCPs are available.

Why learn:

Arriving with context ready reduces friction for the user. Write test prompts only after completing this part — a poor interview creates a skill that covers the happy path and breaks everywhere else.

Key concepts:

edge cases · input/output formats · example files · success criteria · dependencies · parallel research

What it is:

Based on the interview, fill in name (identifier), description (trigger: what it does AND when to use it), and the Markdown body with the instructions. Every "when to use" detail belongs in the description, not the body.

Why learn:

The description is the primary trigger mechanism. Since Claude tends to under-trigger, it should be a little "pushy" — list concrete contexts where the skill should kick in even when the user doesn't explicitly ask for it.

Key concepts:

name · pushy description · what + when · body <500 lines · compatibility (rare)

What it is:

Write in the imperative, use theory of mind, explain why each instruction exists instead of piling on uppercase MUSTs, and keep the skill general instead of tailoring it to the examples.

Why learn:

Today’s LLMs are smart: when you explain why, they go beyond rote instructions and solve the real case. ALL CAPS ALWAYS/NEVER is a yellow flag—rephrase it to explain the reason.

Key concepts:

imperative · theory of mind · explain why · avoid MUSTs · general skill, not narrow

What it is:

Write a first draft without getting stuck, then reread it with fresh eyes and improve it. Draft → review → improve is a cycle within the writing itself.

Why learn:

The first draft is almost never the best. A fresh reread reveals redundant or ambiguous instructions, or ones that tell the model to waste time.

Key concepts:

quick draft · reread with fresh eyes · cut redundancy · clarity · iterate on the writing

What it is:

Decide when to move resources out of SKILL.md: if 3 runs repeat the same script, turn it into scripts/; large docs go in references/ loaded on demand.

Why learn:

Progressive disclosure: metadata always in context, body when triggered, resources only when needed. Bundling a repeated script saves every future invocation from reinventing the wheel.

Key concepts:

scripts/ · references/ · assets/ · progressive disclosure · rule of 3 repetitions · TOC in docs >300 lines

View Full
4.2~50 min

🔁 Test, evaluate, and iterate

The heart of the loop: test realistic prompts, run with-skill vs. baseline, evaluate qualitatively and quantitatively, generalize from feedback, and optimize the description using the test metric.

What it is:

After the draft, create 2-3 realistic test prompts — the kind of thing a real user would type — and show them to the user for validation before running them.

Why learn:

Artificial test prompts produce misleading evaluations. The prompts are in evals/evals.json without assertions yet — the assertions come later, while the runs are running.

Key concepts:

2-3 prompts · real user language · validate with the user · evals.json · no assertions yet

What it is:

For each test case, launch two subagents in the same turn: one with the skill, one without (baseline). Launch them all at once so they finish together.

Why learn:

Without a baseline, you don’t know whether the skill added anything. For a new skill, the baseline = no skill at all. For an existing skill, the baseline = the old version (snapshot before editing).

Key concepts:

with_skill · without_skill · same turn · snapshot of the old version · workspace per iteration

What it is:

Evaluate on two fronts: qualitative (review the outputs in the viewer) and quantitative (verifiable assertions that produce a pass_rate). Subjective skills are evaluated only qualitatively.

Why learn:

Good assertions are objectively verifiable and have descriptive names. Don't force assertions on things that call for human judgment (writing style, design).

Key concepts:

qualitative in the viewer · verifiable assertions · pass_rate · descriptive names · don’t overfit to subjective criteria

What it is:

Generalize from feedback instead of gradually overfitting to a few examples, keep the prompt lean, and read the transcripts (not just the final outputs).

Why learn:

The skill will be used a million times in different prompts. If it only works for the test examples, it’s useless. Avoid fiddly changes and oppressive MUSTs.

Key concepts:

generalize · don't overfit · keep it concise · read transcripts · explain why · repeated script becomes a bundle

What it is:

Apply the improvement → rerun all test cases in a new iteration → review with the user → read the feedback → repeat until it meets the requirements.

Why learn:

The loop stops when the user is happy, all feedback is empty, or you’re no longer making meaningful progress. Each iteration goes in its own directory.

Key concepts:

apply → rerun → review → repeat · iteration-N · previous-workspace · stopping criteria

What it is:

Generate ~20 trigger queries (a mix of should-trigger and should-not-trigger), focus on near-misses, run the optimization loop, and choose the description based on the test set metric.

Why learn:

The description determines whether the skill triggers. Obvious negatives don't test anything — the valuable ones are near-misses that share keywords but need something else.

Key concepts:

20 queries · should-trigger / should-not · near-misses · 60% train / 40% test · best_description based on the test set

View Full
4.3~45 min

⭐ The Best Meta-Skills (for creating skills)

The tools for people who CREATE skills: skill-creator (246k), find-skills (1,8M), and the scaffolding pattern. What each one does, when to use it, and where it fits in the workflow.

What it is:

A skill whose job is to help you work with other skills—discovering, creating, testing, optimizing, and packaging them. It operates one level above an ordinary skill.

Why learn:

Many people create without checking whether something better already exists and without an eval loop. Meta-skills solve this: discover before creating, create methodically, validate with data.

Key concepts:

meta-skill · discover · create · test · optimize · package

What it is:

Anthropic’s meta-skill orchestrates the entire cycle—draft → eval → iterate—and includes a separate description optimizer.

Why learn:

It’s the central axis of the creation workflow. Everything in modules 4.1 and 4.2 comes from it; don’t improvise a parallel process.

Key concepts:

draft → eval → iterate · scripts · aggregate_benchmark · run_loop · package_skill

What it is:

The most installed skill in the catalog (1.802.925, vercel-labs). It finds relevant skills for a task before you create one from scratch.

Why learn:

Prevents the most costly mistake creators make: spending hours writing something that already exists in a better form. It's step zero in the workflow.

Key concepts:

discovery · use / extend / create · step zero · catalog of 39.366 skills

What it is:

A standard way to generate the initial scaffold—SKILL.md with frontmatter and scripts/, references/, and assets/ folders—instead of typing everything by hand.

Why learn:

It speeds up the start and removes the blank page. But create only the folders you’ll use—empty folders confuse the model and violate the "keep it lean" principle.

Key concepts:

skeleton · pre-filled frontmatter · don’t create everything upfront · trim boilerplate

What it is:

The sequence that prevents rework: find-skills (discover) → scaffolding (generate a base) → skill-creator (create and iterate) → description optimizer (refine the trigger).

Why learn:

The three don’t compete; they connect in sequence. Using them out of order is how skills with fewer than 100 installs come about.

Key concepts:

discover → foundation → create/iterate → optimize → package · golden rule

What it is:

A situation → tool cheat sheet: "I need a skill for X" → find-skills; "I’m going to create one" → skill-creator; "it doesn’t trigger when it should" → description optimizer.

Why learn:

Make the right decision at the right time. Install all three and, since Claude tends to under-trigger them, mention them explicitly the first few times until it becomes automatic.

Key concepts:

situation → meta-skill · install all three · pro tip for triggering · package_skill

View Full
4.4~50 min

🛠️ How to Create: Complete Walkthrough with Evals

A complete worked example: intent (4 questions) → draft → test prompts → evals.json with verifiable assertions → run with-skill vs. baseline → iterate. Ready-to-use JSON templates.

What it is:

The 4 questions applied to the “margem-xlsx” case: what does it enable, when does it trigger, what’s the output format, and is it worth testing? Verifiable output → worth testing.

Why learn:

Question 4 determines the rest of the walkthrough. File transformations and code generation are worth evaluating; style/art aren’t.

Key concepts:

what it enables · when it triggers · format · worth testing · extract from conversation · confirm

What it is:

Write the complete draft: name, pushy description (what it does AND when to use it), and an imperative body that explains why.

Why learn:

The description is the trigger. "Works with spreadsheets" under-triggers; listing verbs + contexts ("margin, profit... even without asking for a column") covers near-triggers.

Key concepts:

name · pushy description · imperative body · explain why · don’t guess columns

What it is:

2-3 prompts written as a real user would write them — with backstory, paths, and details — saved in evals/evals.json with no assertions yet. JSON template ready.

Why learn:

Artificial prompts produce misleading evaluations. Validate with the user before running—it’s cheap and avoids running everything for nothing.

Key concepts:

2-3 prompts · natural language · paths and backstory · evals.json · validate first

What it is:

While the runs are in progress, write objective assertions with descriptive names in eval_metadata.json—preferably checked by a script. Ready-to-use template.

Why learn:

"the spreadsheet looks good" is subjective; "values match (C-D)/C" is verifiable and distinguishes with_skill from baseline.

Key concepts:

verifiable assertions · descriptive name · checked by a script · discriminates · not subjective

What it is:

Two subagents in the same turn per test case (with skill / without skill = baseline), output organized by iteration and eval, with total_tokens and duration_ms in timing.json.

Why learn:

Launching them together avoids bias. Timing can only be captured when the notification arrives—process each one right away.

Key concepts:

with_skill / without_skill · same turn · workspace per iteration · timing.json

What it is:

Numbered timeline: grade each run → aggregate the benchmark → open the viewer before you judge → read the feedback and generalize → rerun in iteration-2.

Why learn:

In this case, all 3 runs wrote nearly identical calc_margem.py files — a sign they should be bundled. In iteration 2, the pass_rate went up and token usage went down.

Key concepts:

grading.json · aggregate_benchmark · generate_review · feedback · repeated script bundle

View Full
4.5~50 min

🚀 Advanced Tips: Description Optimization and Benchmarking

The optimization loop (60/40 split, near-misses, best_description based on the test score), how triggering really works, blind comparison, and how to read the benchmark without fooling yourself.

What it is:

run_loop.py splits the eval set into 60% train / 40% test, evaluates the description (3 runs per query), proposes improvements based on what failed, and reevaluates, up to 5x.

Why learn:

Activation is stochastic; 3 runs provide a stable trigger rate. Use the session’s model ID so the test matches what the user experiences.

Key concepts:

run_loop · 60/40 · 3 runs/query · propose → reevaluate · session model ID · background

What it is:

~20 realistic queries, half should-trigger and half should-not. The gold is in the near-misses — phrases that share words but need something else.

Why learn:

"Margin of error in a survey" uses "margin" but isn't spreadsheet profit — it forces the description to distinguish. Obvious negatives teach nothing.

Key concepts:

should-trigger 8-10 · should-not 8-10 · near-misses · avoid obvious ones · trigger-eval.json

What it is:

Skills appear in available_skills with a name + description, and Claude decides whether to consult them—but only for tasks it can’t easily solve on its own.

Why learn:

"Read this PDF" may not trigger even with a perfect description. Your test queries need to be substantive (multi-step, specialized).

Key concepts:

available_skills · non-trivial tasks only · substantive queries · under-triggering · pushy

What it is:

Give two outputs to an independent agent without saying which is which, let it judge their quality, and then analyze why the winner won.

Why learn:

Because it’s blind, it doesn’t favor “the new” just because it’s new. It’s optional and requires subagents—save it for when the question is costly.

Key concepts:

anonymize · judge quality · why it won · optional · subagents · human review is enough

What it is:

The benchmark provides pass_rate, tokens, and time by config, with mean ± standard deviation and delta. The analyst’s pass reveals what the average hides.

Why learn:

An assertion that passes with and without the skill measures nothing and inflates the pass_rate; high deviation may be flaky; a high pass_rate may hide a token explosion.

Key concepts:

pass_rate · delta · non-discriminating · variance/flaky · token/time tradeoff

What it is:

Take the best_description (chosen by the test score, not the train score), apply it to the frontmatter, show the before and after, and package it with package_skill.

Why learn:

Choosing based on the train set would reward a description that memorized the examples. The held-out test simulates the real world—it's what separates generalizing from shining only in the lab.

Key concepts:

best_description · test score > train · optimize only after it's ready · review the eval set · package_skill

View Full

Learning path overview

4.1~45 min
🌱 From intent to first draft

Ask, research, write. A draft that starts with imperative instructions and explains why.

4.2~50 min
🔁 Test, evaluate, and iterate

Run, measure, generalize. The loop that turns a draft into a skill that works a million times.

4.3~45 min
⭐ The Best Meta-Skills

find-skills discovers, skill-creator creates, scaffolding generates the foundation. The tools for people who create skills.

4.4~50 min
🛠️ Complete Walkthrough with Evals

From intent to evals.json with assertions. A worked example from start to finish, JSON ready to copy.

4.5~50 min
🚀 Description Optimization and Benchmarking

60/40 split, near misses, best_description based on the test. Fine-tune the trigger and read the benchmark without fooling yourself.

← Home Track 5 →