PTENES
Skip to content
MODULE 2.2

🎯 From micromanagement to criteria and verification

Stop scripting “do A, then B, then C.” Write objective + guardrails + exit criteria + verification — and let the model choose the path. Verification is the highest-impact lever, and it’s exactly what almost everyone forgets to put in the config.

6
Topics
50
Minutes
Intermediate
Level
Practical
Type
Progress in this module
0%0 of 6
1

🔍 Recognize over-specification

Boris Cherny describes the over-specification — writing the exact sequence of moves the model must make — as one of the most common mistakes he sees. And the uncomfortable detail: it’s more common among people with years or decades of engineering experience, no less. The more experienced the config's author, the more scripted it tends to be.

The reason is historical. Software used to be built by designing the entire system up front, writing a huge suite of unit tests, and treating a re-architecture as a project that took months or years. That reflex—decide everything first, on paper—became the way to write CLAUDE.md. It worked with the old models: they needed you to guide every step. Today, the script does the opposite of what you want: it keeps the model from finding a better path than the one you imagined.

🆕 Four words that will appear throughout the module

  • Guardrail: a limit that can’t be crossed — “never commit without running the tests,” “don’t touch node_modules/”. Say what is forbidden, not how to do it.
  • Exit criterion: the phrase that lets the model know it’s finished. “Done when the build passes and the file exists” is a criterion; “done when it’s good” isn’t.
  • Verification: the concrete form of the model check on your own whether the exit criterion was met — a command, a test, a file comparison, a screenshot.
  • Autonomy: explicit permission for the model to choose the execution strategy, as long as it respects the guardrails and meets the criterion.
RECIPE: RIGID TRACK 12 34 56 78 the better path stays blocked CRITERION: OPEN FIELD guardrailguardrail criterionof output three routes, one gate

What to look at: on the left, there's only one route, and it's the route that you imagined. If there's a better path, the track prohibits it — and the model obeys. Nothing was scripted on the right: the fences are the guardrails (what's not allowed), the gate is the exit criterion (when it's done), and the middle is open. The three arrows show the concrete benefit: the model can choose the shortest route to today's real problem, which isn't necessarily the one you wrote months ago.

✗ Signs of over-specification

  • ✗“First run git status, then git diff, then read the file, then…”
  • ✗Numbering 8 to 12 steps for a task that fits in two sentences
  • ✗“Use tool X, never Y” without saying why (the reason is what gets old)
  • ✗Stacked exceptions: “except when…, but if…, unless…”
  • ✗Instruction that teaches the model to think (“analyze carefully before responding”)

✓ What survives the audit

  • ✓“Deploy happens via git push; the webhook takes care of the rest.” — context the model can’t infer
  • ✓“Never print the value of an API key.” — security guardrail
  • ✓“The keys are in ~/projetos/wifi/.env.” — source of truth, concrete path
  • ✓“Done when npm run build exits with code 0.” — verifiable criterion
  • ✓“Choose the execution strategy.” — stated autonomy

Key concepts

Over-specification

The exact sequence of actions

Old reflex

Design everything first, on paper

Real cost

Blocks the better path

Experience gets in the way

More common among veterans

2

🔄 Turn the recipe into criteria

There’s a canonical conversion, and it fits on one line. Every instruction in the form “do A, then B, then C” becomes: “Produce X. Respect Y. The result must meet Z. Verify using W. Choose the strategy.” Five slots. X is the objective, Y is the guardrails, Z is the exit criterion, W is the verification, and the last sentence is the autonomy.

🎯 The template sentence

Produce X. Respect Y. The result should reach Z. Check using W. Choose the strategy.

If you can’t fill in W, stop: you don’t have an instruction yet, you have a wish. The absence of W is the most common flaw in the configs the audit finds.

A real case: scripted publishing rule

Below is a publishing rule written the way almost everyone CLAUDE.md real writes — eight steps, fixed order, one command per line. It works, but locks the model into a sequence that is no longer the best one (and breaks on the first project with an extra step).

✗ BEFORE — 8 steps, 14 lines

## Publicação

1. Rode `git status` e verifique se há mudanças.
2. Rode `git diff` e leia tudo o que mudou.
3. Rode `git log -5` para ver o estilo das mensagens anteriores.
4. Rode `git config user.email` e confira se é o autor certo.
5. Se estiver errado, rode `git config user.email `.
6. Faça `git add` só dos arquivos que você editou (nunca `git add .`).
7. Escreva a mensagem de commit no padrão `tipo: descrição`.
8. Rode `git push` e depois confira no dashboard se o deploy subiu.

✓ AFTER — 4 lines of criteria

## Publicação

Objetivo: publicar = commit + push no `origin`. O deploy é automático; não é sua responsabilidade.
Guardrails: nunca `git add .`; o autor do commit acompanha a conta de destino do repo (default `inematds`).
Critério de saída: o commit está em `origin` e a mensagem segue `tipo: descrição`.
Verificação: `git log origin/main -1` mostra o seu commit com o autor esperado.

What actually changed: steps 1–3 disappeared because the model already knows how to inspect a repository — they were pure micromanagement. Steps 4–5 became a guardrail (the rule, not the procedure). Steps 6–7 became guardrails and criteria. And step 8 — “check the dashboard” — became a executable verification: a command that gives you the answer without anyone needing to open a website.

💡 The scalpel test

For each numbered step in the instruction, ask: “is this a rule that’s always true, or a move the model would choose on its own?” A rule becomes a guardrail. A movement becomes nothing—delete it. What remains is usually 20–30% of the original, and it’s exactly what the model couldn’t guess.

3

🧱 Use the 6-field template

The template sentence from the previous topic, expanded, becomes a six-field template: Objective · Context · Guardrails · Quality criteria · Verification · Autonomy. It’s not bureaucracy: each field exists because leaving it out causes a specific kind of failure. This applies to standalone prompts and to blocks in the CLAUDE.md and for the body of a skill.

1

Objective — the result, not the process

One sentence saying what should exist at the end.

Typical mistake: describing the activity (“review the code”) instead of the result (“a report with the correctness bugs found in the diff”).

2

Context — only what it can’t discover on its own

Paths, system names, sources of truth, internal conventions.

Typical mistake: dumping the project history. Context is an address, not a biography. If the model can read the file, say where it is instead of summarizing it.

3

Guardrails—the prohibitions, stated negatively

Security, compliance, interface contracts, anything that must never be touched.

Typical mistake: disguising a procedure as a guardrail. “Always run lint before the build” isn’t a limit—it’s a step. A limit is “don’t commit with lint failing.”

4

Quality criteria — what “good” looks like

Preferably measurable: maximum number of lines, output format, coverage, tone.

Typical mistake: adjectives alone—“clear,” “professional,” “well-written.” An adjective without an anchor doesn’t decide anything, and the model fills in the blanks with its own guess.

5

Verification — how it checks by itself

One command, one test, one comparison, one file check. With a stopping condition.

Typical mistake: the field simply doesn't exist. It's the most absent field and the one with the greatest impact — topic 4 is entirely about it.

6

Autonomy — explicit permission

“Choose the execution strategy.” One line, and it changes the behavior.

Typical mistake: thinking it’s redundant. Without this line, a config full of old rules still pushes the model into “follow the script” mode.

📋 Ready-to-paste template

Copy, replace what’s between < > and delete the fields that don't apply—except Verification, which is never removed.

## <nome da tarefa>

**Objetivo:** <o resultado que deve existir no final, em uma frase>

**Contexto:** <caminhos, sistemas e convenções que o modelo não descobre sozinho>

**Guardrails:**
- <o que nunca pode acontecer>
- <limite de segurança / compliance / interface>

**Critérios de qualidade:**
- <algo contável: tamanho, formato, cobertura>
- <algo observável: tom, estrutura, nomes>

**Verificação:** <comando, teste ou comparação que prova o critério>.
Repita produzir → verificar até <condição de parada explícita>.

**Autonomia:** escolha a estratégia de execução; não peça confirmação
para decisões cobertas pelos guardrails acima.

💡 Tip: the template is a measuring stick, not a form

You don't need the six literal headings in the config. You need the six exist somewhere. Use the template to audit an instruction: read it and mark which of the six fields it covers. An instruction that covers only “the path” and none of the six is a direct candidate for REMOVE or SIMPLIFY in the module 2.1 taxonomy.

4

🔬 Install the check (the lever)

They asked Boris which skill matters now that “prompt engineering” has lost traction as a job title. The answer: give the model a task that seems a little too difficult, and make it possible for it to check its own work throughout the process. He called verification “the most important thing people get wrong.” Put in terms of this course: a short prompt with a real way to check it beats a giant prompt with none.

The Electron → Swift example

Boris wanted to see what Claude’s desktop app — built with Electron — would look like as a native app. He connected a macOS runner (a Mac virtual machine), created an empty repository, and gave it a single instruction: rewrite the Electron app in Swift, run the original in the VM, take screenshots, compare them pixel by pixel with the Swift version, and don’t stop until you’re done. On the date of the talk, the task had been running for more than two weeks, creating thousands to tens of thousands of agents. The prompt had no technique at all. It had a complete verification cycle.

THE CYCLE THAT MAKES A SHORT PROMPT WORK FOR WEEKS producerewrite in Swift observerun the original in the VM comparepixel by pixel stopping condition“don’t stop until you’re done” until it passes, go back and produce it again without the cyan feedback, the model delivers its first attempt and stops

What to look at: the cyan return arrow. Almost everyone writes the first three boxes; the return path and the highlighted box are what’s missing. Without an explicit stopping condition, the cycle isn’t a cycle— it’s a straight line that ends with the first delivery. And notice that observe is the box that requires a real-world resource (the VM, the command, the file). Verification without external observation is just the model rereading itself.

Four objective checks in different contexts

1. Test that runs

“Done when pytest tests/test_parser.py passes in full. If it fails, fix it and run it again.” External observation: the test result. Stop condition: zero failures.

2. A command whose exit code proves it

“Run npm run build; consider it complete only with exit code 0.” External observation: the exit code. Stop condition: 0. There’s no room for interpretation.

3. File comparison

“Generate the CSV and run diff saida.csv esperado.csv; repeat until the diff is empty.” External observation: the reference file. Stop condition: empty diff.

4. Screenshot / visual comparison

“Open the page, take a screenshot, compare it with referencia.png; fix it until the difference is below 2%.” External observation: the rendered image. Stop condition: the threshold.

⚠️ Why “review before delivering” is NOT verification

It's the most common line in real-world configs, and it doesn't verify anything. It's missing the two pieces that make the mechanism work:

  • ✗There is no external observation. The model rereads its own text in the same context that produced it—the reasoning that caused the error is still there, pushing it toward “it’s good enough.”
  • ✗There is no stopping condition. “Review” happens once and that’s it. There’s no criterion for whether the review was sufficient, so there’s never a second pass.

Fix: replace it with “run <comando>; while it fails, fix it and run it again.” Same sentence length, completely different mechanism.

The four moments, in order

1

Produce

Generate the output. It’s the only moment every config already has.

2

Observe

Look at something out from the text itself: run a command, open a file, take a screenshot.

3

Compare

Compare the output with the reference and produce a verdict—pass or fail.

4

Stopping condition

The sentence that closes the loop. “Don’t stop until you’re done,” “until the diff is empty,” “until the exit code is 0.”

5

🚀 Aim above the ceiling

Boris’s recommendation is to aim higher than feels comfortable: give the model a task that’s a little harder than you think it can handle. The reason is simple— the ceiling changes with every release. Your estimate of “what it can do” was calibrated to an earlier generation, and you didn’t realize it had become outdated because no one tells you when the ceiling rises.

The direct consequence for the audit: retest your past failures with every new version. That rule you wrote because “it always gets this wrong” may be protecting against a failure that no longer exists. It still costs context on every run, and the reason for it disappeared without notice. This is exactly the mechanism from module 1.1: an instruction is a fix for a specific model’s weakness.

🧬 Empirical science, not theory

Working with a model is more like getting to know a living creature than configuring a system.

Each generation behaves differently and has a slightly different personality. You spend some time learning how it works and then adjust the harness — the set of prompts, tools, and rules around the model — based on what you observed. Notice the order: observe first, write the rule afterward. The opposite is theory, and theories about models go stale in weeks.

💡 The coworker standard

The right level of instruction is what you’d use to guide a capable colleague: context, the goal, the constraints they can’t guess, and how to know when they’re done. You wouldn’t tell a colleague, “open the editor, click File, type the name.” If your instruction wouldn’t pass the test of being read aloud to someone without sounding offensive, it’s micromanagement.

✓ Empirical approach

  • ✓Give it a task that’s harder than you expect, and observe where it gets stuck
  • ✓Keep past failures and rerun with each new model
  • ✓Fixes the difficulty specific observed — nothing beyond it
  • ✓Choose the right remedy: a better prompt, skill, or MCP

✗ Theoretical approach

  • ✗Assumes the previous generation's limit and never tests again
  • ✗Write preventive rules for failures you've never seen happen
  • ✗Fixes “in the area” — three new paragraphs for one specific error
  • ✗Look for the secret trick from a social media influencer

🔁 There's no secret trick—there's a cycle

They asked Boris what separates the top 1% of Claude Code users. His answer was: stop looking for a trick. The method is this, and it’s public:

1. tarefa difícil demais
2. meios de verificar o trabalho
3. observar onde ele trava
4. corrigir AQUILO  →  prompt melhor  (instrução obscura)
                    →  skill          (falta procedimento repetível)
                    →  MCP            (falta contexto que ele não alcança)
5. repetir

Step 4 is where most people get it wrong: when faced with a stumble, the reflex is always “one more rule in the CLAUDE.md”. Three causes, three different remedies—and only one of them is writing text.

6

✍️ Rewrite your worst instruction

Time to apply this. Open your configuration— ~/.claude/CLAUDE.md, o CLAUDE.md of a project or the body of one of your skills—and find the most “cookbook-style” instruction there: the most numbered, the longest, the one that describes steps. You’ll write two versions of it, and the second must include objective verification.

a

50% smaller version

Same structure, half the text. Cut what the model already does on its own, and combine steps that always go together. No actual rule can disappear in this version—this is cutting fat, not muscle.

b

Minimal version, using the 6-field template

A complete rewrite in the form Objective · Context · Guardrails · Criteria · Verification · Autonomy. Verification needs to be a command, test, comparison, or file check—something that produces an observable result.

🧪 Prompt ready to paste into Claude Code

Objective: have Claude Code itself produce both versions side by side, without letting it apply anything to the file.

Read <path to your CLAUDE.md or skill> and locate the block
"<title or first line of the most scripted block>".

DO NOT edit any files. Only produce text in your response.

Return three Markdown blocks:

1. ORIGINAL — the excerpt exactly as it appears today, with the line count.

2. 50% SHORTER VERSION — the same excerpt with half the lines. Cut only what
you would do on your own without the instruction. List below, one per line,
what was cut and why.

3. MINIMAL VERSION — rewrite it using the fields Objective, Context, Guardrails,
Quality criteria, Verification, and Autonomy. Verification must be a command,
a test, a file comparison, or an existence check — nothing like “review before
submitting” — and must include a stopping condition.

At the end, answer in one sentence: could a colleague who doesn't know this
project run the Verification in the minimal version without asking me anything?
If the answer is no, rewrite the Verification.

Exit criterion for this exercise: both versions are written alongside the original, and the minimum version’s verification is executable by a third party without asking you anything. If it depends on context that exists only in your head, it isn’t verification yet.

How to check this in practice: send the minimum version to someone (or paste it into a new session with no history) and ask them only to run the Verification line. If they ask a follow-up question, rewrite it. Keep all three versions — they become direct input for the A/B/C plan in Track 4.

Quick check (doesn't block anything): your minimal instruction ends with “before finishing, carefully review the result and fix anything that's wrong.” Does that count as verification?

Key concepts

Two versions

50% smaller and minimal

Side by side

The original stays visible

Third-party test

Can run without asking you

A/B/C inputs

The three versions become a test

📌 Module Summary

✓
Over-specification is the most common mistake — and more common among people with more years of engineering experience. The script blocks the better path.
✓
The canonical conversion — “Produce X. Respect Y. The result must meet Z. Verify using W. Choose the strategy.”
✓
Six fields — Objective · Context · Guardrails · Criteria · Verification · Autonomy. Use it as a yardstick for auditing, not as a form.
✓
Verification is the lever — produce → observe → compare → stopping condition. Without external observation and a stopping condition, it's not verification.
✓
Aim above the ceiling — and retest your old failures with every new model; the reason for the rule may have disappeared.
✓
No secret trick — difficult task → ways to verify → observe where it gets stuck → fix that → repeat.

Next Module:

3.1 — Running the skill audit-ablacao: install it, choose the scope, and read the 10-section report knowing what each section asks of you.