🏷️ Classify each line into one of the 10 categories
Before deciding whether an instruction stays or goes, you need to say what it is. Without this, the audit becomes a matter of personal taste. The taxonomy has 10 categories — and every line in your config falls into a of them. If you can’t choose a box, that’s already the diagnosis: the instruction is doing two things at once and needs to be split.
🆕 Five words before you continue
- Config: the set of files that instructs the agent —
CLAUDE.md(global and project-level), skills, hooks, andsettings.json. - Skill: a package of instructions that’s loaded only when that task comes up (or when you call it by
/nome)—unlike theCLAUDE.md— which is read in 100% of executions. - Hook: a command that the program runs automatically before or after an agent action. It isn’t text the model reads—it’s code the harness runs.
- Guardrail: a hard limit. It says what can’t happen, never. “Don’t commit to the main branch” is a guardrail.
- Instruction: here, the smallest auditable unit—usually a bullet, a sentence, or a short paragraph with just one rule.
| Category | What it is | Real example of an instruction |
|---|---|---|
| CONTEXT | A fact about your world that the model can’t guess. | "The portal is the project at ~/projetos/portal, domain inema.club, Next.js on Vercel." |
| GUARDRAIL | Hard limit. What must never happen. | "Never commit directly to main: always create a branch first.” |
| QUALITY CRITERION | Defines what a good result looks like without prescribing the steps. | "The page must work offline: no external CDN." |
| VERIFICATION | Give the model an objective way to check its own work. | "Before saying 'done,' run npm run build and paste the output." |
| INTEGRATION/TOOL | Where a credential, service, or binary lives. | "The API keys are in ~/projetos/wifi/.env; load it at runtime; never print the value.” |
| REPEATABLE PROCEDURE | A sequence that genuinely repeats and has value — a candidate for becoming a skill. | "Publish to the portal = edit the catalog, commit to the 3 repos, push." |
| MICROMANAGEMENT | Directs how the model thinks/acts instead of stating the result. | "First read the file, then list the functions, then choose one, then..." |
| REDUNDANCY | The same rule already stated in another file (or twice in the same file). | “No emoji in the output” — present in CLAUDE.md global, in the project one, and in two skills. |
| LEGACY/OBSOLETE | Fixes the weakness of a model that no longer runs here. | "Don't try to edit more than one file at a time; you'll get confused." |
| AMBIGUOUS/UNPROVEN | No one knows what it protects or what breaks without it. | "Be careful and think carefully before responding." |
💡 The six above and the four below
The first six categories describe instructions that can covering its own cost. The last four (micromanagement, redundancy, legacy, ambiguity) are diagnoses: giving a line this label already means it won’t leave the audit the way it came in.
Watch out for the easy reversal: labeling something as CONTEXT isn’t a free pass. Context that only helps with 2% of your
tasks and lives in the CLAUDE.md global still incurs a cost on 100% of executions.
🎚️ Choose among the 6 outputs
Category is a diagnosis; the decision is what you do with it. There are six possible outcomes, and none of the lines
can be left without one. An audit where 90% of the lines come out as KEEP
isn’t a healthy config—it’s an audit that never happened.
What to look at: the stack on the left goes into the funnel in its entirety—no line skips triage. Notice that the only highlighted box is TEST: it’s the default destination when you have no evidence, not a third option for putting off a decision. Also note that ten categories lead to six outcomes — the mapping isn’t one-to-one, and question 3 (topic 3) is usually what decides where the line goes.
KEEP — leave as is
Choose when: you can say what breaks without it, and you've seen that break happen.
Typical of CONTEXT, GUARDRAIL, and INTEGRATION. KEEP requires justification; if your justification is "better not touch it," the right decision is TEST.
SIMPLIFY — the intent stays; the volume disappears
Choose when: the rule is valid, but is written in three paragraphs with four examples.
It's the most common way out of bloated REPEATABLE PROCEDURES and mild MICROMANAGEMENT. The test: can you say this in one sentence a colleague would understand?
MOVE — the rule is good, the location is wrong
Choose when: the instruction only applies to one type of task, but lives in a file read on every run.
From CLAUDE.md global to the project one; from the CLAUDE.md for a skill. Zero cost when the task doesn’t come up.
MERGE — several lines, one intent
Choose when: the same idea appears in scattered pieces, sometimes with minor contradictions.
A natural outcome of REDUNDANCY. When merging, explicitly choose which version wins — otherwise you’ve just put the contradiction in one place.
TEST — remove it in one round and observe
Choose when: you thinks that it helps, but can’t cite a real failure it prevented.
It's the honest way out for AMBIGUOUS/UNPROVEN and much of LEGACY. It doesn't delete anything now: it schedules a measurement (the A/B/C plan from Track 4).
REMOVE — removed, with the cited excerpt and risk documented
Choose when: the line is an exact duplicate, fixes a model you no longer use, or directs reasoning the model already does.
Every removal comes with three things: the passage, the risk, and how to test it. A removal without all three is a guess dressed up as a method.
💡 The golden rule of this module
Without enough evidence, prefer TEST a KEEP.
No config gets bloated because someone deliberately decided to bloat it. It gets bloated because, line by line,
keeping it seemed cheaper than checking. TEST is the antidote: it costs
nothing today and turns "I think" into data next week.
❓ Ask the 7 questions per instruction
The category and decision don’t come out of nowhere: they come from seven questions asked in the same order for every instruction. Asking all seven every time is what keeps the audit from becoming a matter of impression. And one of them — the third — usually resolves the case on its own.
| # | Question | Where the answer points |
|---|---|---|
| 1 | What is it trying to prevent or ensure? | If you can’t answer: AMBÍGUA → TEST. |
| 2 | Does the current model still need it? | Written for an older model: LEGADO → TEST or REMOVE. |
| 3 | Does it say WHAT should happen, or try to direct HOW the model thinks or executes? | "How" → MICROGERENCIAMENTO → SIMPLIFY. The most discriminating one. |
| 4 | Is it duplicated in another file? | REDUNDÂNCIA → MERGE or REMOVE. |
| 5 | Unnecessarily limits autonomy? | Rules out better options without a reason → SIMPLIFY or REMOVE. |
| 6 | Is there a shorter way to preserve the intent? | Almost always yes → SIMPLIFY. |
| 7 | What breaks if it disappears? | Concrete response → KEEP. Silence or “who knows” → TEST. |
Why 3 is the most discriminating. Instructions that describe the result age well: “the build must pass” remains true with any model. Instructions that describe the internal process age poorly because they were tuned to a specific weakness in a specific generation. When a better model arrives, the instruction about “how” doesn't become neutral—it becomes a shackle, blocking the better path the model could now find on its own. Here, autonomy is exactly that: the space you leave for the model to choose a strategy after you've defined the target.
✓ Says WHAT (ages well)
- ✓"The result must run offline — no requests to external hosts."
- ✓"Before saying you're finished, run the tests and paste the output."
- ✓"Never push with the wrong author; check
git config user.emailbefore.” - ✓"Publish = commit + push. Deploy is automatic and isn't your responsibility."
✗ Directs HOW (ages poorly)
- ✗"First read the file, then list the functions, then choose one, then edit."
- ✗"Think step by step and explain your reasoning before acting."
- ✗"Don't use more than two tools per response."
- ✗“Always reread what you wrote twice before continuing.”
🧾 The 7 questions applied to a real line
EXCERPT “Before editing any file, read the entire file,
list the functions you find, choose the target function,
and only then apply the edit.”
1 prevents/ensures what? blind edits to a file the agent hasn’t read
2 still needed? no — the current model already reads before editing
3 WHAT or HOW? HOW ← determines the case
4 duplicated? yes, a shorter version already exists in the refactor skill
5 limits autonomy? yes: prohibits direct edits even when they’re obvious
6 shorter form? “don’t edit a file you haven’t read in this session”
7 what breaks? nothing observed in recent months
CATEGORY MICROMANAGEMENT
DECISION SIMPLIFY (use the form from question 6)
🛡️ Protect what the model can’t infer
There’s a class of instruction that you never cuts on reflex— no matter how attractive the cut may look in the line count: what the model has no way to figure out on its own. It can reason; it can’t guess that your domain is inema.club,
that the source of truth for prices is a specific spreadsheet, or that your company prohibits sending customer data outside the company.
✓ Context only you know—protected
- ✓Project identity: what it is, who it's for, and what the right name is.
- ✓File paths: where what lives on your disk and in the repo.
- ✓Sources of truth: which file takes precedence when two disagree.
- ✓Branding: palette, tone of voice, what never appears in the brand.
- ✓Security and compliance: what can't be removed, what needs approval.
- ✓Interface contracts: payload format, field names, versions.
- ✓Integrations: which service, which credential, which usage limit.
- ✓Internal conventions: “here we call this X,” commit convention.
✗ Generic reasoning — the model already does this
- ✗"Write readable, well-named code."
- ✗“Handle errors and edge cases.”
- ✗"Break large problems into smaller parts."
- ✗"Explain what a function does before rewriting it."
- ✗"Consider the alternatives before choosing one."
- ✗“Use general security best practices.”
- ✗“Check whether the response makes sense.”
- ✗An entire skill teaching "how to debug a problem."
🧭 The new colleague test
Imagine a capable professional who joined your team today. They know how to code, write, and research—but they don't know your setup. Every instruction you'd need to give them because it wouldn’t have had any way of knowing is protected context. Any instruction that would insult its intelligence is a candidate for removal.
Question 3 and this test are based on the same idea: “read the file before editing” insults a new colleague; “the course CSS comes
from assets/curso.css, don’t repeat inline” is exactly the kind of thing it
would be grateful to know on day one.
⚠️ The costly mistake of doing ablation poorly
The costliest failure isn’t keeping a useless line — it’s cutting irreplaceable context and only finding out three weeks later, when the agent published to the wrong repository, with the wrong author, following a convention no one documented anymore. A useless line costs context; lost context costs rework and trust. That’s why the course order is diagnosis → proposal → test, never cut on impulse.
⚖️ Optimize the right function
Here’s the part almost everyone gets wrong when they learn about ablation: the goal isn’t about reducing as much as possible. A zero-line config isn’t the goal. The goal is a ratio—and it has three things on top and one on the bottom.
What to look at: is a fraction, not a reduction target. The two red boxes below show the two ways to get it wrong—and they err in opposite directions. Notice that verifiability is in the numerator: add a verification instruction increases the result even if it costs lines, because the gain above outweighs the cost below. Ablation isn’t just subtraction.
Quality — is the result good enough?
The delivered work meets your needs without requiring you to fix it afterward. Measure it by human corrections per task, not by how it feels.
Autonomy — how much can it handle without you?
Room to choose the strategy after the target is defined. Instructions about "how" limit autonomy; exit criteria preserve it.
Verifiability — can it check its own work?
There’s a command, a test, a comparison that objectively says “passed” or “failed.” It’s the term most people forget to audit — and the only one that tends to be missing instead of being left over.
Complexity—the denominator
Lines read on every run, rules that compete with each other, accumulated exceptions, files no one fully understands. Every line of the CLAUDE.md global is charged on 100% of tasks, including those unrelated to it.
💡 Two audits that fail
- "I cut 90% and the agent got lost": the numerator dropped too. You cut irreplaceable context or the only check that existed. Rejected — even with the biggest reduction in the group.
- "I didn't change anything; everything seemed important": the denominator remains inflated and you don't have a single data point. Rejected too — just quietly, which is how sediment survives.
- Approved: fewer lines, the same quality measured on real tasks, more autonomy, and at least one objective check where there wasn’t one before.
🎯 Look for the signs in your config
Theory is over. Now you apply the taxonomy to your own config. First, the hunt list—the 13 patterns that appear in almost every accumulated configuration. Go through it with the file open: each item you recognize is a candidate whose category is almost decided.
Hunting checklist
CLAUDE.md and skillsMOVE🧪 Exercise: 15 classified lines
Objective: extract 15 real instructions from your config and assign a category and decision to each. Copy and run this in the terminal.
# 1) Extraia as instruções numeradas/bulletadas do CLAUDE.md global grep -n "^-\|^[0-9]\." ~/.claude/CLAUDE.md | head -40 # 2) Faça o mesmo no CLAUDE.md do projeto que você mais usa grep -n "^-\|^[0-9]\." ./CLAUDE.md | head -40 # 3) Tamanho de cada skill (skill grande demais é sinal do checklist) wc -l ~/.claude/skills/*/SKILL.md | sort -rn | head -15 # 4) Caça rápida a redundância: uma palavra-chave sua em toda a config grep -rn "commit\|deploy\|emoji" ~/.claude/CLAUDE.md ~/.claude/skills/ | head -20
Now fill out this table with 15 lines. One instruction per line, no grouping.
| Passage | Category | Decision | Question that decided | What breaks if it disappears |
|---|---|---|---|---|
| "No emoji in the output" | REDUNDANCY | MERGE | 4 — duplicated | nothing: it stays in the merged version |
| "First read, then list, then…" | MICROMANAGEMENT | SIMPLIFY | 3 — directs HOW | nothing observed |
"API keys in ~/projetos/wifi/.env" | INTEGRATION | KEEP | 7 — breaks immediately | the agent asks the user for the key |
| "Be careful and think carefully" | AMBIGUOUS | TEST | 1 — no clear intent | unknown—that’s why you test |
| … | … | … | … | … |
Exit criterion: 15 completed lines,
all with a category e decision, and at least one in TEST.
Why the TEST requirement:
if all 15 came out as KEEP, the most likely explanation isn’t that your config is perfect—
it’s that you classified it based on comfort. Go back to questions 1 and 7: if you can’t say what the line prevents or what breaks without it, it isn’t KEEP.
Quick check (doesn't block anything): you find a line that says "always reread the file twice before editing." You can't recall any failure it has prevented. What's the most defensible category/decision pair?
Key concepts
Can't choose? It's doing two things
WHAT ages well, HOW it ages poorly
Without evidence, schedule the measurement
The only item you can resolve by adding
📌 Module Summary
Next Module:
2.2 — From micromanagement to criteria and verification: how to rewrite "do A, then B, then C" as an objective, guardrails, an exit criterion, and a real way to check.