🔍 How to diagnose
How do you go line by line and decide what stays, what goes, and what needs testing? This track gives you the taxonomy—10 categories, 6 decisions, 7 questions—and the move that most improves quality: replace recipes with criteria and verification.
What to look at: the funnel doesn't have two outcomes ("stays" / "goes") — it has six. The highlighted box isn't the decision: it's the 7 questions— because the question decides, not taste. And the output highlighted in cyan is TEST: when you can’t answer “what breaks if it disappears?”, the line’s destination is the ablation test, never KEEP for comfort — KEEP without an answer is just fear dressed up as a criterion.
Track map
Detailed content
🔍 The taxonomy: what each line is
Classify each instruction into a category and a decision, based on 7 questions. When in doubt, TEST — never KEEP for comfort.
Every instruction in your config falls into one of ten boxes: CONTEXT, GUARDRAIL, QUALITY CRITERION, VERIFICATION, INTEGRATION/TOOL, REPEATABLE PROCEDURE, MICROMANAGEMENT, REDUNDANCY, LEGACY/OBSOLETE, and AMBIGUOUS/UNPROVEN. "Deployments go through git" is CONTEXT; "never force push to main" is a GUARDRAIL; "think step by step" is MICROMANAGEMENT.
Without a name, every line seems equally important — and you end up defending everything. The category is what separates what the model there’s no way to know (context, integrations) for what it already does better on its own (micromanagement). This is the step that makes the audit debatable with someone else.
Category describes what the line é, not what to do with it. The last four (micromanagement, redundancy, legacy, ambiguity) are the natural suspects; the first six carry the value. A line can fit two categories—that’s already a sign it needs to be split in two.
After the category comes the decision. KEEP = leave as is. SIMPLIFY = keep it, but in half the words. MOVE = it’s in the wrong place (a project rule in the global config, or vice versa). MERGE = two lines say the same thing. TEST = you don’t know, so prove it through ablation. REMOVE = you know it’s dead weight.
The audit fails when there are only two buttons, “stays” and “goes” — then everything becomes “stays.” Six outcomes allow for an honest move: most lines don’t need to disappear; they need to be shortened, moved to another file, or merged with the one next to them.
When in doubt, TEST — never KEEP. An audit with no lines in TEST almost always means the author is protecting their own config. TEST doesn’t delete anything: it only schedules the proof.
What does this line prevent or ensure? Does the current model still need it? Does it say what should happen or directs how Does the model stumble over it? Is it duplicated somewhere else? Does it limit autonomy for no reason? Is there a shorter version? What breaks if it disappears?
They turn an opinion (“I think this helps”) into an auditable record. The most discriminating is the third: an instruction that says what usually survives; an instruction that directs how to think is the one that aged along with the old model.
Always write down the question you decided and the answer to “what breaks if it’s removed?” If the answer is “I don’t know,” the decision is already made: TEST. This pair—question + consequence—is what Track 4 will use as the hypothesis for the A/B/C test.
Project identity, paths and sources of truth, branding, security, compliance, interface contracts, integrations, and internal conventions. None of this is in the model's weights: it can't guess that your deploys happen through git, that the key is in a certain `.env` file, or that the commit's destination account varies by repository.
Ablation turns into damage when someone cuts things out on a whim. This list is the brake: these are the lines whose absence doesn’t gradually degrade quality—they break things all at once, sometimes in production.
Quick test: if the information is specific to your world and not inferable from the repository, it's KEEP (SIMPLIFY at most). "Don't cut" doesn't mean "don't shorten"—true context can also fit in one line.
The audit’s objective function is to maximize quality + autonomy + verifiability ÷ complexity. Line count appears only in the denominator—and never by itself.
Anyone who focuses only on the denominator produces a beautiful config that’s worse: it cuts verification (which was cheap and valuable) along with micromanagement. An audit can legitimately end with the config larger— if what you added was a verifiable exit criterion.
Adding checks raises the numerator. Removing step-by-step instructions lowers the denominator e increases autonomy. They’re the two highest-return moves—and that’s why module 2.2 exists.
Patterns that appear in almost every mature configuration: rules duplicated between `CLAUDE.md` and skills, oversized skills, too many examples, rigid formatting without a reason, accumulated exceptions ("except when…" repeated three times), contradictions between files, global context that only serves two tasks — and, the most costly of all, missing checks.
They’re reading shortcuts: instead of rereading everything with the same attention, you scan the config for these five or six signs and quickly find the candidates. Duplication and contradiction are the ones that quietly sabotage things the most—the model picks one of the two versions, and you never know which.
The costliest signal is the one that isn’t there. Accumulated exceptions are fossils from an old model; contradiction is a silent bug; missing verification is the instruction that would have let everything else work without a babysitter.
🎯 From micromanagement to criteria and verification
Stop scripting "do A, then B." Write the objective, guardrails, exit criteria — and a real way for the model to check its own work.
Tell the model every step along the way instead of telling it where to end up. "Open the file, find the function, edit line 12, run the build, then…" — a script that only describes one path, what you would have followed.
This is the number one failure mode—and, oddly, it's more common among people with years of engineering experience, because specifying things well has always been a virtue. With older models, the script helped; with current ones, it blocks a better route the model would find on its own.
The right level of instruction is what you’d give a capable colleague who just joined: context and criteria, not step-by-step directions. If your instruction wouldn’t hold up to a “why?”, it’s a script.
Replace "do A, then B, then C" with "Produce X. Follow Y. The result must achieve Z. Verify using W. Choose the strategy." Same intent, different form: what was a sequence becomes a target plus a constraint plus proof.
It's the rewrite that usually shrinks the instruction and improves the result at the same time—because it gives the model a choice of approach, but not of standards. You stay in control of what matters: the Y and the Z.
If, when converting, you can’t write Z ("the result must achieve…"), the prompt was never the problem: you haven’t defined what you want yet. And without W, the instruction once again depends on you checking each output.
Objective (what to produce) · Context (what the model won't infer) · Guardrails (what never to do) · Quality criteria (what good looks like) · Verification (how to check) · Autonomy (what it can decide on its own and when to stop and ask).
Filling in all six fields exposes the gap right away. Almost every bad instruction you have today is strong on Objective, bloated with step-by-step details, and empty under Verification and Autonomy—the two fields that determine whether you’ll need to stay nearby.
Each field maps directly to the taxonomy in module 2.1: Context = CONTEXT, Guardrails = GUARDRAIL, Criteria = QUALITY CRITERION, Verification = VERIFICATION. Anything that doesn't fit into one of the six fields is very likely micromanagement.
Give the model a real way to check its own work. Cherny’s example: rewrite an Electron app in Swift, run the original in a VM, take screenshots of both, and compare them pixel by pixel, without stopping until they match. Produce → observe → compare → stop.
A short prompt with verification beats a giant prompt without verification, and it’s not even close. Without the comparison step, the model has no way to know it made a mistake—and you become the verifier, manually, forever.
Good verification is executable by a third party: a command, a test, a comparison, a file check. “Review carefully” isn’t verification — it’s wishful thinking. And every verification needs an explicit stop condition.
Give the model a task that’s a little harder than feels comfortable — and keep a list of requests that failed, so you can run them again with each new model.
Most of the defensive instructions in your config exist because of a failure that no longer happens today. Without retesting, you’ll never find out—and you’ll keep paying the complexity cost for a problem that was solved two models ago.
This is empirical science, not theoretical: there’s no secret trick, just a difficult task → a way to verify → observe where it gets stuck → fix it that → repeat. The list of past failures is your personal benchmark, and it’s what drives the cycle in Track 4.
Choose the most "by-the-book" instruction in your config or one of your skills and write two versions next to the original: (a) 50% shorter, preserving the intent; (b) minimal, using the 6-field template, with an objective verification step.
Reading about conversion changes nothing; doing it once does. The two versions side by side reveal how much of the original text was process and how much was criteria — and the answer is usually surprising.
Exercise exit criterion: the minimum version has a check that someone else could execute without asking you anything. If it needs you to explain, it isn’t verification yet. Keep all three versions — they become the A/B/C arms of the ablation test.