📊 Prove and maintain
Cutting without measuring is a guess dressed up as a decision. This path completes the method: you build three versions of the same configuration, run the same real tasks on all three, compare them across nine dimensions — and turn the result into a routine that keeps the sediment from coming back with the next model.
What to look at: the only thing that changes across the three columns is the configuration—the task and model are identical, so the difference in output can only come from the instructions. The highlighted box is the verdict: if A and C tie, what A has and C doesn’t is proven dead weight. And the cyan return path is what separates a cleanup from a method: the result doesn’t end with the report; it becomes a test case and feeds into the next round.
Track map
Detailed content
📊 The A/B/C ablation plan
Three config versions, 5 to 10 real tasks, and nine comparison dimensions. Without testing, cutting is guesswork.
A is your current, untouched config. B keeps approximately half the instructions—the ones the audit classified as most defensible. C is the minimal version: project context, objective, nonnegotiable guardrails, quality criteria, and verification. Nothing else.
Two versions only answer “did it get better or worse?” Three answer “where is the curve’s knee?”—if B and C tie with A, you can cut further than you thought; if C drops and B doesn’t, the limit is between them.
Version all three in git before you start. C isn't the empty config: it's the config that contains only what no model can guess on its own — your context and your rules.
The test task set comes from your real work history: the bug fix you made last week, the tedious refactor, the page that took three rounds of back-and-forth. Five is the minimum to avoid deciding by chance; ten is tiring, and you stop taking notes properly.
A toy task (“write a function that adds two numbers”) passes with any configuration. It doesn’t distinguish A from C—it only creates the false confidence that you “tested.” A test is only valuable if the tasks can fail.
Cover the kinds of work you actually do, not just the most common one. Include at least one task where the current config has already caused a problem — that’s where the difference shows up.
Nine columns in your table: result quality, adherence to your rules, autonomy (how many times it had to ask), necessary human corrections, consistency across runs, time to delivery, tool use, complexity of the proposed solution, and self-verification (did it check its own work?).
A single score hides tradeoffs. It’s common for version C to deliver the same quality in less time and with a simpler solution—and that only becomes visible when time and complexity are separate columns.
Human corrections are the most honest dimension: they count the work left for you. Solution complexity usually gets worse with too many instructions, not better.
Same task, same model, same repository state across all three versions. Before running anything, write down what a “good” answer to that task would look like. Only then read the outputs.
If you define the criterion after seeing the answer, the criterion bends to fit the answer — you approve what showed up instead of evaluating what needed to show up. That’s the bias that ruins most homegrown prompt tests.
Switching models mid-run invalidates the entire comparison. If you can, review the outputs without knowing which version produced them—and note it right away, not from memory at the end of the day.
A tie between A and C means everything in A that isn’t in C is dead weight: it costs context and attention without adding anything. A one-off decline on a task is noise. A repeated decline in the same dimension across different tasks triggers the reintroduction rule.
The reintroduction rule prevents two opposite mistakes: putting everything back after the first scare, or sticking with the minimal version while ignoring a real failure. Restore one instruction at a time, as specifically as possible, citing the failure that justifies it.
An instruction returned without a recorded failure is faith, not evidence. And the version that comes back isn't the old one: it's the old one minus everything the test has proven unnecessary.
Every failure you observe in a test becomes a saved case: the task, what went wrong, and what the correct result would be. This collection is your eval suite—and it grows from real work, not from a generic catalog.
This is how the test stops being a one-time event. The next time you switch models, you don't start from scratch: run the battery and see, in minutes, what the new model can already handle on its own.
A saturated eval—that always passes, in every version—gets retired. It's become dead weight, just like the instruction that led to it; keeping everything forever recreates the problem in another file.
🔁 The continuous cycle
The 5-step cycle, re-audit triggers, the empirical mindset — and your personal ablation policy.
Audit the configuration; cut what the diagnosis marked as removable; use the trimmed version for real work for a few days; note any failures that come up; repeat. Five steps, no extra step.
The step almost everyone skips is the third one. Cutting and rereading the file proves nothing—only using it in real work reveals whether anything was missing, and it reveals that in days, not minutes.
The cycle is short on purpose. A small, frequent iteration catches problems while they’re still cheap to undo.
Three objective signals: a new model came out; your CLAUDE.md exceeded approximately 150 lines; two or more skills began competing for the same activation trigger.
Without an explicit trigger, the re-audit happens when the frustration builds up — in other words, too late. A new model is the strongest trigger: half your instructions were written to fix weaknesses it may no longer have.
150 lines isn't a rule; it's an alarm. Skills competing for the same trigger aren't a description problem: it's a sign that you have too many skills for the same job.
Treat your own configuration as empirical science: you don’t know what the model needs, so you test. You observe what happens. And you retest what failed before, because yesterday’s failure may have been resolved by today’s model.
Instructions almost always start with a specific frustration, become permanent, and are never checked again. Most of the dead weight in a CLAUDE.md was useful — two models ago.
Writing more instructions is the easy reflex, and almost always the wrong one. The right question is "is this still necessary?", and the only way to answer it is to run it again.
There are three areas where the minimal version still tends to struggle: deep, tightly coupled systems; distributed architectures with state spread around; and fine visual verification (a pixel out of place, poor contrast).
Knowing where the method struggles prevents two mistakes: cutting context that is genuinely necessary in these cases, and concluding that "ablation doesn't work" because you tested it on the hardest terrain.
In these cases, what stays isn't micromanagement: it's factual context the model can't discover on its own—and verification criteria, not step-by-step instructions.
A Routine is a scheduled task whose content is a single prompt. At Anthropic, teams run 20 to 30 of them a day — issue triage, build checks, summaries — and each fits in one sentence.
It's the practical proof of the entire course: real, recurring work running with a single instruction. If this is the professional standard, yours CLAUDE.md of 400 lines needs to justify every one of them.
A good candidate for a routine: a repetitive task with clear success criteria and a low cost of failure. The periodic re-audit itself could become one.
Ten lines from you answering: when I re-audit, what I never cut, how I decide to reintroduce something, what my test task set is, and when an eval is retired.
Without a written policy, every re-audit starts the discussion from scratch, and decisions vary with the mood of the day. With one, the criterion comes before the specific case — exactly what prevents bias toward approving what’s already there.
Save it where you'll review it—a dedicated file, a card, or the top of the release checklist. Never as one more paragraph inside the CLAUDE.md: it would mean becoming exactly the kind of sediment the course taught you to cut.
🎓 Final project
The course ends with four deliverables based on your actual configuration — the same one you used throughout all four tracks. This isn’t a paper exercise: it’s material you’ll reuse the next time you switch models.
Skill report
The complete output of audit-ablacao run on your config, with all ten sections filled out — inventory, line-by-line classification, redundancies, micromanagement, and a minimal-version proposal.
CLAUDE.md before and after
The two versions and the diff between them. The diff makes every cut auditable later—including by you, three months from now, when you don't remember why that line disappeared.
A/B/C table
The test results from at least two real tasks, with all nine dimensions filled in for all three versions and the verdict written in one sentence.
Personal ablation policy
Your ten lines of criteria, saved outside the CLAUDE.md, somewhere you’ll actually reread when the next model comes out.
Approval criterion
- 1.All four deliverables are required — three aren't enough.
- 2.Every applied removal has a cited passage, an identified risk, and a way to test it.
- 3.Every returned instruction has a recorded recurring failure to justify it.
- 4.The final version has at least one objective check where there wasn’t one before.