PTENES
Skip to content
PATH 4

📊 Prove and maintain

Cutting without measuring is a guess dressed up as a decision. This path completes the method: you build three versions of the same configuration, run the same real tasks on all three, compare them across nine dimensions — and turn the result into a routine that keeps the sediment from coming back with the next model.

each observed failure becomes an eval—and the cycle starts again 5 a 10 real tasks A · current config B · half C · minimal verdict 9 dimensions compared same task · same model · three configurations tie = the extra instructions were dead weight

What to look at: the only thing that changes across the three columns is the configuration—the task and model are identical, so the difference in output can only come from the instructions. The highlighted box is the verdict: if A and C tie, what A has and C doesn’t is proven dead weight. And the cyan return path is what separates a cleanup from a method: the result doesn’t end with the report; it becomes a test case and feeds into the next round.

2
Modules
12
Topics
~1h35
Duration
Advanced
Level
Track progress
0% 0 of 0

Track map

Detailed content

4.1~50 min

📊 The A/B/C ablation plan

Three config versions, 5 to 10 real tasks, and nine comparison dimensions. Without testing, cutting is guesswork.

0% 0 of 6
What it is:

A is your current, untouched config. B keeps approximately half the instructions—the ones the audit classified as most defensible. C is the minimal version: project context, objective, nonnegotiable guardrails, quality criteria, and verification. Nothing else.

Why learn:

Two versions only answer “did it get better or worse?” Three answer “where is the curve’s knee?”—if B and C tie with A, you can cut further than you thought; if C drops and B doesn’t, the limit is between them.

Key concepts:

Version all three in git before you start. C isn't the empty config: it's the config that contains only what no model can guess on its own — your context and your rules.

What it is:

The test task set comes from your real work history: the bug fix you made last week, the tedious refactor, the page that took three rounds of back-and-forth. Five is the minimum to avoid deciding by chance; ten is tiring, and you stop taking notes properly.

Why learn:

A toy task (“write a function that adds two numbers”) passes with any configuration. It doesn’t distinguish A from C—it only creates the false confidence that you “tested.” A test is only valuable if the tasks can fail.

Key concepts:

Cover the kinds of work you actually do, not just the most common one. Include at least one task where the current config has already caused a problem — that’s where the difference shows up.

What it is:

Nine columns in your table: result quality, adherence to your rules, autonomy (how many times it had to ask), necessary human corrections, consistency across runs, time to delivery, tool use, complexity of the proposed solution, and self-verification (did it check its own work?).

Why learn:

A single score hides tradeoffs. It’s common for version C to deliver the same quality in less time and with a simpler solution—and that only becomes visible when time and complexity are separate columns.

Key concepts:

Human corrections are the most honest dimension: they count the work left for you. Solution complexity usually gets worse with too many instructions, not better.

What it is:

Same task, same model, same repository state across all three versions. Before running anything, write down what a “good” answer to that task would look like. Only then read the outputs.

Why learn:

If you define the criterion after seeing the answer, the criterion bends to fit the answer — you approve what showed up instead of evaluating what needed to show up. That’s the bias that ruins most homegrown prompt tests.

Key concepts:

Switching models mid-run invalidates the entire comparison. If you can, review the outputs without knowing which version produced them—and note it right away, not from memory at the end of the day.

What it is:

A tie between A and C means everything in A that isn’t in C is dead weight: it costs context and attention without adding anything. A one-off decline on a task is noise. A repeated decline in the same dimension across different tasks triggers the reintroduction rule.

Why learn:

The reintroduction rule prevents two opposite mistakes: putting everything back after the first scare, or sticking with the minimal version while ignoring a real failure. Restore one instruction at a time, as specifically as possible, citing the failure that justifies it.

Key concepts:

An instruction returned without a recorded failure is faith, not evidence. And the version that comes back isn't the old one: it's the old one minus everything the test has proven unnecessary.

What it is:

Every failure you observe in a test becomes a saved case: the task, what went wrong, and what the correct result would be. This collection is your eval suite—and it grows from real work, not from a generic catalog.

Why learn:

This is how the test stops being a one-time event. The next time you switch models, you don't start from scratch: run the battery and see, in minutes, what the new model can already handle on its own.

Key concepts:

A saturated eval—that always passes, in every version—gets retired. It's become dead weight, just like the instruction that led to it; keeping everything forever recreates the problem in another file.

View Full
4.2~45 min

🔁 The continuous cycle

The 5-step cycle, re-audit triggers, the empirical mindset — and your personal ablation policy.

0% 0 of 6
What it is:

Audit the configuration; cut what the diagnosis marked as removable; use the trimmed version for real work for a few days; note any failures that come up; repeat. Five steps, no extra step.

Why learn:

The step almost everyone skips is the third one. Cutting and rereading the file proves nothing—only using it in real work reveals whether anything was missing, and it reveals that in days, not minutes.

Key concepts:

The cycle is short on purpose. A small, frequent iteration catches problems while they’re still cheap to undo.

What it is:

Three objective signals: a new model came out; your CLAUDE.md exceeded approximately 150 lines; two or more skills began competing for the same activation trigger.

Why learn:

Without an explicit trigger, the re-audit happens when the frustration builds up — in other words, too late. A new model is the strongest trigger: half your instructions were written to fix weaknesses it may no longer have.

Key concepts:

150 lines isn't a rule; it's an alarm. Skills competing for the same trigger aren't a description problem: it's a sign that you have too many skills for the same job.

What it is:

Treat your own configuration as empirical science: you don’t know what the model needs, so you test. You observe what happens. And you retest what failed before, because yesterday’s failure may have been resolved by today’s model.

Why learn:

Instructions almost always start with a specific frustration, become permanent, and are never checked again. Most of the dead weight in a CLAUDE.md was useful — two models ago.

Key concepts:

Writing more instructions is the easy reflex, and almost always the wrong one. The right question is "is this still necessary?", and the only way to answer it is to run it again.

What it is:

There are three areas where the minimal version still tends to struggle: deep, tightly coupled systems; distributed architectures with state spread around; and fine visual verification (a pixel out of place, poor contrast).

Why learn:

Knowing where the method struggles prevents two mistakes: cutting context that is genuinely necessary in these cases, and concluding that "ablation doesn't work" because you tested it on the hardest terrain.

Key concepts:

In these cases, what stays isn't micromanagement: it's factual context the model can't discover on its own—and verification criteria, not step-by-step instructions.

What it is:

A Routine is a scheduled task whose content is a single prompt. At Anthropic, teams run 20 to 30 of them a day — issue triage, build checks, summaries — and each fits in one sentence.

Why learn:

It's the practical proof of the entire course: real, recurring work running with a single instruction. If this is the professional standard, yours CLAUDE.md of 400 lines needs to justify every one of them.

Key concepts:

A good candidate for a routine: a repetitive task with clear success criteria and a low cost of failure. The periodic re-audit itself could become one.

What it is:

Ten lines from you answering: when I re-audit, what I never cut, how I decide to reintroduce something, what my test task set is, and when an eval is retired.

Why learn:

Without a written policy, every re-audit starts the discussion from scratch, and decisions vary with the mood of the day. With one, the criterion comes before the specific case — exactly what prevents bias toward approving what’s already there.

Key concepts:

Save it where you'll review it—a dedicated file, a card, or the top of the release checklist. Never as one more paragraph inside the CLAUDE.md: it would mean becoming exactly the kind of sediment the course taught you to cut.

View Full

🎓 Final project

The course ends with four deliverables based on your actual configuration — the same one you used throughout all four tracks. This isn’t a paper exercise: it’s material you’ll reuse the next time you switch models.

Deliverable 1

Skill report

The complete output of audit-ablacao run on your config, with all ten sections filled out — inventory, line-by-line classification, redundancies, micromanagement, and a minimal-version proposal.

Deliverable 2

CLAUDE.md before and after

The two versions and the diff between them. The diff makes every cut auditable later—including by you, three months from now, when you don't remember why that line disappeared.

Deliverable 3

A/B/C table

The test results from at least two real tasks, with all nine dimensions filled in for all three versions and the verdict written in one sentence.

Deliverable 4

Personal ablation policy

Your ten lines of criteria, saved outside the CLAUDE.md, somewhere you’ll actually reread when the next model comes out.

Approval criterion

  • 1.All four deliverables are required — three aren't enough.
  • 2.Every applied removal has a cited passage, an identified risk, and a way to test it.
  • 3.Every returned instruction has a recorded recurring failure to justify it.
  • 4.The final version has at least one objective check where there wasn’t one before.