🔄 Run the 5-step cycle
Everything you learned in the first three tracks fits into a five-step cycle that repeats. It doesn’t end: each round produces a smaller configuration than the last, and each new model restarts the cycle. The point of the module isn’t to do the audit once — is setting up the mechanism that gets you to do it again six months from now, without relying on motivation or a crisis.
What to look at: the dotted cyan line is the only edge that closes the loop—without it, you’ve done a cleanup, not installed a cycle. And notice the emphasis on node 4: that’s where most people fail, because returning an instruction after the first failure seems careful, but in practice it’s how the configuration starts to bloat again. The entire discipline of the course lives in this box.
Run /audit-ablacao in the selected scope
Global, one project, or a set of skills. The skill only diagnoses: it reads, classifies, and proposes—it doesn't edit, move, or delete. The deliverable is a 10-section report.
Cut the Top 10 by impact ÷ risk — in a separate session
The skill doesn’t apply anything. Applying changes is your decision, in a new request, with the config under git so you can roll back. Start with redundancies and conflicts (almost zero risk), then legacy instructions, then micromanagement.
Use it on real work for a few days
Not in a hypothetical test. An invented test exercises what you imagine matters; real work exercises what actually matters—and that's where the failure appears.
Document the failures and restore only when they repeat
This is the step almost everyone skips. Record the failure; only add an instruction back when the same class of failure if it repeats — in the shortest form possible. A single failure is noise; a repeated failure is a signal.
Repeat every ~6 months or with each major model release
What survived the last pass may not survive the next: the instruction that corrected a real weakness becomes dead weight when that weakness no longer exists.
💡 Why step 4 gets skipped the most
Because one painful mistake already seems like enough justification. The model makes a mistake on a task, you write the rule right away — and you’ve just traded a one-off mistake for a permanent cost: that line will be read in 100% of future runs, including ones that have nothing to do with it. Taking notes and waiting for it to happen again costs three days of patience and saves years of burned context.
Key concepts
Return, not a one-time event
Audit ≠ apply
A hypothetical test doesn’t count
Only trigger for reintroduction
🚨 Recognize the re-audit triggers
The six-month schedule is the floor, not the ceiling. There are signs that say “audit now, don’t wait for the date” — and almost all of them are easy to observe: a number that crossed a threshold, behavior that became unpredictable, or something someone said to you. The table below connects each trigger to the scope that it asks for: not every signal calls for auditing everything.
| Observed trigger | Scope to audit | Why |
|---|---|---|
| New model released | Everything: global + projects + skills | Every instruction written to fix the previous model’s weakness became a candidate for dead weight. |
CLAUDE.md exceeded ~150 lines | That file only | Beyond that, it’s almost always procedure disguised as a global rule—material for MOVE for a skill. |
| 3+ skills competing for the same trigger | The set of skills, not the CLAUDE.md | The trigger became a lottery. The solution is usually MERGE, a narrower scope or direct invocation. |
| "No one knows what this rule is protecting anymore" | The block that contains the rule | A rule no one dares delete is a rule with no owner. It goes to TEST, never for KEEP out of fear. |
| You explained the same rule to someone else twice | The rule and its neighbors | If it needs a human explanation to be understood, it’s poorly written — or shouldn’t exist. |
📌 The most underestimated trigger
The last one in the table. When you explain the same rule to a colleague for the second time, it has failed as text—and if it fails with a human who can ask you a follow-up question, it fails even more with the model, which won’t ask. This trigger doesn’t show up in any metric; it shows up in what you say. Pay attention to it.
✓ Re-audit now
- ✓A new model came out, and your config is the same as it was two models ago
- ✓You've added rules in recent months and never removed any
- ✓Two rules in your config contradict each other, and you don’t know which one takes precedence
✗ Not a re-audit trigger
- ✗A task went wrong today—that calls for diagnosis, not a general cleanup
- ✗You read a post saying that "a large CLAUDE.md is bad" — measure yours, not the average
- ✗Wanting to tinker on a Friday with no real work to test it afterward
🧪 Unlearn overengineering
Boris keeps coming back to the same topic: working with a model has become empirical, not theoretical. You don’t infer what the model needs from principles — you test, observe where it struggles, and adjust based on what saw. This means letting go of two kinds of baggage: what you learned about how earlier models behaved, and assumptions from computer science theory about what “should” be difficult.
The third requirement is the most uncomfortable: stay willing to retest ideas that failed in the past, because the reason they failed may have disappeared. Boris describes overengineering as a failure mode recurring — and unlearning this habit as a real journey for experienced builders. The more years of engineering experience you have, the stronger the reflex to specify everything, and the more work it takes to undo it.
🆕 Two words before moving on
- Overengineering: overengineering—solving a problem that called for less with too much structure, too many rules, and too many steps. In an agent config, it appears as an instruction that scripts every step instead of giving an objective and criteria.
- Eval: abbreviation for evaluation — a test case for your agent: a task with an expected result that you rerun with each model version to see whether it improved or got worse. Evals age too: retire the ones the model always aces.
✗ Old mindset (theoretical)
- ✗"This is hard for models" — a deduction based on what you saw two models ago
- ✗"In theory, this problem is intractable, so I need to decompose it by hand"
- ✗"I tried this once and it didn't work" — archived forever
- ✗Writing the rule before to see the failure, just in case
- ✗Measure prompt quality by its length and level of detail
✓ Empirical mindset
- ✓“Let’s see” — run the task and observe where the model actually stumbles
- ✓Give the whole task first; break it down only if observation calls for it
- ✓Scheduled retesting of past failures with each new version
- ✓Write the rule afterward from the second occurrence of the same failure
- ✓Measure prompt quality by the verified result
🎯 The detail that makes a difference: retest what failed
The list of “things AI can’t get right” that you carry around in your head was built using models that no longer exist. Every item on it is a disproven hypothesis, not a fact. Pick one item from that list with each new model and retest it—it’s the cheapest way to discover new capabilities, and the only one that doesn’t depend on someone telling you.
🧭 Calibrate what’s still unresolved
Boris has said publicly that programming is solved — and added the caveat that tends to disappear from the quote: solved for the kind of programming it does, not for everyone. The caveat matters to you for a practical reason: in areas where the model still struggles, the remedy isn’t to write a longer instruction. It’s to tighten verification.
| Area that’s still difficult | How this looks in practice | Verification that pays off |
|---|---|---|
| Very deep system codebases | A plausible change that violates an invariant buried in another layer. | Integration test that exercises the layer below, not just the one you changed. |
| Distributed systems | Works on the machine, fails under concurrency, message ordering, or network partition. | Run under load and with injected failures; check the final state, not just the return value. |
| Detailed visual verification | Something shifted by one pixel passes as “the same.” Opus 5 made a major leap in vision and computer use, but it still isn’t perfect. | Automated image comparison with a threshold, a reference screenshot, and a human review at the end. |
⚠️ Warning: the trap in this section
Read “this is still difficult here” and respond with three more instruction paragraphs in the CLAUDE.md is exactly the reflex the entire course tried to dismantle. A long instruction can't fix a capability limit; it just burns context while the limit remains. In these areas: stronger verification, narrower scope, human review at the right point. Never longer prose.
🔭 Calibration isn't giving up
This list is a snapshot of today, and the rule from section 3 applies here too: retest. Visual verification, which used to be a dead end, has advanced a lot from one model to the next. Treat every item as "not yet" and set a review date—not as "never."
⚙️ Automate with one-sentence routines (bonus)
This section is a bonus: it solves a different problem from the rest of the course. Loops and routines aren't about a large task divided into parts—they're about a a repetitive task run on a schedule. A loop is essentially a cron job running Claude locally. A routine is the same thing in the cloud, so you can close your laptop.
🆕 Four words before you continue
- Cron job: a scheduled task that the system runs on its own at a fixed time—“every day at 3 a.m.,” “every hour.” It comes from
cron, the classic Unix scheduler. - Routine: the same schedule, but running in the cloud — it doesn't depend on your computer being on.
- PR (pull request): a proposed code change opened in a repository, with the diff visible, for someone to review and approve before it goes in. It’s how the routine delivers work without applying anything on its own.
- Scaffolding: scaffolding—temporary code that exists only to support something during a phase (for example, toggling an experiment on and off) and becomes junk when the phase ends.
Two properties matter: each run doesn’t share context with the previous one, though it can share memory; and scheduling can be every five minutes, every hour, or daily. Anthropic currently runs 20 to 30 daily routines on its own products—and the detail that changes everything: each one is a single prompt, often one sentence. The model figures out the implementation details on its own.
What to look at: notice the size of the boxes in the middle — it’s the entire prompt for each routine, one sentence. Nobody wrote "use static and dynamic analysis" in the dead code routine; the model chose those methods on its own. And look at the label on the right: the deliverable is a PR, not an applied change. It’s the same discipline as the audit skill — propose and leave the decision to a human.
The five routines Anthropic runs
- Clean up dead code. Runs daily, uses static and dynamic analysis, and opens a PR removing what it finds. No one explicitly requested these methods — the model figured them out on its own.
- Publish completed experiments. Finds experiments already released to 100% of users, removes the experiment scaffolding, and publishes the result.
- Write missing tests in areas of the codebase with low coverage.
- Delete useless tests — including low-value tests added by older models or by people over time.
- Abstraction police. Finds nearly duplicated abstractions that diverged over time and merges them back into one.
The pattern worth copying: the highest-leverage routines are obvious, repetitive maintenance tasks that are easy to put off. Auditing your own config is exactly this kind of task.
🧪 Copy and run: your re-audit as a periodic reminder
Objective: turn step 5 of the cycle into something that happens without relying on your memory. Two options—the one-line version, which works on any machine, and the routine version, for those who already use cloud scheduling.
# --- Opção A: cron job local (roda no seu computador) --- # abre o editor de agendamentos: crontab -e # cola esta linha: dia 1, a cada 6 meses (janeiro e julho), 9h 0 9 1 1,7 * echo "ABLACAO: rodar /audit-ablacao na config global" >> ~/ablacao-lembretes.txt # --- Opção B: routine (roda na nuvem, notebook fechado) --- # o prompt inteiro da routine é UMA frase: "Rode uma auditoria de ablação em ~/.claude/CLAUDE.md e ~/.claude/skills/, e me entregue o Top 10 por impacto ÷ risco. Não altere nenhum arquivo." # --- Opção C: sem nada instalado --- # evento recorrente semestral no calendário, com este título: "Ablação: rodar /audit-ablacao + abrir ablacao-diario.md"
How to verify:
- 1. Run
crontab -land confirm that the line appears in the listing. - 2. Temporarily change the schedule to 2 minutes from now and check that
~/ablacao-lembretes.txtgot one line. Then return to the semiannual review. - 3. If you used option B, confirm that the routine’s sentence explicitly prohibits changing files — an audit that starts making changes is one you can’t trust.
Now replace it with yours: replace the paths with
<a config que você realmente audita> and the month by
<seus dois meses do ano>. If you have more than one active project, make one line per scope — auditing everything together turns into a report nobody reads.
📎 Why “one sentence” and not a manual
The routine is the final test of the course’s thesis: if a one-sentence prompt, with no step-by-step instructions, produces useful PRs daily in a codebase the size of Claude Code’s, then your 12-step procedure wasn’t ensuring quality—it was just making you feel in control.
📝 Write your ablation policy
Final module and course exercise: write your ablation policy personal in ≤10 lines. It needs four things: when to re-audit, what you never cut, the reintroduction rule, and where the logs and evals live. Less than that isn’t a policy; more than that won’t get read.
And it will where it will be reread — as a separate skill or note, with a date set,
not as another paragraph in the CLAUDE.md. If you paste the policy there, you’ve just
created zombie line number 1 for the next audit: text about ablation being loaded in 100% of runs,
including those that have nothing to do with auditing configuration.
⚠️ The irony that closes the course
Finish eight modules on not inflating the CLAUDE.md and celebrating by adding ten lines to it would be a perfect ending—for a comedy. A policy is a procedure with a known trigger, so it belongs in a skill (or a note with a reminder). You already know the rule: what's always true stays in the CLAUDE.md; what’s true only when the task is that specific becomes a skill.
📋 Copy and fill in: policy template (≤10 lines)
Objective: save as ~/.claude/skills/politica-ablacao/SKILL.md (or as a pinned note, if you prefer). Fill in everything between < > with your real-world case.
## Política de ablação — <seu nome> 1. Re-auditar: a cada 6 meses (próxima: <DD/MM/AAAA>) e a cada modelo novo. 2. Gatilho extra: CLAUDE.md > 150 linhas, 3+ skills no mesmo gatilho, regra sem dono. 3. Escopo padrão: <global | projeto X | conjunto de skills Y>. 4. Nunca cortar sem teste: identidade, caminhos/fontes de verdade, segurança, compliance, contratos de interface, integrações, convenções internas. 5. Reintrodução: só após a MESMA falha repetir 2x em trabalho real, e na forma mais curta possível. 6. Aplicar cortes sempre em sessão separada da auditoria, com a config sob git. 7. Diário de falhas: <caminho/ablacao-diario.md>. 8. Evals pessoais: <caminho/evals/> — aposentar as saturadas a cada re-auditoria. 9. Toda remoção aplicada registra: trecho citado + risco + como testar. 10. Toda versão nova precisa ter pelo menos uma verificação objetiva.
How to verify:
- 1. None remain
< >in the file—if it’s still there, the policy is still a template, not yours. - 2. Line 1 has a concrete date, not “in a few months.”
- 3. The file is NOT inside the
CLAUDE.md. Rungrep -c "Política de ablação" ~/.claude/CLAUDE.md— needs to return0. - 4. The paths in lines 7 and 8 really exist:
ls <caminho>works without errors.
🎓 Final course project
You deliver four items, all about your own configuration — nothing hypothetical:
The 10 sections, saved in .md (module 3.1)
With the diff visible (module 3.2)
≥2 real tasks × 3 versions (module 4.1)
≤10 lines, outside CLAUDE.md (this module)
Approval criteria
- •All four items are present.
- •Every applied removal has quoted excerpt + risk + how to test.
- •Every returned instruction has repeated failure recorded — no reintroducing things out of caution.
- •The final version has at least one objective check where there wasn't one before.
✓ Exit criterion for this module
Written policy, with the date of the next scheduled re-audit — and you can tell, by looking at the CLAUDE.md final, why each remaining line survived. If there’s a line where the answer is “who knows, it’s always been there,” it’s the first line of your next audit.
Quick check (doesn't block anything): you finished the audit, cut the Top 10, and wrote your policy in 8 lines. Where should it live?
🏁 Module Summary and end of course
The entire course, in four lines
- T1Why delete — configuration gets stale: each instruction fixes a specific model’s weakness and becomes dead weight when that weakness goes away.
- T2How to diagnose — 10 categories, 6 decisions, 7 questions; and turning micromanagement into an objective + guardrails + criteria + verification.
- T3Audit for real — the skill that reads, classifies, and proposes without touching anything; and turning the Top 10 into safe cuts, with skills as the unit for trimming.
- T4Prove and maintain — A/B/C on real tasks to prove it didn’t get worse, and the continuous cycle to keep the sediment from returning.
The four deliverables of the final project
10 sections, saved in .md
CLAUDE.md with diff
≥2 real tasks
≤10 lines, dated
This is the end of the course. You came in with a config that kept growing on its own and leave with one you can defend line by line—and, more importantly, with the habit of reviewing that defense when the world changes. The hard part wasn't deleting things; it was resisting the urge to write a rule after the first failure. Save the date of your next re-audit somewhere that will hold you accountable. It'll come sooner than you think.