PTENES
Skip to content
MODULE 1.2

🧪 The ablation method

Ablation means deleting to measure. You delete the configuration, use the tool on real work, observe where it stumbles—and only restore an instruction after seeing the same failure repeat. Rebuild based on evidence, never prediction.

6
Topics
45
Minutes
Basic
Level
Method
Type
Progress in this module
0%0 of 6
1

🔬 Understand what ablation means

Ablation is a term borrowed from research: you remove a component of a system for measure the impact it had. It's not cleanup, and it's not opinion — it's an experiment. The ablation always answers the same question: what changes when this is no longer here?

Applied to prompts and configuration, the method is literal: the entire system prompt is deleted, then each line is brought back one at a time to find out what each line actually does. That’s how the Claude Code team discovered that more than 80% of its own system prompt wasn’t doing anything—the model was already handling it on its own.

🆕 New here? Three words from this module

  • Ablation: remove a piece of the system on purpose to measure how much it contributed.
  • Hook: a Claude Code automation hook — a command that runs automatically before or after an action (e.g., run the formatter every time a file is saved).
  • Baseline: the baseline measurement—how the system behaves without nothing on top. That’s what you compare everything else against.
Predict (what almost everyone does) read the config “I think it’s needed” keep everything nothing measured Measure (ablation) remove the line actually use observe the result measured impact the difference isn’t the care you take—it’s whether you have a measurement at the end

What to look at: both routes have the same number of steps, and the top one even seems more responsible (“I read everything carefully”). But look at the last box in each row: only the bottom one ends in a data. Reading the prompt to review it produces an opinion; removing it and using it produces evidence. Also notice that the step in cyan — “actually use it” — the only thing that doesn’t happen in your head.

✓ This is ablation

  • ✓Remove the instruction and run the same real task again
  • ✓Compare the result with and without it, on the same task
  • ✓Conclude “it’s a tie” and leave it removed

✗ This is not an ablation

  • ✗Reread the CLAUDE.md and think line 42 is important
  • ✗Ask the model to “review my prompt” — it’s guessing too
  • ✗Cut everything at once and never use the trimmed config for real work
2

📅 Schedule when to delete it

Boris Cherny’s advice for anyone who uses Claude Code — and it doesn't build agentic products — is uncomfortably direct: approximately every six months, and especially with a major model launch, delete the CLAUDE.md, delete the skills, delete the hooks—and see what the model does without them. This isn’t rhetoric. It’s the literal recommendation.

Two reasons support this schedule, and it’s worth understanding both because they point in the same direction.

📊 The two reasons, in numbers

  • You’re terrible at predicting. Even the team that built Claude Code got it wrong: when they measured, more than 80% of the system prompt was dead weight. If the person who wrote the harness was 80% wrong, your intuition about your own CLAUDE.md isn’t better.
  • Each line costs you on every run. A rule in the CLAUDE.md isn’t read “when needed.” It enters the context in 100% of runs, including the 95% where it’s irrelevant. The cost isn’t writing the line—it’s rereading it forever.
  • The trigger isn’t just the calendar. A major model release matters more than the date: that’s exactly when old weaknesses disappear and the fixes written for them become dead weight.

💡 "Delete" here is reversible

Deleting doesn’t mean destroying. It means take out of circulation for a real period of work, with the old version saved and recoverable — in git, or simply by renaming the file. If you can’t go back in 30 seconds, you’re not running an experiment: you’re taking a risk. Topic 4 shows how to set up that safety net.

3

🔁 Follow the 4 reconstruction steps

Deleting is only the first half. Rebuilding is what determines whether the new configuration will actually be smaller or just different. There are four steps, and the fourth is the one that usually gets skipped — which is exactly why most configs return to their original size within two weeks.

The cycle, step by step

  1. 1

    Delete the configuration

    CLAUDE.md— skills, hooks. Out of circulation, with the old version saved. No middle ground: cutting “only the worst ones” doesn’t measure anything, because you’re still trusting your prediction about which ones were the worst.

  2. 2

    Use it on real work — never in a hypothetical test

    The real product, the real codebase, the task you’d do today anyway. An invented test produces an invented failure: you end up putting back instructions for problems you’d never encounter in practice.

  3. 3

    Observe what works well and where it stumbles

    Record both sides. Recording only the stumbles skews your read—you end the experiment convinced everything got worse, when 90% stayed the same. Getting it right without instructions is the most valuable evidence there is.

  4. 4

    Bring back an instruction only after you see the failure REPEAT

    This is the step that matters. An isolated failure may be noise: poor context, an ambiguous request, an off day. Only repeated same class of a failure proves there’s a structural gap—and only a structural gap justifies paying for context forever.

1 · delete the config 2 · real work 3 · observe 4 · return only if it repeated didn’t repeat → run again, return nothing rare output the usual path is the cyan loop, not box 4

What to look at: the arrow cyan dashed is the widest in the diagram for a reason. In practice, most observations end there—the stumble doesn’t happen again, and you simply keep working without writing any rule. Box 4, with its glow, looks like the flow’s destination, but it’s marked “rare exit.” A config that grows every week is a config where that loop is never used.

4

⚙️ Build your baseline

There are two ways to run your own ablation. The first is the system prompt flag at Claude Code startup: it replaces the system prompt with whatever you want—including nothing. The second is the environment variable CLAUDE_CODE_SIMPLE=1, poorly documented, that removes all the system prompts — including the prompts associated with each tool. This is the ablation baseline used internally at Anthropic.

One environment variable is a value you set in the terminal that the program reads on startup; write it before the command (VAR=1 comando) applies only to that run— nothing is configured permanently. That’s why it’s the perfect tool for an experiment.

🧪 Copy and run: the baseline

Objective: open a session with no system prompt at all, then a session without your CLAUDE.md. Both are reversible.

# 1) Linha de base total — remove TODOS os prompts de sistema,
#    inclusive os das ferramentas. Vale só para esta sessão.
CLAUDE_CODE_SIMPLE=1 claude

# 2) Alternativa mais segura — mantém o produto normal,
#    mas tira só a SUA configuração do caminho.
mv ~/.claude/CLAUDE.md ~/.claude/CLAUDE.md.bak
claude          # trabalhe normalmente por alguns dias

# 3) Restaurar quando quiser (ou quando o experimento acabar)
mv ~/.claude/CLAUDE.md.bak ~/.claude/CLAUDE.md

How to verify: three signs that the ablation is actually active—(1) ls ~/.claude/CLAUDE.md returns “No such file” during the experiment; (2) in the session, ask “summarize in one sentence the rules you received in this project” — if it repeats your usual rules, something wasn’t removed (check the CLAUDE.md of the project, in addition to the global one); (3) the behavior changes in something — if nothing changed at all, great: that’s already the first result of the experiment.

Replace: the global path through ./CLAUDE.md if what you want to ablate is a specific project’s config.

⚠️ Do this with the config under git

Before moving or deleting anything, make sure ~/.claude/ (or the folder .claude/ of the project) is versioned and everything is committed. A mv A typo in a directory without git can wipe out months of skills and hooks without warning — and turn the experiment into a loss.

Practical rule: if you can’t restore everything with one command, don’t start. Skills and hooks also count toward this, not just the CLAUDE.md.

💡 The counterintuitive detail

Without the system prompts, the model is slightly smarter. You didn’t misread it: fewer instructions, a slightly better result. The prompt costs attention, and some of it pushes the model toward paths it no longer needs to take.

So why are the prompts still there? Because they make Claude Code behave as the person expects when using a product: predictable output format, permissions, tone, tool use. The practical conclusion for you is good: if even the official prompt doesn’t exist because of a capability limitation, much less your 300 lines of CLAUDE.md.

5

🎯 Retire your evals at the right time

If instructions age quickly, what lasts? Evals last longer—but not forever. An eval typically survives for one to three model generations before the model saturate it; at that point, it’s discarded and replaced with a harder one.

🆕 New here? Evals and saturation

  • Eval: a test case for the model. A concrete task with a clear way to tell whether the result passed. It’s the “automated test” for your work with the agent—it measures capability, not behavior correction.
  • Saturate an eval: the model starts getting it right every time. From then on, it stops telling you anything: whether the model is good or great, the result is the same. A saturated eval is noise dressed up as a metric.

✓ An eval worth keeping

  • ✓It came from a point where you actually observed the model to stumble
  • ✓It still fails sometimes—the result varies, so it’s still measuring something
  • ✓Use a real task from your job, with the pass criteria written down beforehand

✗ Eval to retire

  • ✗It passes 100% of the time across three model generations
  • ✗Invented in theory (“it would be good to test this”), with no observed failure behind it
  • ✗Test a weakness in a model you don’t even use anymore

📊 The life cycle of an eval

  • Generation 0: you see the failure in real work and turn the case into an eval. It fails often.
  • Generation 1 to 3: the pass rate rises, but fluctuates. This is the useful phase—the eval can still distinguish a good model from a great one.
  • Saturation: a consistent 100% success rate. The eval can’t distinguish anything anymore; keeping it only wastes execution time and gives a false sense of coverage.
  • Retirement: archive (don’t delete the history—it documents what has already been difficult) and write a new one based on the latest stumbling block you observed.

💡 An eval isn't an instruction

Common confusion: “if the model gets X wrong, I’ll write a rule about X.” But evals and instructions solve different problems. The eval detects the problem and costs zero context — it runs when you run it. The instruction tries to fix the problem and costs context on every run. When a new failure occurs, the eval comes first; the instruction only comes back if the reintroduction rule in the next topic allows it.

6

📓 Apply the reintroduction rule

A removed instruction doesn’t come back because you missed it. It comes back when it passes four conditions, all together. If any one is missing, the instruction is still out — and the work continues normally without it.

✓ The 4 conditions for handing it back

  • ✓1. Real failure: happened in your own work, not in a made-up test
  • ✓2. Repetition: a same class of the failure has occurred at least a second time
  • ✓3. A specific instruction solves: you can point to which sentence would have prevented that
  • ✓4. Shortest possible form: one line, not a 12-step procedure

✗ Reasons that don't count

  • ✗“I feel safer with this rule there”
  • ✗Made a mistake once, on a request you wrote poorly yourself
  • ✗“Since I’m already working on this, I might as well restore the other three”
  • ✗The failure was different, and you brought back the old rule by association

📝 Exercise: create your ablacao-diario.md

Objective: record a real task done with the config ablated. Without a log, step 4 of the cycle is impossible — you have no way to know whether the failure repeated.

# Diário de ablação

Config ablacionada em: 2026-08-18
O que saiu: CLAUDE.md global · skills X e Y · hook de pre-commit
Como restauro: `mv ~/.claude/CLAUDE.md.bak ~/.claude/CLAUDE.md`

---

## Tarefa 1 — <o que você realmente precisava fazer>
- **Data:** 2026-08-18
- **Tarefa real:** refatorar o módulo de export do projeto Z
- **Acertos (foi bem sem instrução):**
  - encontrou os arquivos certos sozinho
  - rodou os testes sem eu pedir
- **Tropeços (onde falhou):**
  - commitou sem eu autorizar
- **Classe do tropeço:** permissão / commit não solicitado
- **Repetiu?** ainda não (1ª ocorrência)
- **Instrução devolvida:** NENHUMA — aguardando repetição

How to verify: the field “Type of stumble” is what makes the journal work. It needs to be generic enough for you to recognize the same fails a week from now in a different context — “committed without authorization” is a category; “committed the utils.ts file at 2 p.m.” is an anecdote. If you can't name the category, the stumble was probably noise.

Exit criterion for this module: daily with ≥1 real task recorded, with successes, stumbles, and “repeated?” filled in — and no instructions returned yet. Reintroducing it at the first occurrence is the mistake this entire module is designed to prevent.

You deleted the CLAUDE.md and on the first real task, the model committed without authorization. What does the reintroduction rule say to do?

📌 Module Summary

✓
Ablation means deleting to measure — remove a component to discover its real impact. Rereading the prompt doesn’t measure anything.
✓
Every ~6 months and with a major model release — delete it CLAUDE.md— skills and hooks, and see what the model does without them.
✓
Four steps, and the fourth is the one that matters — delete it, use it in real work, observe, and restore it only after the failure repeats.
✓
CLAUDE_CODE_SIMPLE=1 is the baseline — remove all system prompts; without them, the model gets slightly smarter.
✓
Evals last 1 to 3 generations — create them where you observed the stumble; retire the ones the model has outgrown.
✓
Reintroducing it requires all 4 conditions — a real, repeated failure, with a specific instruction, in the shortest form possible.

Next Module:

2.1 — The taxonomy: what each line is. Ten categories, six decisions, and the seven questions that classify any instruction in your config.