🧪 Build the three versions
Ablation is a term borrowed from research: removing part of a system to measure what it actually contributed. Here, the removed part is instructions from your config. And measurement requires at least two configs running the same thing—in practice, three: A, B e C.
A is your current config, intact—the baseline (the honest point of comparison, what you have today and must not let get worse without noticing). B is the simplified version, with about 50% fewer instructions: you kept what seemed essential and cut the rest. C is the minimal version: essential context + objective + guardrails + quality criteria + verification. Nothing more.
What to look at: the columns shrink from left to right, but the cyan bands run across all three with the same length. That detail is what makes the test meaningful: only one variable changes (the config). If you also change the task, model, or input prompt, the result no longer tells you anything about the cut.
✓ What GOES INTO version C
- ✓Essential context — project identity, paths, sources of truth, internal conventions. The model can't infer this.
- ✓Objective — the result to be produced, in one sentence.
- ✓Guardrails — the non-negotiable limits: security, compliance, what must never be touched.
- ✓Quality criteria — what “good” looks like.
- ✓Verification — how the model checks its own work before finishing.
✗ What is LEFT OUT of version C
- ✗Step-by-step execution ("first do X, then Y, then Z").
- ✗A rule repeated in three different places.
- ✗Behavior correction for a model that no longer runs.
- ✗Excessive examples and rigid formatting without a stated reason.
- ✗Instructions that teach the model to think instead of saying what to deliver.
💡 C isn't "empty config"
The most common mistake when building C is confusing minimal with none. If you delete the repository path, domain name, git account, or output file format, the model has no way to guess. That’s not dead weight; it’s context that exists only in your head and your files. C cuts execution instructions; C preserves context the model can't infer.
Key concepts
Remove it to measure the impact
A is the point of comparison
Context + objective + guardrails + criteria + verification
Non-inferable context stays
🎯 Choose 5 to 10 real tasks
A real task is one you’ve already done or will do anyway this week: publish a post, refactor a module, generate a report, build a page. A hypothetical task—“ask the model to write a Fibonacci function”—doesn’t test any of your config, because your config wasn’t written for Fibonacci.
Five tasks is the minimum to keep the result from being a matter of luck; ten is the practical ceiling before the round becomes a project in itself. The set needs to be representative: if 70% of your work is writing content and 30% is working with code, the grid should reflect roughly that proportion.
✓ A good task grid has
- ✓Coverage of the kinds of work you actually do
- ✓At least one task where the current config clearly helps — without it, the test starts out rooting for removal
- ✓At least one difficult task, a little beyond what you think the model can handle
- ✓Concrete inputs: the files, links, and real data the task uses
✗ Signs of a biased rubric
- ✗All tasks are easy—any version passes, so the test distinguishes nothing
- ✗None affect the rules you want to keep
- ✗You chose the tasks after you'd already decided to cut
- ✗They’re all the same type (all code, all text)
⚠️ Warning: the toy-task bias
Bias here, any test choice that pushes the result toward the side you already preferred. Testing with a toy task is the most common and quietest way to “prove” what you already wanted to believe: on a trivial task, A, B, and C always tie—and the tie seems like permission to cut everything.
The antidote is to choose the tasks before of building B and C, and deliberately include one where you bet the current config makes a difference. If it still ties, then you’ve learned something.
Key concepts
Minimum versus luck
Reflect your actual work
Test that already knows what it wants to prove
This is where the versions diverge
📏 Measure the nine dimensions
"It got better" isn't a measure. Nine dimensions cover what matters, and each has a simple way to score it: a scale of 1 to 5 when judgment is unavoidable, objective counts whenever possible. Counts are always preferable — "3 fixes" leaves no room for debate; "quality 4" does.
Three of these dimensions are often ignored, and they’re exactly the ones that reveal bloated config: autonomy (how far the model gets on its own before stopping to ask), adherence (whether it followed the rules you wanted keep — not all the rules that existed) and the ability to verify its own work.
| Dimension | What it is | How to score |
|---|---|---|
| Quality | The result meets the criteria you defined | 1–5 (1 = I'd redo it from scratch, 5 = I'd ship it as is) |
| Adherence | Followed the rules you wanted to keep | Count: how many of the N target rules were followed |
| Autonomy | How far did you get without asking for help | Count: number of questions/interruptions to the human |
| Human corrections | How many times did you have to step in and correct it | Direct count (0, 1, 2, …) |
| Consistency | Same task, two runs: does it produce the same kind of result? | Same / similar / different—run the task 2× |
| Time | From submission to usable delivery | Minutes, by the clock |
| Unnecessary use of tools | Searches, reading, and commands that were of no use | Count of disposable calls |
| Complexity | Size of the config that produced that | Number of lines / number of rules in the version |
| Self-check | Did it check its own work before finishing? | No / checked superficially / checked with evidence |
📊 The objective function, in one line
You're not looking for the shortest config. You're looking for the greatest value of (quality + autonomy + verifiability) ÷ complexity.
That’s why complexity appears in the table as a measured dimension, not a separate goal: cutting 200 lines and losing autonomy is a terrible trade; cutting 200 lines while keeping everything the same is pure context savings.
Key concepts
Met the target rules
Fewer interruptions = better
Objective whenever possible
The most overlooked dimension
📝 Record it without fooling yourself
Three things stay frozen in every run: same task, same model, same input prompt. Only the config changes. Comparing yesterday’s A, on the old model, with today’s C on a similar task isn’t ablation — it’s anecdotal.
And there’s a rule that separates those who measure from those who convince themselves: write down what "good" would look like before looking at the answer. If you define the criteria after reading the output, your brain adjusts the criteria to the output — you’ll think what came back is great and won’t notice that you moved the goalposts. Write down the criteria first, then run it.
✓ Honest protocol
- ✓Write down "what good looks like" in the rubric before running it
- ✓Run A, B, and C in the same work session, on the same day
- ✓Fill in the row in the grid right after each round, not at the end of the day
- ✓Save the link/file for each run's output so you can review it later
✗ How you fool yourself
- ✗Compare yesterday’s A with today’s C across different tasks
- ✗Switching models mid-run
- ✗Rewrite the “just a little” input prompt in version C
- ✗Assess memory quality two days later, without notes
📋 Copy: the tracking grid
Objective: create the file ablacao-grade.md alongside yours ablacao-diario.md (from module 1.2). One table per task, three rows per table — A, B, and C.
# Grade de ablação A/B/C Modelo usado: <nome do modelo> Data: <aaaa-mm-dd> Regras-alvo (para aderência): <liste as N regras que você quer manter> ## Tarefa 1 — <sua tarefa aqui> Prompt de entrada (idêntico nas 3): <cole o prompt> O que seria BOM (escrito ANTES de rodar): <seu critério aqui> | Versão | Qualid. 1-5 | Aderência n/N | Autonomia (interrupções) | Correções | Consistência | Tempo (min) | Ferram. inúteis | Complexidade (linhas) | Autoverificação | |---|---|---|---|---|---|---|---|---|---| | A (atual) | | | | | | | | | | | B (50% menos) | | | | | | | | | | | C (mínima) | | | | | | | | | | Observações (o que quebrou, em qual versão, com trecho): - <anote aqui> ## Tarefa 2 — <sua tarefa aqui> (repita o bloco acima)
How to verify: the line "What would be GOOD" is filled in before of any cell in the table. If you filled in the table first and the criterion afterward, delete the criterion and redo the task—the record is contaminated.
💡 Consistency costs an extra round
The “consistency” dimension requires running the same task twice on the same version. Do this at least on version C and for the difficult task — that’s where instability shows up. In the other cells, if time is short, mark “not measured” instead of guessing. An empty cell is honest; a guessed cell distorts the picture.
Key concepts
Task, model, prompt
Criterion before the response
Where the test becomes a record
Better than guessing
⚖️ Read the result and decide
Reading the grid has three outcomes, and only three. Where C tied with A, the instructions that were in A and disappeared in C were dead weight: they stay out, permanently. Where C performed worse once, that doesn’t count—a bad run is noise, and restoring an instruction because of it is like rewriting the entire config every time the model yawns.
Where C performed worse repeatedly, that’s when the reintroduction rule comes in. And there’s a fourth case that comes up often and surprises people: B beating A and C. When this happens, you’ve found the sweet spot—that’s usually where the config stays.
What to look at: there are four possible outcomes, and only a returns instructions for the config. Notice that “C got worse” alone isn’t an outcome—it has to pass through the second diamond. This flowchart exists to help you resist the urge to rewrite the config at the first frustration.
A real failure occurred
Not “it seemed a little weird”: something you had to correct, or that broke a criterion written down before the run. Failure recorded in the rubric, with an excerpt.
The same class of failure occurred again
Two occurrences of the same type—not two different errors. “Forgot to run the test” twice is a repeated class; “forgot the test” and “used the wrong author” are two classes with one occurrence each.
It's clear that a specific instruction solves
If you can’t write the sentence that would have prevented the failure, the problem isn’t a lack of instructions — maybe you need a skill (procedure) or context the model can’t access.
It fits in the shortest form possible
One sentence, one criterion, one check. If the instruction you restore has eight steps, you didn’t restore a rule—you brought back the micromanagement you had just cut.
📊 When B wins, B stays
If B ties A on quality and adherence, but wins on autonomy and complexity — and C repeatedly loses on two tasks — the answer isn’t to force C. It’s to adopt B as the new baseline and run the next ablation from there in six months. Ablation is iterative: today’s C is the next round’s A.
Key concepts
It was dead weight
Returns nothing
Reintroduction rule
The common sweet spot
🔁 Turn failures into evals
One eval is a test case for a model: a concrete task with a success criterion you can check. Every failure that showed up in your grid is already a ready-made eval— just write "this task, this criterion, the model failed here" and save it.
Evals age too: they typically last 1 to 3 model generations. When an eval passes every time, in every version, it got saturated — no longer separates anything and only costs execution time. Retire it and create new ones where you see the current model stumble.
| Task | A (current) | B (50% less) | C (minimal) | Reading |
|---|---|---|---|---|
| Publish a blog post | Qual. 4 · 2 fixes | Qual. 4 · 1 fix | Qual. 4 · 1 fix | Tie → cut |
| Refactor a legacy module | Qual. 4 · 1 fix | Qual. 4 · 1 fix | Qual. 2 · 4 fixes (2×) | Repeatedly worse → give back 1 rule |
| Generate monthly report | Qual. 3 · 3 fixes | Qual. 4 · 1 fix | Qual. 3 · 2 fixes | B won → B becomes the baseline |
| Fix the reported bug | Qual. 5 · 0 fixes | Qual. 5 · 0 fixes | Qual. 4 · 1 fix (1×) | Just once → noise, keep C |
How to read this example table: it’s illustrative—the numbers are made up. What matters is the "Reading" column: four tasks, four different outcomes. A real grid rarely points to a single verdict, and that’s exactly why it’s worth more than your intuition.
🧪 Exercise: run your rubric
Objective: build the A/B/C grid with 5 real tasks and run it at least 2 of them across the three versions— filling in the table. Paste the prompt below into Claude Code, one version at a time, without changing anything but the config.
Rodada de ablação — versão <A | B | C> Tarefa: <sua tarefa aqui, exatamente como você a pediria num dia normal> Entradas: <arquivos, links, dados reais que a tarefa usa> Faça a tarefa até o fim. Ao terminar, responda também: 1. Quais critérios você usou para decidir que estava pronto? 2. Como você verificou o próprio trabalho? Cite a evidência (comando, arquivo conferido, teste rodado). 3. Em que ponto você ficou em dúvida e teria perguntado a um humano? Não peça confirmação no meio do caminho a menos que seja bloqueante.
How to verify (exit criterion):
ablacao-grade.md with ≥2 tasks × 3 versions filled in; and,
for each instruction that returned to the config, the failure that justified it is recorded with an excerpt and
if it happened at least twice. If an instruction came back without a recorded repeated failure, delete it
and run it again.
Now replace it with yours: the model’s answers 1, 2, and 3 feed directly into three columns of the grid—quality criteria, self-checking, and autonomy. Paste them into the task’s “Observations” field.
Quick check (doesn't block anything): in the "refactor legacy module" task, version C forgot to run the tests — once. On the other tasks, C tied with A. What do you do?
💡 The eval born from failure
Each "Notes" row in your grid becomes a one-sentence eval: “task X, with config C, must run the tests before wrapping up — evidence: test command output in the log”. Keep these cases in one file and rerun them whenever a new model comes out. That way, the next ablation starts with a ready-made benchmark.
Key concepts
Verifiable test case
Nothing observed is lost
If it always passes, retire it
Typical validity period of an eval
📌 Module Summary
Next Module:
4.2 — The continuous cycle: turn the audit into a regular routine, with re-audit triggers and a written ablation policy.