PTENES
Skip to content
MODULE 4.1

📊 The A/B/C ablation plan

You cut it. Now prove it. This module turns your opinion about the cut into evidence: three config versions, real tasks from your work, nine dimensions measured the same way across all three. Without this test, cutting is a guess — and adding the instruction back is too.

6
Topics
50
Minutes
Advanced
Level
Practical
Type
Progress in this module
0%0 of 6
1

🧪 Build the three versions

Ablation is a term borrowed from research: removing part of a system to measure what it actually contributed. Here, the removed part is instructions from your config. And measurement requires at least two configs running the same thing—in practice, three: A, B e C.

A is your current config, intact—the baseline (the honest point of comparison, what you have today and must not let get worse without noticing). B is the simplified version, with about 50% fewer instructions: you kept what seemed essential and cut the rest. C is the minimal version: essential context + objective + guardrails + quality criteria + verification. Nothing more.

A — current configB — 50% lessC — minimal everything that exists todayhalf the instructions context + objective +guardrails + criteria + verification the SAME 5 real tasks run through A, B, and C the only thing that changes between the columns is the config — the task is identical in all three

What to look at: the columns shrink from left to right, but the cyan bands run across all three with the same length. That detail is what makes the test meaningful: only one variable changes (the config). If you also change the task, model, or input prompt, the result no longer tells you anything about the cut.

✓ What GOES INTO version C

  • ✓Essential context — project identity, paths, sources of truth, internal conventions. The model can't infer this.
  • ✓Objective — the result to be produced, in one sentence.
  • ✓Guardrails — the non-negotiable limits: security, compliance, what must never be touched.
  • ✓Quality criteria — what “good” looks like.
  • ✓Verification — how the model checks its own work before finishing.

✗ What is LEFT OUT of version C

  • ✗Step-by-step execution ("first do X, then Y, then Z").
  • ✗A rule repeated in three different places.
  • ✗Behavior correction for a model that no longer runs.
  • ✗Excessive examples and rigid formatting without a stated reason.
  • ✗Instructions that teach the model to think instead of saying what to deliver.

💡 C isn't "empty config"

The most common mistake when building C is confusing minimal with none. If you delete the repository path, domain name, git account, or output file format, the model has no way to guess. That’s not dead weight; it’s context that exists only in your head and your files. C cuts execution instructions; C preserves context the model can't infer.

Key concepts

Ablation

Remove it to measure the impact

Baseline

A is the point of comparison

C version

Context + objective + guardrails + criteria + verification

Minimal ≠ empty

Non-inferable context stays

2

🎯 Choose 5 to 10 real tasks

A real task is one you’ve already done or will do anyway this week: publish a post, refactor a module, generate a report, build a page. A hypothetical task—“ask the model to write a Fibonacci function”—doesn’t test any of your config, because your config wasn’t written for Fibonacci.

Five tasks is the minimum to keep the result from being a matter of luck; ten is the practical ceiling before the round becomes a project in itself. The set needs to be representative: if 70% of your work is writing content and 30% is working with code, the grid should reflect roughly that proportion.

✓ A good task grid has

  • ✓Coverage of the kinds of work you actually do
  • ✓At least one task where the current config clearly helps — without it, the test starts out rooting for removal
  • ✓At least one difficult task, a little beyond what you think the model can handle
  • ✓Concrete inputs: the files, links, and real data the task uses

✗ Signs of a biased rubric

  • ✗All tasks are easy—any version passes, so the test distinguishes nothing
  • ✗None affect the rules you want to keep
  • ✗You chose the tasks after you'd already decided to cut
  • ✗They’re all the same type (all code, all text)

⚠️ Warning: the toy-task bias

Bias here, any test choice that pushes the result toward the side you already preferred. Testing with a toy task is the most common and quietest way to “prove” what you already wanted to believe: on a trivial task, A, B, and C always tie—and the tie seems like permission to cut everything.

The antidote is to choose the tasks before of building B and C, and deliberately include one where you bet the current config makes a difference. If it still ties, then you’ve learned something.

Key concepts

5 to 10 tasks

Minimum versus luck

Representative

Reflect your actual work

Bias

Test that already knows what it wants to prove

Difficult task

This is where the versions diverge

3

📏 Measure the nine dimensions

"It got better" isn't a measure. Nine dimensions cover what matters, and each has a simple way to score it: a scale of 1 to 5 when judgment is unavoidable, objective counts whenever possible. Counts are always preferable — "3 fixes" leaves no room for debate; "quality 4" does.

Three of these dimensions are often ignored, and they’re exactly the ones that reveal bloated config: autonomy (how far the model gets on its own before stopping to ask), adherence (whether it followed the rules you wanted keep — not all the rules that existed) and the ability to verify its own work.

Dimension What it is How to score
QualityThe result meets the criteria you defined1–5 (1 = I'd redo it from scratch, 5 = I'd ship it as is)
AdherenceFollowed the rules you wanted to keepCount: how many of the N target rules were followed
AutonomyHow far did you get without asking for helpCount: number of questions/interruptions to the human
Human correctionsHow many times did you have to step in and correct itDirect count (0, 1, 2, …)
ConsistencySame task, two runs: does it produce the same kind of result?Same / similar / different—run the task 2×
TimeFrom submission to usable deliveryMinutes, by the clock
Unnecessary use of toolsSearches, reading, and commands that were of no useCount of disposable calls
ComplexitySize of the config that produced thatNumber of lines / number of rules in the version
Self-checkDid it check its own work before finishing?No / checked superficially / checked with evidence

📊 The objective function, in one line

You're not looking for the shortest config. You're looking for the greatest value of (quality + autonomy + verifiability) ÷ complexity.

That’s why complexity appears in the table as a measured dimension, not a separate goal: cutting 200 lines and losing autonomy is a terrible trade; cutting 200 lines while keeping everything the same is pure context savings.

Key concepts

Adherence

Met the target rules

Autonomy

Fewer interruptions = better

Count > rating

Objective whenever possible

Self-check

The most overlooked dimension

4

📝 Record it without fooling yourself

Three things stay frozen in every run: same task, same model, same input prompt. Only the config changes. Comparing yesterday’s A, on the old model, with today’s C on a similar task isn’t ablation — it’s anecdotal.

And there’s a rule that separates those who measure from those who convince themselves: write down what "good" would look like before looking at the answer. If you define the criteria after reading the output, your brain adjusts the criteria to the output — you’ll think what came back is great and won’t notice that you moved the goalposts. Write down the criteria first, then run it.

✓ Honest protocol

  • ✓Write down "what good looks like" in the rubric before running it
  • ✓Run A, B, and C in the same work session, on the same day
  • ✓Fill in the row in the grid right after each round, not at the end of the day
  • ✓Save the link/file for each run's output so you can review it later

✗ How you fool yourself

  • ✗Compare yesterday’s A with today’s C across different tasks
  • ✗Switching models mid-run
  • ✗Rewrite the “just a little” input prompt in version C
  • ✗Assess memory quality two days later, without notes

📋 Copy: the tracking grid

Objective: create the file ablacao-grade.md alongside yours ablacao-diario.md (from module 1.2). One table per task, three rows per table — A, B, and C.

# Grade de ablação A/B/C

Modelo usado: <nome do modelo>      Data: <aaaa-mm-dd>
Regras-alvo (para aderência): <liste as N regras que você quer manter>

## Tarefa 1 — <sua tarefa aqui>
Prompt de entrada (idêntico nas 3): <cole o prompt>
O que seria BOM (escrito ANTES de rodar): <seu critério aqui>

| Versão | Qualid. 1-5 | Aderência n/N | Autonomia (interrupções) | Correções | Consistência | Tempo (min) | Ferram. inúteis | Complexidade (linhas) | Autoverificação |
|---|---|---|---|---|---|---|---|---|---|
| A (atual)      |   |   |   |   |   |   |   |   |   |
| B (50% menos)  |   |   |   |   |   |   |   |   |   |
| C (mínima)     |   |   |   |   |   |   |   |   |   |

Observações (o que quebrou, em qual versão, com trecho):
- <anote aqui>

## Tarefa 2 — <sua tarefa aqui>
(repita o bloco acima)

How to verify: the line "What would be GOOD" is filled in before of any cell in the table. If you filled in the table first and the criterion afterward, delete the criterion and redo the task—the record is contaminated.

💡 Consistency costs an extra round

The “consistency” dimension requires running the same task twice on the same version. Do this at least on version C and for the difficult task — that’s where instability shows up. In the other cells, if time is short, mark “not measured” instead of guessing. An empty cell is honest; a guessed cell distorts the picture.

Key concepts

Three frozen

Task, model, prompt

Benchmark first

Criterion before the response

ablacao-grade.md

Where the test becomes a record

Not measured

Better than guessing

5

⚖️ Read the result and decide

Reading the grid has three outcomes, and only three. Where C tied with A, the instructions that were in A and disappeared in C were dead weight: they stay out, permanently. Where C performed worse once, that doesn’t count—a bad run is noise, and restoring an instruction because of it is like rewriting the entire config every time the model yawns.

Where C performed worse repeatedly, that’s when the reintroduction rule comes in. And there’s a fourth case that comes up often and surprises people: B beating A and C. When this happens, you’ve found the sweet spot—that’s usually where the config stays.

Did C tie with A?look at the row in the grid YES → it was dead weightdefinitively out NO → did it repeatedly get worse?the SAME class of failure, 2×+ Just once → it was noisedoesn’t count, keep C It happened again → return the instructionin the shortest form that fixes that failure only the path below, with repeated failure, authorizes returning an instruction

What to look at: there are four possible outcomes, and only a returns instructions for the config. Notice that “C got worse” alone isn’t an outcome—it has to pass through the second diamond. This flowchart exists to help you resist the urge to rewrite the config at the first frustration.

1

A real failure occurred

Not “it seemed a little weird”: something you had to correct, or that broke a criterion written down before the run. Failure recorded in the rubric, with an excerpt.

2

The same class of failure occurred again

Two occurrences of the same type—not two different errors. “Forgot to run the test” twice is a repeated class; “forgot the test” and “used the wrong author” are two classes with one occurrence each.

3

It's clear that a specific instruction solves

If you can’t write the sentence that would have prevented the failure, the problem isn’t a lack of instructions — maybe you need a skill (procedure) or context the model can’t access.

4

It fits in the shortest form possible

One sentence, one criterion, one check. If the instruction you restore has eight steps, you didn’t restore a rule—you brought back the micromanagement you had just cut.

📊 When B wins, B stays

If B ties A on quality and adherence, but wins on autonomy and complexity — and C repeatedly loses on two tasks — the answer isn’t to force C. It’s to adopt B as the new baseline and run the next ablation from there in six months. Ablation is iterative: today’s C is the next round’s A.

Key concepts

A tie means cut

It was dead weight

One time = noise

Returns nothing

4 conditions

Reintroduction rule

B wins

The common sweet spot

6

🔁 Turn failures into evals

One eval is a test case for a model: a concrete task with a success criterion you can check. Every failure that showed up in your grid is already a ready-made eval— just write "this task, this criterion, the model failed here" and save it.

Evals age too: they typically last 1 to 3 model generations. When an eval passes every time, in every version, it got saturated — no longer separates anything and only costs execution time. Retire it and create new ones where you see the current model stumble.

Task A (current) B (50% less) C (minimal) Reading
Publish a blog postQual. 4 · 2 fixesQual. 4 · 1 fixQual. 4 · 1 fixTie → cut
Refactor a legacy moduleQual. 4 · 1 fixQual. 4 · 1 fixQual. 2 · 4 fixes (2×)Repeatedly worse → give back 1 rule
Generate monthly reportQual. 3 · 3 fixesQual. 4 · 1 fixQual. 3 · 2 fixesB won → B becomes the baseline
Fix the reported bugQual. 5 · 0 fixesQual. 5 · 0 fixesQual. 4 · 1 fix (1×)Just once → noise, keep C

How to read this example table: it’s illustrative—the numbers are made up. What matters is the "Reading" column: four tasks, four different outcomes. A real grid rarely points to a single verdict, and that’s exactly why it’s worth more than your intuition.

🧪 Exercise: run your rubric

Objective: build the A/B/C grid with 5 real tasks and run it at least 2 of them across the three versions— filling in the table. Paste the prompt below into Claude Code, one version at a time, without changing anything but the config.

Rodada de ablação — versão <A | B | C>

Tarefa: <sua tarefa aqui, exatamente como você a pediria num dia normal>
Entradas: <arquivos, links, dados reais que a tarefa usa>

Faça a tarefa até o fim. Ao terminar, responda também:
1. Quais critérios você usou para decidir que estava pronto?
2. Como você verificou o próprio trabalho? Cite a evidência (comando,
   arquivo conferido, teste rodado).
3. Em que ponto você ficou em dúvida e teria perguntado a um humano?

Não peça confirmação no meio do caminho a menos que seja bloqueante.

How to verify (exit criterion): ablacao-grade.md with ≥2 tasks × 3 versions filled in; and, for each instruction that returned to the config, the failure that justified it is recorded with an excerpt and if it happened at least twice. If an instruction came back without a recorded repeated failure, delete it and run it again.

Now replace it with yours: the model’s answers 1, 2, and 3 feed directly into three columns of the grid—quality criteria, self-checking, and autonomy. Paste them into the task’s “Observations” field.

Quick check (doesn't block anything): in the "refactor legacy module" task, version C forgot to run the tests — once. On the other tasks, C tied with A. What do you do?

💡 The eval born from failure

Each "Notes" row in your grid becomes a one-sentence eval: “task X, with config C, must run the tests before wrapping up — evidence: test command output in the log”. Keep these cases in one file and rerun them whenever a new model comes out. That way, the next ablation starts with a ready-made benchmark.

Key concepts

Eval

Verifiable test case

Failure → eval

Nothing observed is lost

Saturated

If it always passes, retire it

1 to 3 generations

Typical validity period of an eval

📌 Module Summary

✓
Three versions — A, current; B, 50% less; C, minimal. C isn’t empty: non-inferable context remains.
✓
Real tasks — 5 to 10, representative ones, including a difficult one and one where the current config clearly helps.
✓
Nine dimensions — objective counts whenever possible; self-verification is the most often forgotten.
✓
A benchmark before the response — same task, same model, same prompt; and “what good looks like” written down before running it.
✓
A tie means cut; noise doesn’t count — only repeated failures bring an instruction back, in the shortest form possible.
✓
Failure becomes an eval — the personal set of cases survives 1 to 3 generations; retire the ones that have saturated.

Next Module:

4.2 — The continuous cycle: turn the audit into a regular routine, with re-audit triggers and a written ablation policy.