🏅 The gold field message.model
Every comparison starts by separating each model’s corpus, and the key is the field message.model.
O extract_corpus.py scans the entire history (all projects in
~/.claude/projects) and combine them into a joint corpus all turns from a model.
# default: claude-fable-5 python extract_corpus.py # um modelo específico, com pasta de saída python extract_corpus.py --model claude-opus-4-8 --out ./corpus_opus # só listar os modelos que existem no histórico python extract_corpus.py --list
📦 What it produces
- •
transcript.md— a lightweight transcript with only that model's turns. - •
stats.json— all metrics measured. - •A readable report printed in the terminal.
💡 Start with the --list
Before filtering, run --list to see the exact names of the models in your history. Names change between versions—the --list keeps you from filtering for a model that doesn't exist and getting an empty corpus.
message.model
the entire history
models present
transcript + stats
🔗 LOGICAL TURN: the right unit
In the previous module, we saw that each block is one line. If you measure by line, a response turn with 5 blocks becomes “5 turns”—the signal gets diluted. The right unit is the logical turn: 1 human prompt until the next, with everything the model did in between. All metrics in this module are per logical turn.
🎯 Why this is the right unit
The logical turn answers the question that matters: "to answer a this request, did the model think first? How many tools did it use? In what order?” That’s the granularity of work pace — neither the line (too fine-grained) nor the session (too coarse-grained).
prompt → prompt
all blocks
the actual pace
per logical turn
📐 The measured metrics
O extract_corpus.py doesn’t give impressions—it gives numbers. Each corpus is summarized by:
logical turns, % that thought first, tools per turn (mean and median), read-before-edit and test-after-edit.
This is the shared vocabulary used to calculate the delta.
Logical turns & % that think first
How many logical turns the model had, and in what fraction of them there was a reasoning block before of the first action.
Tools / turn (mean & median)
Action density per turn. Average and median together—because a handful of huge turns can pull up the average without changing what’s typical.
Read before editing & test after editing
Discipline heuristics: did you read the file before editing it? Did you run a test after editing? Measured by tool sequence.
💡 Numbers, not impressions
The trick is to replace “I thought Fable thinks more” with “Fable thought first in 99% of turns.” That’s the number you can defend, transfer, and—in track 3—turn into a playbook rule.
presence 🧠
mean + median
discipline
discipline
⚖️ Compare 2 models
With two measured corpora, the compare_models.py put the models
side by side and prints the column Δ — the distance between the two paces.
This is where the difference that becomes a playbook is visible in a single table.
# fable-5 vs opus-4-8 (default) python compare_models.py # explícito, gravando o resultado python compare_models.py --a claude-fable-5 --b claude-opus-4-8 --out compare.json
métrica claude-fable-5 claude-opus-4-8 Δ ──────────────────────────────────────────────────────────────── % turnos c/ raciocínio 99% 54% +45% ferramentas/turno (média) 6.57 7.86 -1.29 sessões 7 1114 turnos do assistente 69 ...
📄 O compare.json
With --out compare.json, the result becomes a structured file—easy to version and feed into make_playbook.py in track 3. The terminal table is for reading; the JSON is for the machine to follow the pipeline.
both models
table + Δ
compare.json
becomes a playbook
🎯 The actual measured delta
Here’s the finding that drives the course. Fable-5 thought beforehand in 99% of logical turns against 54% from Opus — a delta of +45 points, the strong, transferable signal. In tools/turns: 6.57 vs 7.86 — Fable was more economical.
⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.
✓ What the number says
- ✓Fable thinks before acting almost every time (99%).
- ✓+45 points is a large, clear delta.
- ✓Fable uses fewer tools per turn (6,57).
✗ What NOT to infer
- ✗“Fewer tools = worse” — it could be density.
- ✗“More tools = better” — it could be thrashing.
- ✗That the number of tools alone determines the ranking.
💡 Honesty
More tools no is better on its own: it can be density (solved it with a few right actions) or thrashing (tried several things until it got it right). The solid, transferable delta here is the think-before-acting — this becomes an anchor rule in track 3.
99%
54%
+45 pts
6.57 vs 7.86
⚠️ The limits
An honest course also tells you where NOT to step. The Fable sample is small (7 sessions / 69 turns), which makes read-before-edit and test-after-edit noisy — don't draw robust conclusions from them. And the reasoning text comes encrypted: presence is measured, never content.
⚠️ What the small sample costs
With 7 Fable sessions versus 1114 Opus sessions, any low-frequency metric (few edits, few tests) fluctuates a lot. The think-before-acting survives because it appears in almost every turn; read-before-edit e test-after-edit are good practice measured using a conservative heuristic, not a defensible delta.
Small Fable sample
7 sessions make rare metrics unstable. Treat read/test as a weak signal, not as a delta.
Encrypted reasoning
The thinking isn’t in the logs. You measure the presence (🧠), never the content of the thought.
Where Fable was weak
Overthinking the trivial (thinking too much even for simple tasks) and verbosity. The delta isn't "Fable is better at everything".
💡 The golden rule of honesty
Only “think before acting” is a firm delta. State that confidently; treat the rest as context. That discipline is what makes the Track 3 playbook defensible—you inject what you measured, not what you hoped was true.
7 sessions
read/test
presence, not text
think-first only
📊 Module Summary
extract_corpus.py filter by message.model — the entire history; --list shows the models.compare_models.py gives you the table + Δ — and the compare.json for the pipeline.Next Track:
Track 3 — From Delta to Playbook: turn the number into an injectable rule.