PTENES
Skip to content
MODULE 2.2

📊 Corpus, Numbers, and the Fable vs Opus Delta

Now come the numbers. In this module: extracting a model’s corpus from the full history, why the logical turn is the right unit, the metrics measured, the side-by-side comparison, and the real delta — Fable thought ahead 99% of the time vs. Opus at 54%—with honesty about what NOT to claim.

6
Topics
30
Minutes
Practical
Level
Hands-on
Type
history all projects filters .model extract_corpus corpus · fable-5 corpus · opus-4-8 Δ per logical turn compare_models
1

🏅 The gold field message.model

Every comparison starts by separating each model’s corpus, and the key is the field message.model. O extract_corpus.py scans the entire history (all projects in ~/.claude/projects) and combine them into a joint corpus all turns from a model.

# default: claude-fable-5
python extract_corpus.py

# um modelo específico, com pasta de saída
python extract_corpus.py --model claude-opus-4-8 --out ./corpus_opus

# só listar os modelos que existem no histórico
python extract_corpus.py --list

📦 What it produces

  • •transcript.md — a lightweight transcript with only that model's turns.
  • •stats.json — all metrics measured.
  • •A readable report printed in the terminal.

💡 Start with the --list

Before filtering, run --list to see the exact names of the models in your history. Names change between versions—the --list keeps you from filtering for a model that doesn't exist and getting an empty corpus.

Filter by

message.model

Scope

the entire history

--list

models present

Output

transcript + stats

2

🔗 LOGICAL TURN: the right unit

In the previous module, we saw that each block is one line. If you measure by line, a response turn with 5 blocks becomes “5 turns”—the signal gets diluted. The right unit is the logical turn: 1 human prompt until the next, with everything the model did in between. All metrics in this module are per logical turn.

prompt 1 logical turn = unit of the metric line: thinking 🧠 line: text line: tool_use (Read) line: tool_use (Edit) line: tool_use (Bash) next prompt

🎯 Why this is the right unit

The logical turn answers the question that matters: "to answer a this request, did the model think first? How many tools did it use? In what order?” That’s the granularity of work pace — neither the line (too fine-grained) nor the session (too coarse-grained).

Define

prompt → prompt

Group

all blocks

Measure

the actual pace

Every metric

per logical turn

3

📐 The measured metrics

O extract_corpus.py doesn’t give impressions—it gives numbers. Each corpus is summarized by: logical turns, % that thought first, tools per turn (mean and median), read-before-edit and test-after-edit. This is the shared vocabulary used to calculate the delta.

›

Logical turns & % that think first

How many logical turns the model had, and in what fraction of them there was a reasoning block before of the first action.

›

Tools / turn (mean & median)

Action density per turn. Average and median together—because a handful of huge turns can pull up the average without changing what’s typical.

›

Read before editing & test after editing

Discipline heuristics: did you read the file before editing it? Did you run a test after editing? Measured by tool sequence.

💡 Numbers, not impressions

The trick is to replace “I thought Fable thinks more” with “Fable thought first in 99% of turns.” That’s the number you can defend, transfer, and—in track 3—turn into a playbook rule.

% think-first

presence 🧠

tools/turn

mean + median

read→edit

discipline

edit→test

discipline

4

⚖️ Compare 2 models

With two measured corpora, the compare_models.py put the models side by side and prints the column Δ — the distance between the two paces. This is where the difference that becomes a playbook is visible in a single table.

# fable-5 vs opus-4-8 (default)
python compare_models.py

# explícito, gravando o resultado
python compare_models.py --a claude-fable-5 --b claude-opus-4-8 --out compare.json
métrica                       claude-fable-5   claude-opus-4-8     Δ
────────────────────────────────────────────────────────────────
% turnos c/ raciocínio                   99%              54%   +45%
ferramentas/turno (média)              6.57             7.86   -1.29
sessões                                   7             1114
turnos do assistente                     69             ...

📄 O compare.json

With --out compare.json, the result becomes a structured file—easy to version and feed into make_playbook.py in track 3. The terminal table is for reading; the JSON is for the machine to follow the pipeline.

--a / --b

both models

Output

table + Δ

--out

compare.json

Δ

becomes a playbook

5

🎯 The actual measured delta

Here’s the finding that drives the course. Fable-5 thought beforehand in 99% of logical turns against 54% from Opus — a delta of +45 points, the strong, transferable signal. In tools/turns: 6.57 vs 7.86 — Fable was more economical.

⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.

thought first — claude-fable-5 99% thought first — claude-opus-4-8 54% +45 points

✓ What the number says

  • ✓Fable thinks before acting almost every time (99%).
  • ✓+45 points is a large, clear delta.
  • ✓Fable uses fewer tools per turn (6,57).

✗ What NOT to infer

  • ✗“Fewer tools = worse” — it could be density.
  • ✗“More tools = better” — it could be thrashing.
  • ✗That the number of tools alone determines the ranking.

💡 Honesty

More tools no is better on its own: it can be density (solved it with a few right actions) or thrashing (tried several things until it got it right). The solid, transferable delta here is the think-before-acting — this becomes an anchor rule in track 3.

Fable thinks first

99%

Opus thinks first

54%

Δ think-before

+45 pts

tools/turn

6.57 vs 7.86

6

⚠️ The limits

An honest course also tells you where NOT to step. The Fable sample is small (7 sessions / 69 turns), which makes read-before-edit and test-after-edit noisy — don't draw robust conclusions from them. And the reasoning text comes encrypted: presence is measured, never content.

⚠️ What the small sample costs

With 7 Fable sessions versus 1114 Opus sessions, any low-frequency metric (few edits, few tests) fluctuates a lot. The think-before-acting survives because it appears in almost every turn; read-before-edit e test-after-edit are good practice measured using a conservative heuristic, not a defensible delta.

!

Small Fable sample

7 sessions make rare metrics unstable. Treat read/test as a weak signal, not as a delta.

!

Encrypted reasoning

The thinking isn’t in the logs. You measure the presence (🧠), never the content of the thought.

!

Where Fable was weak

Overthinking the trivial (thinking too much even for simple tasks) and verbosity. The delta isn't "Fable is better at everything".

💡 The golden rule of honesty

Only “think before acting” is a firm delta. State that confidently; treat the rest as context. That discipline is what makes the Track 3 playbook defensible—you inject what you measured, not what you hoped was true.

Fable Sample

7 sessions

Noisy

read/test

Encrypted

presence, not text

Solid delta

think-first only

📊 Module Summary

✓
extract_corpus.py filter by message.model — the entire history; --list shows the models.
✓
The logical turn is the unit — from prompt to prompt; every metric is per logical turn.
✓
Measured, not estimated, metrics — % think-before, tools/turn, read→edit, edit→test.
✓
compare_models.py gives you the table + Δ — and the compare.json for the pipeline.
✓
Real delta: 99% vs. 54% (+45 pts) — and 6.57 vs 7.86 tools/turn (Fable is more efficient).
✓
Honesty about the limits — small sample, opaque reasoning; only think-before is a solid delta.

Next Track:

Track 3 — From Delta to Playbook: turn the number into an injectable rule.