🛠️ Hands-on
Now that you know where the gold is, you can mine it. This path runs the actual scripts: the debloat that distills a heavy session into a lightweight transcript, and the extract/compare that extracts each model's corpus and measures the Fable vs Opus delta in real numbers.
Track map
Detailed content
🧪 Debloat: Distilling the Transcript
What to throw away, what to keep, the script that does it, and the ~74% that disappear—without losing momentum.
Most of a JSONL session is bloat: tool_result with the echoed output, full file dumps, command output, and base64 attachments. Loading the raw data into the context is wasteful.
Without distillation, any analysis drowns the signal in opaque bytes—and wastes context.
Echoed output + blobs = weight, not behavior; the raw data won’t fit in context.
Preserves your prompts, the assistant’s text, the message.model, 1 line per tool_use (name + short target) and the PRESENCE of reasoning (🧠 marker, since the text comes encrypted).
That’s exactly the gold—the decisions, action order, and cadence—that the playbook will use.
Keep the pace, not the output; mark 🧠 = thought (without revealing what).
Throw out the payloads from tool_result, attachment blobs and harness bookkeeping— usage, uuids, isMeta, sidechain.
It’s what adds bulk without carrying behavioral signal—discarding it doesn’t change the measured pace.
Echoed output and harness metadata are removed; the “what was done” stays.
Runs in the demo_session.jsonl by default; accepts -o (output), --no-thinking e --no-open; prints the size before and after.
It’s the first hands-on tool—you’ll see the lightweight format in practice.
Default = demo; -o chooses the output; shows before → after.
In a typical session, debloating cuts about 74% of the weight — filler made up most of the file.
Sets expectations: the signal fits in a small, readable file.
−74% typical; the remaining ~26% is the gold you analyze.
Each assistant block is a separate LINE; debloat preserves the order, so you see the assistant in several consecutive headers. That's normal.
That’s why the analysis groups by LOGICAL TURN—otherwise the signal gets diluted per line.
Preserve order; multiple headers = one turn; group by human prompt.
📊 Corpus, Numbers, and the Fable vs Opus Delta
Extract the corpus by model, measure the pace using real numbers, compare two models, and interpret the delta honestly.
extract_corpus.py --model claude-fable-5 extracts ALL turns from one model across the entire history (all projects). --list shows the models present.
It’s what separates each model’s corpus—the basis for all comparison.
Filter by message.model; scans entire projects; --list reveals the models.
A logical turn is one human prompt through the next one. Since each block is one line, measuring by line skews the results—which is why the metrics are based on logical turns.
It’s the unit that reveals “how much the model did to answer that prompt”—the actual pace.
Prompt → prompt; groups blocks; every metric is per logical turn.
Logical turns, % that thought first, tools/turn (mean and median), read-before-edit, and test-after-edit — all measured, not estimated by guesswork.
Replacing impressions with numbers is what makes the finding defensible and transferable.
Reasoning presence + tool density + order (read→edit→test).
compare_models.py --a claude-fable-5 --b claude-opus-4-8 prints the side-by-side table and the Δ column; --out compare.json records the result.
The Δ is the playbook's raw material—the difference between the two rhythms.
Two columns + Δ; each row is a metric; saves to compare.json.
Fable-5 thought first in 99% of turns vs. Opus 54% (+45 points). Tools/turn: 6.57 vs 7.86 (Fable is more efficient).
⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.
It’s the strong, transferable finding—it becomes the anchor rule “think before you act.”
Honesty: more tools ≠ better (density vs. thrashing).
Fable's sample is small (7 sessions), which makes read-before-edit/test-after-edit noisy; and the reasoning text is encrypted — presence is measured, not content.
Knowing what NOT to claim is what keeps the course honest and defensible.
Where Fable was weak: overthinking trivial tasks, verbosity.