PTENES
Skip to content
TRACK 2

🛠️ Hands-on

Now that you know where the gold is, you can mine it. This path runs the actual scripts: the debloat that distills a heavy session into a lightweight transcript, and the extract/compare that extracts each model's corpus and measures the Fable vs Opus delta in real numbers.

session.jsonl raw / heavy 100% debloat distills the gold transcription lightweight / readable ~26% logical turn · thought first? tools / turn read-before-edit Δ Fable vs Opus
2
Modules
12
Topics
~55 min
Duration
Practical
Level
Track 2 Progress 0%
0 of 0

Track map

Detailed content

2.1 ~25 min

🧪 Debloat: Distilling the Transcript

What to throw away, what to keep, the script that does it, and the ~74% that disappear—without losing momentum.

What it is:

Most of a JSONL session is bloat: tool_result with the echoed output, full file dumps, command output, and base64 attachments. Loading the raw data into the context is wasteful.

Why learn:

Without distillation, any analysis drowns the signal in opaque bytes—and wastes context.

Key concepts:

Echoed output + blobs = weight, not behavior; the raw data won’t fit in context.

What it is:

Preserves your prompts, the assistant’s text, the message.model, 1 line per tool_use (name + short target) and the PRESENCE of reasoning (🧠 marker, since the text comes encrypted).

Why learn:

That’s exactly the gold—the decisions, action order, and cadence—that the playbook will use.

Key concepts:

Keep the pace, not the output; mark 🧠 = thought (without revealing what).

What it is:

Throw out the payloads from tool_result, attachment blobs and harness bookkeeping— usage, uuids, isMeta, sidechain.

Why learn:

It’s what adds bulk without carrying behavioral signal—discarding it doesn’t change the measured pace.

Key concepts:

Echoed output and harness metadata are removed; the “what was done” stays.

What it is:

Runs in the demo_session.jsonl by default; accepts -o (output), --no-thinking e --no-open; prints the size before and after.

Why learn:

It’s the first hands-on tool—you’ll see the lightweight format in practice.

Key concepts:

Default = demo; -o chooses the output; shows before → after.

What it is:

In a typical session, debloating cuts about 74% of the weight — filler made up most of the file.

Why learn:

Sets expectations: the signal fits in a small, readable file.

Key concepts:

−74% typical; the remaining ~26% is the gold you analyze.

What it is:

Each assistant block is a separate LINE; debloat preserves the order, so you see the assistant in several consecutive headers. That's normal.

Why learn:

That’s why the analysis groups by LOGICAL TURN—otherwise the signal gets diluted per line.

Key concepts:

Preserve order; multiple headers = one turn; group by human prompt.

View Full Version
2.2 ~30 min

📊 Corpus, Numbers, and the Fable vs Opus Delta

Extract the corpus by model, measure the pace using real numbers, compare two models, and interpret the delta honestly.

What it is:

extract_corpus.py --model claude-fable-5 extracts ALL turns from one model across the entire history (all projects). --list shows the models present.

Why learn:

It’s what separates each model’s corpus—the basis for all comparison.

Key concepts:

Filter by message.model; scans entire projects; --list reveals the models.

What it is:

A logical turn is one human prompt through the next one. Since each block is one line, measuring by line skews the results—which is why the metrics are based on logical turns.

Why learn:

It’s the unit that reveals “how much the model did to answer that prompt”—the actual pace.

Key concepts:

Prompt → prompt; groups blocks; every metric is per logical turn.

What it is:

Logical turns, % that thought first, tools/turn (mean and median), read-before-edit, and test-after-edit — all measured, not estimated by guesswork.

Why learn:

Replacing impressions with numbers is what makes the finding defensible and transferable.

Key concepts:

Reasoning presence + tool density + order (read→edit→test).

What it is:

compare_models.py --a claude-fable-5 --b claude-opus-4-8 prints the side-by-side table and the Δ column; --out compare.json records the result.

Why learn:

The Δ is the playbook's raw material—the difference between the two rhythms.

Key concepts:

Two columns + Δ; each row is a metric; saves to compare.json.

What it is:

Fable-5 thought first in 99% of turns vs. Opus 54% (+45 points). Tools/turn: 6.57 vs 7.86 (Fable is more efficient).

⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.

Why learn:

It’s the strong, transferable finding—it becomes the anchor rule “think before you act.”

Key concepts:

Honesty: more tools ≠ better (density vs. thrashing).

What it is:

Fable's sample is small (7 sessions), which makes read-before-edit/test-after-edit noisy; and the reasoning text is encrypted — presence is measured, not content.

Why learn:

Knowing what NOT to claim is what keeps the course honest and defensible.

Key concepts:

Where Fable was weak: overthinking trivial tasks, verbosity.

View Full Version