🍔 The bloat
Most of a log’s weight is bloat: bytes that say nothing about how the model works.
The champion is the tool_result — each tool’s output echoed back. Along with it come the full file dumps, command output, attachment blobs (base64), and harness accounting (usage, sidechain, isMeta).
🗑️ What is fluff (discard)
- ✗
tool_result— the tool output, echoed in full. - ✗Full file dumps (Read contents).
- ✗Command output (long stdout/stderr).
- ✗Attachment blobs (base64 images, PDFs).
- ✗Harness accounting:
usage,sidechain,isMeta.
💡 Practical tip
Discard the tool_result don’t delete the decision of using the tool — the tool_use (gold) remains. You discard the response, not the request.
echoed output
entire files
base64
usage/meta
💎 The gold
The gold is small but dense: your prompts, o assistant text,
a reasoning presence, a sequence of tool_use
and the timestamps. Put it all together and you have the material that reveals how each model really works.
💎 What is gold (keep)
- ✓Your prompts—what you asked for.
- ✓The assistant's text—the visible response.
- ✓Whether reasoning is present—whether there was thinking or not.
- ✓The sequence of
tool_use— the order of actions. - ✓The timestamps—the cadence over time.
🧭 The rule of thumb
Fat is echoed output and opaque bytes. Gold is decision, order, and cadence. If the byte counts what the model chose to do, is gold. If it counts what the tool answered, is fluff.
the request
the speech
the sequence
the cadence
🔐 The MYTH of mineable reasoning
This is the honest finding that defines the course. Intuition says: "if there’s reasoning in the logs, I’ll read the model’s thoughts".
Wrong. The thinking comes
EMPTY or encrypted — only the signature survives.
You no mines the model's literal thinking.
{"type":"thinking",
"thinking":"", # vazio
"signature":"EqoBCkgI...assinado..." # só prova que pensou
}
🧨 Debunk the myth
What you CAN'T extract from the logs:
- ✗The literal contents of the reasoning (the words the model "thought").
- ✗The internal step-by-step inference chain.
✅ What you actually mine
- •A presence of reasoning (there was a thinking block, attested by the signature).
- •O visible text that the model chose to show.
- •A tool sequence — the order in which it acted.
✓ mineable
✓ mineable
✓ mineable
✗ encrypted
🎵 Why this is still powerful
If literal thinking is out of reach, why is it still worth it? Because the pace is what transfers. The presence of reasoning + tool cadence + action order reveal as a model approaches the work—and that is exactly what becomes a playbook rule.
Reasoning presence
“Does this model think before acting in 99% of turns?” is an answerable question — and transferable as a rule: think before you act.
Tool cadence
How many tools per turn? Dense or economical? Reveals discipline (or thrashing).
Order of actions
Do you read before editing? Test after editing? The sequence is a signature of good practice.
💡 The key idea
You don’t need the literal thought process to copy the habit. "Think before acting" is a simple instruction — and the log proves that Fable follows it 99% of the time. That’s enough to make it a rule.
⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.
think-first
tools/turn
read→edit→test
transferable
📉 How much you can shrink it
In practice, removing the fat shrinks a typical session by about 74%. It’s not a magic number — it directly reflects that the bulk really is most of the file. The gold that remains fits in a lightweight, readable transcript.
📊 The real number
reduction in a typical session after debloating.
The bar shows the fraction that was filler. What's left (26%) is gold.
💡 Why it matters
A transcript that is 74% shorter is cheaper to read, analyze, and pass to the next stage. And it’s honest: it shows exactly what the model decided, without the noise from tool output.
~74%
the gold (~26%)
lightweight transcription
easy to analyze
🔏 Ethics and privacy
One reminder that must not be missed: the logs have YOUR code and YOUR data. File paths, source snippets, sometimes secrets pasted into a prompt. Treat the corpus as sensitive e draft before sharing.
✓ Best practices
- ✓Treat the corpus as personal data.
- ✓Redact secrets/keys before exporting.
- ✓Share only the gold (rhythm), not the data.
✗ Avoid
- ✗Paste raw logs somewhere public.
- ✗Assume it’s “just metadata.”
- ✗Upload the corpus without reviewing the content.
🔒 Remember
The course goal is to mine pace, not data. Whenever possible, derive the metrics and discard the raw content—you almost never need the original code to measure behavior.
your code/data
before sharing
pace, not data
the raw data
⛏️ Module Summary
Next Track:
Track 2 — Hands-On: debloating in practice, corpus, and the Fable vs. Opus delta.