PTENES
Skip to content
MODULE 2.1

🧪 Debloat: Distilling the Transcript

Before measuring, you need to distill. In this module: why raw logs are large, what the debloat_jsonl.py keeps and what it discards, how to run it, the ~74% that disappear, and the schema gotcha that justifies grouping everything into a logical turn.

6
Topics
25
Minutes
Practical
Level
Hands-on
Type
Raw JSONL gold + overhead mixed debloat filter ✓ prompts + text ✓ message.model ✓ tool_use (1 line) ✓ 🧠 reasoning (presence) ✗ tool_result / blobs ✗ usage / uuids / meta lightweight transcription ~26% of the size
1

🍔 The problem: heavy logs

Open a real JSONL session, and the first thing that jumps out is the weight. Most of it isn’t conversation: it’s tool_result with the output echoed back into the context, full file dumps, command output and base64 attachments. Loading the raw data into context to analyze it is pure waste — you pay tokens for opaque bytes.

⚖️ Where the weight comes from

  • •tool_result — the Read/Bash/Grep output echoed back to the log in full.
  • •File dumps—the complete contents of each opened file.
  • •Command output—stdout/stderr from everything that ran.
  • •Attachments — images/files in base64 (megabytes per line).
{"type":"user","message":{"content":[{"type":"tool_result",
   "content":"…3.412 linhas do arquivo inteiro ecoadas aqui…"}]}}
{"type":"user","message":{"content":[{"type":"tool_result",
   "content":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…"}]}}
Greater weight

echoed tool_result

Dumps

entire files

Attachments

base64

Cost

unnecessary context

2

💎 What debloat KEEPS

The debloat rule is surgical: keep what reveals behavior; discard what only adds weight. What remains is exactly the gold — your prompts, the assistant’s text, the model that wrote each turn, one concise line per tool call, and the presence of reasoning (marked with 🧠, since the text is encrypted in the logs).

✓ What to KEEP (the gold)

  • ✓My user prompts.
  • ✓The assistant's text.
  • ✓The model (message.model).
  • ✓1 line per tool_use (name + short target).
  • ✓THE PRESENCE of reasoning (🧠 marker).

✗ What it does NOT store

  • ✗The tool’s complete output.
  • ✗The contents of the file that was read.
  • ✗The literal text of the thinking (is encrypted).

💡 Why only the PRESENCE of reasoning

The block thinking reaches the logs empty/encrypted — only with the signature. You don’t have the literal thought, but you know that it existed. Debloat records this as a 🧠 marker: enough to measure "thought before acting?" without making up content.

## 👤 humano
refatora o parser e roda os testes

## 🤖 claude-fable-5  🧠
Vou ler o arquivo antes de mexer.
  → Read parser.py
  → Edit parser.py
  → Bash pytest -q
3

🗑️ What debloat DISCARDS

The other side of the rule: everything that is weight without a signal go away. The payloads from tool_result, the attachment blobs and harness accounting — usage, uuids, isMeta, sidechain. None of this changes the pace measured, so it can be dropped without guilt.

✗

tool_result (payload)

The tool’s echoed output. The fact that the tool ran stays (the line for tool_use); the output itself, no.

✗

Attachment blobs

Base64 images and files. Megabytes that say nothing about how the model works.

✗

Harness accounting

usage (tokens), uuids, isMeta, sidechain — internal bookkeeping, not behavior.

🔎 The mental test

For each field, ask: "does this change the answer to 'did the model think first? how many tools did it use? in what order?'". If it doesn't change the rhythm, it's dead weight. That question is the entire logic of the debloat_jsonl.py.

tool_result

payload outside

attachments

base64 blobs

usage/uuid

harness

Criterion

does it change the pace?

4

🐍 The script debloat_jsonl.py

Theory becomes practice with a single command. With no arguments, the debloat_jsonl.py runs in demo_session.jsonl that lives next to it, prints the size before/after and opens the clean transcript — exactly the course flow, so you can check the format the first time.

# roda no demo_session.jsonl (ao lado do script)
python debloat_jsonl.py

# num arquivo seu, escolhendo a saída
python debloat_jsonl.py minha_sessao.jsonl -o limpa.md

# sem a marca de raciocínio, sem abrir o editor
python debloat_jsonl.py s.jsonl --no-thinking --no-open

🎛️ The flags

  • •(sem args) — use the demo_session.jsonl and opens the result.
  • •-o ARQUIVO — choose where to save the lightweight transcript.
  • •--no-thinking — omits the 🧠 reasoning marker.
  • •--no-open — don't try to open it in the editor (useful in a pipeline).

💡 Practical tip

É pure stdlib — no dependencies to install. Start with the demo to see the format; then point to one of your own sessions at ~/.claude/projects/<projeto>/. In automation, always use --no-open.

Default

demo_session.jsonl

-o

chooses output

Prints

before/after

Deps

pure stdlib

5

📉 The result: ~74% reduction

The number the script prints is almost always surprising: in a typical session, debloating cuts about 74% of the weight. Filler really did make up most of the file — what's left, the gold, fits in a small file that's easy to read from beginning to end.

before (raw) 100% afterward (debloat) ~26% −74% bloat
$ python debloat_jsonl.py
antes:  1.84 MB
depois: 0.48 MB   # ~74% menor
escrito em: demo_session.clean.md

💡 What to expect

The exact % varies with how much tool output the session had—sessions full of file reads shrink even more. But the range of ~74% is typical: the signal lives in a small fraction of the file.

Before

100% (raw)

After

~26%

Capped

~74%

Remaining

the gold

6

🧩 The schema gotcha

When you open the lightweight transcript for the first time, one detail seems odd: the assistant appears in several consecutive headers. It’s not a bug. Each assistant block is a Separate LINE in the JSONL, and the debloat preserves the order — so a single response turn becomes several entries. That’s exactly why the analysis (in the next module) groups everything into logical turn.

1
line ## 👤 humano — the prompt
2
line ## 🤖 claude-fable-5 🧠 — thinking
3
line ## 🤖 claude-fable-5 — text
4
line → Read / → Edit / → Bash — tool_use
5
line ## 👤 humano — next prompt → closes the turn

✓ Why it's right

  • ✓Preserves the actual order of events.
  • ✓It doesn't invent turn boundaries too early.
  • ✓Leave grouping to the analysis (logical turn).

✗ The common mistake

  • ✗Thinking “5 headers = 5 responses.”
  • ✗Counting by header inflates the numbers.
  • ✗Measure without grouping by human prompt.
1 block

= 1 line

Order

preserved

Multiple ## 🤖

= 1 turn

Group

logical turn

🧪 Module Summary

✓
Raw logs are bulky — echoed tool_result, dumps, command output, and base64 dominate the file.
✓
Debloating KEEPS the gold — prompts, text, model, 1 line per tool_use, and the 🧠 reasoning marker.
✓
AND DISCARD the weight — tool_result payloads, blobs, and harness accounting.
✓
debloat_jsonl.py runs in the demo by default — flags -o, --no-thinking, --no-open; pure stdlib.
✓
~74% reduction — the bulk was most of it; the signal fits in a small file.
✓
1 block = 1 line — multiple assistant headers are normal; that's why they're grouped into a logical turn.

Next Module:

2.2 — Corpus, Numbers, and the Fable vs. Opus Delta