📜 What a playbook is
A playbook is a single file .md actionable. It captures the delta you
measured not as a narrative ("Fable seemed more careful") or a video impression, but as RULES: imperative instructions
that the target model can follow — “think before acting”, “read before editing”. A rule can be injected; an impression can’t.
🎯 What an honest playbook does
- •1 file: fits in context without bloating the session.
- •Imperative: each line is an instruction, not a description.
- •Actual numbers embedded: the +45 points come from YOUR corpus.
- •Honest: says what does NOT transfer and what NOT to copy.
✓ It’s a playbook
- ✓“For a non-trivial task, plan in 2–5 lines before the 1st tool.”
- ✓“Never edit a file you haven’t read in this context.”
- ✓Every rule checkable and followable.
✗ Is NOT a playbook
- ✗“Fable is a very reflective model.” (impression)
- ✗An essay on how the model "thinks".
- ✗Blindly copy Fable, flaws included.
⚙️ make_playbook.py — generates from compare.json
The generator is the make_playbook.py. It has two modes: either you provide the compare.json
already measured in Track 2 (via --from-json, faster), or it measures the history on the fly. Either way, the key point is
the same: the real NUMBERS go in the rule text — not generic impressions.
# reusa o compare já medido (rápido): python make_playbook.py --from-json compare.json \ --out ../playbook/fable-mindset-playbook.md # ou mede o histórico na hora: python make_playbook.py --target claude-opus-4-8 \ --model-fable claude-fable-5
--from-json compare.json
Reuse the measurement already taken by compare_models.py. Same number, without rescanning thousands of sessions.
Numbers injected into the text
O build_playbook() writes "99% against 54% (+45 points)” by reading the values from the metrics dict — not hard-coded constants.
💡 Practical tip
Generate the compare.json once (Path 2), and version it along with the playbook. That way anyone can regenerate the .md with --from-json in seconds, without access to your raw history.
📶 Sort by signal STRENGTH
The order of the rules doesn’t follow the order of the video or the code—it follows measured signal strength. The robust delta is the think-before-acting (+45 pts): comes first. Read-before-edit and test-after come later, as good practices to adopt, with the honest caveat of a small sample — not as a statistically robust delta.
Think before you act firm delta · +45 pts
Fable reasoned in 99% of turns vs. Opus’s 54%. The biggest lever, the one with the most transferable impact — at the top of the list.
⚠ Update: measured later on a large sample (4.892 steps), the honest number is ~85% (not 99%) and the gap drops to +31 pts — and a hidden gap shows up in test-after-edit (41% vs 2%). See the Track 4 · The Real-World Test.
Purposeful tool use density
6.57 vs 7.86 tools/turn: Fable was more economical. Density, not thrashing.
Read before editing best practice
~34% / ~38% in both: modest in both. Discipline is RISING to 100%, not copying.
Close the loop: test/build best practice
Noisy sample (0% / 3%). Include it as a habit to adopt, with the caveat.
⚠️ Small Fable sample (7 sessions)
Treat read-before-edit and test-after-edit as BEST PRACTICES to adopt, not as a firm delta. The make_playbook.py even warns when there are fewer than 30 sessions. The only robust signal here is reasoning-before-acting.
🔄 Fix, don’t copy
Here’s what distinguishes an honest playbook from a naive copy: it reverses Fable’s weak habits instead of replicating them. Reasoning through 99% of turns is great for difficult tasks — and wasteful for a one-line rename. A playbook that copies everything would also import the flaws.
✗ Fable’s weak habit
- ✗Overthinking the trivial — reasoning even for a rename.
- ✗Verbosity — narrates too much of what it’s going to do.
- ✗A plan that turns into rehearsal — planning for too long.
✓ Correction in the playbook
- ✓Reasoning proportional to difficulty — mechanical work goes straight through.
- ✓Act; the summary comes afterward, kept brief, with paths and what was verified.
- ✓The plan fits in a few lines; depth belongs in execution.
🔎 Why this matters
The goal isn’t to become Fable—it’s to improve the target model’s execution. Importing a flaw along with a strength cancels out the gain. Correcting it keeps only what helps: the good pace, without the friction.
🧭 Transferable rules
These are the rules the target model actually adopts—the heart of the playbook. They all converge on one canonical sequence of work.
🧩 The six rules that transfer
- 1.Think first in the non-trivial task (2–5 lines of plan before the 1st tool).
- 2.Purposeful tool use: density, not busyness; independent calls in the same turn.
- 3.Read before editing → 100%: never edit a file you haven't read in this context.
- 4.Close the loop with a test/build and report the REAL result.
- 5.Canonical sequence: understand → plan → read → edit → test → report.
- 6.Stop when you have enough to act — don't re-derive facts that are already established.
✅ The one-line checklist
The entire playbook boils down to a single line you paste at the top of coding tasks. It’s a minimal reminder, with negligible context cost, that carries the canonical sequence and the corrections. This is the final checklist, exactly as it comes out of the generator:
entender → plano curto → ler → editar → testar → relatar (com output) ·
pense-antes-no-não-trivial · ferramenta-com-propósito · read-before-edit→100% ·
teste-depois-de-edit · sem-verbosidade · sem-over-think-no-trivial
top of the coding task
negligible (1 line)
sequence + corrections
inject automatically (3.2)
💡 Practical tip
This checklist is the summary; the full playbook (with the delta table and the “DO NOT copy this” section) is what you inject. In Module 3.2, we'll see how to make this file enter the context automatically at every session.
📜 Module Summary
make_playbook.py injects the actual numbers — from the compare.json or by measuring in real time.Next Module:
3.2 — Injecting into the Model (and Without Your Own Data)