🎯 The two evidence-backed levers
Of everything the delta indicated, only two levers survive the large sample and actually transfer: thinking before acting (Fable 85% vs Opus 54%) e close the loop with a test after editing (Fable 41% vs Opus 2%). The rest is noise or a tie.
✓ Worth copying
- ✓Think before acting on non-trivial tasks— 85 vs 54, a solid gap at 4.9k steps.
- ✓Test after editing — 41 vs 2, the biggest hidden lever.
- ✓Closed loop: edited → ran → confirmed.
✗ Is NOT an advantage
- ✗“Save on tools”: Fable uses 10,45/turn vs. Opus 7.96 — does MORE.
- ✗read-before-edit: ~37–40% for both — tie.
- ✗The Test 1 "+45pp" — it was a sample of 7 sessions.
💡 Practical tip
Resist the temptation to bundle ten rules. Two well-supported levers beat ten guesses. The corrected playbook fits in one paragraph: think before acting on non-trivial tasks; after editing, run the test.
🪝 Inject the focused rule
The corrected rule is applied via a hook SessionStart that reads a .md separate — so you can edit the rule
without changing settings. The script is fail-open (if the file disappears, the session continues) and the hook BLENDS in the array of
SessionStart existing — doesn't overwrite the hooks you already have.
// ~/.claude/settings.json — MESCLE no array existente, não sobrescreva { "hooks": { "SessionStart": [ { "matcher": "startup", "hooks": [ { "type": "command", // lê regra-focada.md e devolve via additionalContext; fail-open "command": "bash ~/projetos/fablelite/hooks/inject-regra.sh" } ] } ] } }
🔎 Three properties that matter
- . separate .md — edit the rule without touching the settings JSON.
- fail-open — without the file, prints
{"continue":true}and exits with 0. - merges — adds to the array of
SessionStart; preserves your other hooks.
💡 Practical tip
Keep the rule short: the TWO levers. A additionalContext concise content gets read; a long manifesto gets ignored. Point to another file via FABLE_REGRA= when you want to test variations.
⚖️ Measure honestly: balance the sample
The correction that prevents the +45pp false: cap the sample by number of steps (e.g., ~950 from each side) before comparing, and warn when the sizes diverge significantly. Without this, 4,892 Fable steps versus 42k Opus steps can produce any narrative you want.
| Test | Sample | Thinking (Fable vs Opus) | Verdict |
|---|---|---|---|
| 1 — local | 7 sessions / ~950 | 99% vs. 54% (+45 pp) | misleading |
| 2 — capped | ~950 each (4 sessions) | Opus jumps to 94% | unstable |
| 3 — HF | 4.892 steps / 30 sessions | 85% vs 54% (+31pp) | solid |
💡 Practical tip
Capping helps, but even an equal small sample is still noise (Test 2 with 4 sessions was unstable). The goal is to cap it AND have enough steps on both sides — Test 3 (4.9k) is what gave us the defensible number of 85 vs 54.
🤗 Reliable reference via HF (without the datasets library)
The large sample’s anchor is the raw dump from Glint-Research/Fable-5-traces. They are files .jsonl from Claude Code (events, not
flattened chat). Download them directly through the HF API — without installing the library datasets — and run it in the fable_lib, the same meter.
# 1. lista os .jsonl crus do repo via API do HF (nada de `pip install datasets`) REPO=Glint-Research/Fable-5-traces curl -s "https://huggingface.co/api/datasets/$REPO/tree/main?recursive=1" \ | python3 -c 'import sys,json; [print(f["path"]) for f in json.load(sys.stdin) if f["path"].endswith(".jsonl")]' \ > files.txt # 2. baixa cada arquivo cru (resolve/main = o blob, não a página) mkdir -p corpus_fable_hf while read f; do curl -sL "https://huggingface.co/datasets/$REPO/resolve/main/$f" \ -o "corpus_fable_hf/$(basename "$f")" done < files.txt # 3. roda o MESMO medidor no fable_lib sobre os .jsonl crus python3 -m fable_lib.measure --glob "corpus_fable_hf/*.jsonl" --cap-steps 950
Fable-5-traces (HF)
. event .jsonl files
lib datasets
fable_lib
💡 Practical tip
Use resolve/main/<path> to get the raw blob (the URL blob/ brings the page's HTML). Since there are already .jsonl from Claude Code, the fable_lib reads directly — no need to flatten chat or map the schema.
🔁 Replay everything (and set the baseline)
The real proof is reproducible: run the same commands and save a DATED baseline in one place durable (not /tmp, which disappears).
Without the saved baseline, you can't tell later whether the playbook moved the number.
extract_corpus.py
Reads the .jsonl and produces the event corpus by model.
the HF meter
Run the same exercise on the 4.892 steps of Fable-5-traces (reference side).
compare_models.py
Compares Fable vs Opus, already capped, and records the dated, durable baseline.
✓ Baseline well preserved
- ✓Dated:
baseline-2026-06-15.json. - ✓Durable: in the project repo/folder, versioned.
- ✓With the recorded sample (steps/sessions per side).
✗ Baseline that evaporates
- ✗In
/tmp— disappears on the next boot. - ✗No date—you can’t compare “before/after.”
- ✗Without the sample size recorded.
📊 Why date and save it
A dated baseline turns "I think it improved" into a number. Measure BEFORE, apply the playbook, measure AFTER — with a sufficient sample — and the delta speaks for itself. Without the BEFORE saved, there's no proof.
🧭 What does NOT transfer
Honesty over hype, to wrap up: what we measure is presence, not content (the text of the thinking comes encrypted in the logs).
The model weights don't carry over — you imitate the rhythm; you don't clone Fable. And the playbook isn't forever: iterate with new data.
✓ What you gain
- ✓The good PACE measured (think + test).
- ✓One yardstick for measuring any model using logs.
- ✓A baseline to prove change with numbers.
✗ What does NOT transfer
- ✗Reasoning content—presence ≠ content (encrypted).
- ✗The model weights—the power doesn't change hands.
- ✗“Become Fable” — the identity doesn’t come with it.
⚠️ Presence ≠ content
We measure WHETHER the model thought, not WHAT it thought — the thinking comes encrypted. The small sample was misleading in both directions: it inflated "thinking" (99→85) AND hid "testing after editing" (0→41%). The larger benchmark corrects this; it doesn’t read minds.
💡 Practical tip
Treat the playbook as a living document: as you accumulate sessions, remeasure against the dated baseline and update the levers. Honesty over hype — only promote a number when the large sample supports it.
🛠️ Module Summary
SessionStart reads one .md, fail-open, MERGE in the array..jsonl raw data from the Fable-5-traces through the API and run it in the fable_lib.extract_corpus.py + compare_models.py + dated, durable baseline.The yardstick, in one sentence:
Measure in a balanced way, inject only what has evidence behind it, and prove it with a baseline. Honesty over hype.