PTENES
Skip to content
MODULE 4.2

🛠️ Applicable Solutions

Once the sample trap is out of the way, what really remains? Here are the two levers with solid evidence, the corrected playbook, and how measure honestly, inject the rule e reproduce the experiment—including downloading the raw reference from Hugging Face and pinning a dated baseline.

6
Topics
~28
Minutes
Applied
Level
Practice
Type
① measure honestlybalanced sample ② inject the ruleSessionStart hook ③ reproduce dated baseline new logs feed back in — remeasure against the baseline
1

🎯 The two evidence-backed levers

Of everything the delta indicated, only two levers survive the large sample and actually transfer: thinking before acting (Fable 85% vs Opus 54%) e close the loop with a test after editing (Fable 41% vs Opus 2%). The rest is noise or a tie.

✓ Worth copying

  • ✓Think before acting on non-trivial tasks— 85 vs 54, a solid gap at 4.9k steps.
  • ✓Test after editing — 41 vs 2, the biggest hidden lever.
  • ✓Closed loop: edited → ran → confirmed.

✗ Is NOT an advantage

  • ✗“Save on tools”: Fable uses 10,45/turn vs. Opus 7.96 — does MORE.
  • ✗read-before-edit: ~37–40% for both — tie.
  • ✗The Test 1 "+45pp" — it was a sample of 7 sessions.

💡 Practical tip

Resist the temptation to bundle ten rules. Two well-supported levers beat ten guesses. The corrected playbook fits in one paragraph: think before acting on non-trivial tasks; after editing, run the test.

2

🪝 Inject the focused rule

The corrected rule is applied via a hook SessionStart that reads a .md separate — so you can edit the rule without changing settings. The script is fail-open (if the file disappears, the session continues) and the hook BLENDS in the array of SessionStart existing — doesn't overwrite the hooks you already have.

// ~/.claude/settings.json — MESCLE no array existente, não sobrescreva
{
  "hooks": {
    "SessionStart": [
      {
        "matcher": "startup",
        "hooks": [
          { "type": "command",
            // lê regra-focada.md e devolve via additionalContext; fail-open
            "command": "bash ~/projetos/fablelite/hooks/inject-regra.sh" }
        ]
      }
    ]
  }
}

🔎 Three properties that matter

  • . separate .md — edit the rule without touching the settings JSON.
  • fail-open — without the file, prints {"continue":true} and exits with 0.
  • merges — adds to the array of SessionStart; preserves your other hooks.

💡 Practical tip

Keep the rule short: the TWO levers. A additionalContext concise content gets read; a long manifesto gets ignored. Point to another file via FABLE_REGRA= when you want to test variations.

3

⚖️ Measure honestly: balance the sample

The correction that prevents the +45pp false: cap the sample by number of steps (e.g., ~950 from each side) before comparing, and warn when the sizes diverge significantly. Without this, 4,892 Fable steps versus 42k Opus steps can produce any narrative you want.

Test Sample Thinking (Fable vs Opus) Verdict
1 — local 7 sessions / ~950 99% vs. 54% (+45 pp) misleading
2 — capped ~950 each (4 sessions) Opus jumps to 94% unstable
3 — HF 4.892 steps / 30 sessions 85% vs 54% (+31pp) solid

💡 Practical tip

Capping helps, but even an equal small sample is still noise (Test 2 with 4 sessions was unstable). The goal is to cap it AND have enough steps on both sides — Test 3 (4.9k) is what gave us the defensible number of 85 vs 54.

4

🤗 Reliable reference via HF (without the datasets library)

The large sample’s anchor is the raw dump from Glint-Research/Fable-5-traces. They are files .jsonl from Claude Code (events, not flattened chat). Download them directly through the HF API — without installing the library datasets — and run it in the fable_lib, the same meter.

# 1. lista os .jsonl crus do repo via API do HF (nada de `pip install datasets`)
REPO=Glint-Research/Fable-5-traces
curl -s "https://huggingface.co/api/datasets/$REPO/tree/main?recursive=1" \
  | python3 -c 'import sys,json; [print(f["path"]) for f in json.load(sys.stdin) if f["path"].endswith(".jsonl")]' \
  > files.txt

# 2. baixa cada arquivo cru (resolve/main = o blob, não a página)
mkdir -p corpus_fable_hf
while read f; do
  curl -sL "https://huggingface.co/datasets/$REPO/resolve/main/$f" \
    -o "corpus_fable_hf/$(basename "$f")"
done < files.txt

# 3. roda o MESMO medidor no fable_lib sobre os .jsonl crus
python3 -m fable_lib.measure --glob "corpus_fable_hf/*.jsonl" --cap-steps 950
Source

Fable-5-traces (HF)

Format

. event .jsonl files

Without

lib datasets

Runs on

fable_lib

💡 Practical tip

Use resolve/main/<path> to get the raw blob (the URL blob/ brings the page's HTML). Since there are already .jsonl from Claude Code, the fable_lib reads directly — no need to flatten chat or map the schema.

5

🔁 Replay everything (and set the baseline)

The real proof is reproducible: run the same commands and save a DATED baseline in one place durable (not /tmp, which disappears). Without the saved baseline, you can't tell later whether the playbook moved the number.

1

extract_corpus.py

Reads the .jsonl and produces the event corpus by model.

2

the HF meter

Run the same exercise on the 4.892 steps of Fable-5-traces (reference side).

3

compare_models.py

Compares Fable vs Opus, already capped, and records the dated, durable baseline.

✓ Baseline well preserved

  • ✓Dated: baseline-2026-06-15.json.
  • ✓Durable: in the project repo/folder, versioned.
  • ✓With the recorded sample (steps/sessions per side).

✗ Baseline that evaporates

  • ✗In /tmp — disappears on the next boot.
  • ✗No date—you can’t compare “before/after.”
  • ✗Without the sample size recorded.

📊 Why date and save it

A dated baseline turns "I think it improved" into a number. Measure BEFORE, apply the playbook, measure AFTER — with a sufficient sample — and the delta speaks for itself. Without the BEFORE saved, there's no proof.

6

🧭 What does NOT transfer

Honesty over hype, to wrap up: what we measure is presence, not content (the text of the thinking comes encrypted in the logs). The model weights don't carry over — you imitate the rhythm; you don't clone Fable. And the playbook isn't forever: iterate with new data.

✓ What you gain

  • ✓The good PACE measured (think + test).
  • ✓One yardstick for measuring any model using logs.
  • ✓A baseline to prove change with numbers.

✗ What does NOT transfer

  • ✗Reasoning content—presence ≠ content (encrypted).
  • ✗The model weights—the power doesn't change hands.
  • ✗“Become Fable” — the identity doesn’t come with it.

⚠️ Presence ≠ content

We measure WHETHER the model thought, not WHAT it thought — the thinking comes encrypted. The small sample was misleading in both directions: it inflated "thinking" (99→85) AND hid "testing after editing" (0→41%). The larger benchmark corrects this; it doesn’t read minds.

💡 Practical tip

Treat the playbook as a living document: as you accumulate sessions, remeasure against the dated baseline and update the levers. Honesty over hype — only promote a number when the large sample supports it.

🛠️ Module Summary

✓
Two well-supported levers — thinking (85 vs 54) + testing after editing (41 vs 2). Only these transfer.
✓
Inject the focused rule — hook SessionStart reads one .md, fail-open, MERGE in the array.
✓
Measure honestly — cap the sample (~950) and flag when the sizes diverge; that’s what killed the false +45pp.
✓
Reference via HF without datasets — download the .jsonl raw data from the Fable-5-traces through the API and run it in the fable_lib.
✓
Reproduce everything — extract_corpus.py + compare_models.py + dated, durable baseline.
✓
What DOESN’T transfer — presence ≠ content; weights don't transfer; iterate with new data.

The yardstick, in one sentence:

Measure in a balanced way, inject only what has evidence behind it, and prove it with a baseline. Honesty over hype.