🛡️ The Real Proof
Before believing any number, you need to trust the benchmark. This path shows how we measured three times — 7 local sessions, then an equal-sized sample, then 4,892 steps from the open dataset — and found that a small sample misleads in both directions. And what solutions to apply to measure, inject, and reproduce without fooling yourself.
Track map
Detailed content
🪤 The Sample Trap
How a small sample misleads in both directions—the journey through three tests to a defensible number.
The measurement journey in three rounds: 7 local sessions → matched sample (~950 steps) → 4,892 steps from the open HF dataset.
Each round corrects a bias from the previous one; seeing the sequence teaches you not to trust the first measurement.
Measure three times; each sample larger; from exciting (misleading) to defensible (solid).
7 local Fable sessions / ~950 steps yielded 99% of “think before acting” versus 54% from Opus — a delta of +45pp.
The number is exciting, but it comes from only 7 sessions: it’s the kind of result that looks like proof but is just noise.
99% vs. 54% (+45 pp); sample of 7 sessions; exciting ≠ reliable.
With the size matched (~950 steps each), Opus jumps to 94% in a sample of just 4 sessions.
Capping the size helps with comparisons, BUT 4 sessions is still a small sample: the number fluctuates depending on the slice.
Matched sample; 4-session slice; unstable; equal size isn’t enough without volume.
4.892 steps / 30 sessions from the open dataset Glint-Research/Fable-5-traces: Fable 85% vs Opus 54% (+31pp).
A large sample from both sides is what makes the number defensible—the benchmark worth carrying forward.
85% vs 54% (+31pp); 4.892 steps / 30 sessions; open dataset; defensible.
The small sample INFLATED "think before acting" (99→85) AND HID "test after editing" (locally it showed 0%; the real figure is 41%).
It’s the trail’s central point: a small sample isn’t just wrong on the high side—it can be wrong in either direction.
Inflate (99→85) and hide (0→41%); the bias goes in two directions; presence, not content.
Never conclude from a small sample; balance the sizes before comparing; confirm with a large sample before believing.
It’s the practical distillation of the pitfall—the discipline that separates a number from an opinion.
Don't conclude too early; balance the sizes; confirm with a large sample; stay skeptical until the volume is there.
🛠️ Applicable Solutions
The corrected playbook and how to measure, inject, and reproduce properly—honesty over hype.
Only two rules transfer with evidence: think before acting (85 vs 54) and close the loop with a test after editing (41 vs 2).
Focus matters more than volume: a two-rule playbook backed by evidence beats a ten-rule one without proof.
Think-first (85 vs 54); test-after (41 vs 2); only what truly transfers.
A hook SessionStart reads one .md separate (editable without changing settings), is fail-open and MERGES into the hooks array — does not overwrite.
{
"hooks": {
"SessionStart": [
{ "hooks": [
{ "type": "command",
"command": "cat ~/.claude/fable-playbook.md 2>/dev/null || true" }
] }
]
}
}
Separate the rule into a .md lets you iterate on the playbook without touching settings, and the || true ensures the session never breaks.
SessionStart; .md editable; fail-open; merge (don't overwrite) the array.
Balance the sample (cap by number of steps, e.g., ~950) and warn when the sizes differ too much between models.
That’s exactly what prevents the false +45pp: without balancing, the model with more data looks “better.”
Cap by steps; warn about divergence; compare only when sizes are close.
Download the RAW dump from Glint-Research/Fable-5-traces (are .jsonl from Claude Code, not a flattened chat) via the HF API — without the datasets — and run it in the fable_lib.
BASE=https://huggingface.co/datasets/Glint-Research/Fable-5-traces/resolve/main
curl -sL "$BASE/<arquivo>.jsonl" -o sessao.jsonl
python -m fable_lib.measure sessao.jsonl
It’s how to get a large reference set even without your own data—and preserve the event format the measurement tool needs.
Raw dump; .jsonl of events; HF API; no datasets; run in fable_lib.
The commands that reproduce the measurement (extract_corpus.py, compare_models.py, the HF meter) and save a DATED baseline somewhere durable—not in /tmp.
Without a durable, dated baseline, you have nothing to compare against later—and you lose proof that anything changed.
extract_corpus / compare_models; HF meter; dated baseline; durable location (not /tmp).
Presence ≠ content (reasoning is encrypted in the logs); model weights don’t transfer; the playbook must be iterated with new data.
Knowing the limits is what keeps the method honest—you copy the pace, not the model’s brain.
Encrypted reasoning; weights don’t transfer; iterate with new data; honesty over hype.
🧰 Examples of usefulness
Where this is genuinely useful—choosing a model, diagnosing your own, onboarding, proving a change, and cutting through hype.
Sonnet thinks only 10% and uses few tools (great for quick, mechanical, low-cost tasks); Opus and Fable for work that requires planning and closing the loop.
The yardstick becomes a routing criterion: you choose the model based on the nature of the task, with the numbers in hand.
Sonnet 10% / few tools; Opus and Fable to plan+close the loop; model per task.
Measure your own logs and see where you fail. E.g., if your Opus tests after editing only 2%, the biggest lever is closing the loop—not "thinking more".
Turns generic advice into a personal diagnosis: you fix the failure YOUR data shows.
Measure your own logs; find the real failure; lever = where the number is lowest.
Inject the good rhythm by default (hook SessionStart or CLAUDE.md) so every dev starts the session with think-before + test-after.
Standardizes good behavior without relying on each person to remember—the playbook becomes culture, not folklore.
Default via hook/CLAUDE.md; think first + test afterward; don't rely on individual memory.
Date the baseline BEFORE, apply the playbook, measure AFTER with a sufficient sample — turns “I think it improved” into a number.
It's the difference between impressions and evidence. Watch for dilution: isolate the new sessions so the effect doesn't disappear in the average.
Before/after; dated baseline; sufficient sample; isolate new sessions (anti-dilution).
The same method measures Codex and open-source models—you just need logs in event format. The field model is the key.
The yardstick isn’t vendor-specific: wherever there are events with model, you measure.
Codex and open source; event format; field model as a key; method-agnostic.
Use the yardstick to require a large sample size before believing any “model X is better than Y” claim.
That’s exactly how we debunked the "+45pp": requiring a sample is what separates marketing from measurement.
Require a large sample; be skeptical of claims; that’s how the +45pp fell apart.