PTENES
Skip to content
TRACK 4

🛡️ The Real Proof

Before believing any number, you need to trust the benchmark. This path shows how we measured three times — 7 local sessions, then an equal-sized sample, then 4,892 steps from the open dataset — and found that a small sample misleads in both directions. And what solutions to apply to measure, inject, and reproduce without fooling yourself.

7 sessions ±unstable sample 1 ~950 steps matched · noisy sample 2 4.892 steps 30 sessions · HF solid sample 3 85% vs 54% +31pp · defensible
3
Modules
18
Topics
~75 min
Duration
Advanced
Level
Track 4 Progress 0%
0 of 0

Track map

Detailed content

4.1 ~25 min

🪤 The Sample Trap

How a small sample misleads in both directions—the journey through three tests to a defensible number.

What it is:

The measurement journey in three rounds: 7 local sessions → matched sample (~950 steps) → 4,892 steps from the open HF dataset.

Why learn:

Each round corrects a bias from the previous one; seeing the sequence teaches you not to trust the first measurement.

Key concepts:

Measure three times; each sample larger; from exciting (misleading) to defensible (solid).

What it is:

7 local Fable sessions / ~950 steps yielded 99% of “think before acting” versus 54% from Opus — a delta of +45pp.

Why learn:

The number is exciting, but it comes from only 7 sessions: it’s the kind of result that looks like proof but is just noise.

Key concepts:

99% vs. 54% (+45 pp); sample of 7 sessions; exciting ≠ reliable.

What it is:

With the size matched (~950 steps each), Opus jumps to 94% in a sample of just 4 sessions.

Why learn:

Capping the size helps with comparisons, BUT 4 sessions is still a small sample: the number fluctuates depending on the slice.

Key concepts:

Matched sample; 4-session slice; unstable; equal size isn’t enough without volume.

What it is:

4.892 steps / 30 sessions from the open dataset Glint-Research/Fable-5-traces: Fable 85% vs Opus 54% (+31pp).

Why learn:

A large sample from both sides is what makes the number defensible—the benchmark worth carrying forward.

Key concepts:

85% vs 54% (+31pp); 4.892 steps / 30 sessions; open dataset; defensible.

What it is:

The small sample INFLATED "think before acting" (99→85) AND HID "test after editing" (locally it showed 0%; the real figure is 41%).

Why learn:

It’s the trail’s central point: a small sample isn’t just wrong on the high side—it can be wrong in either direction.

Key concepts:

Inflate (99→85) and hide (0→41%); the bias goes in two directions; presence, not content.

What it is:

Never conclude from a small sample; balance the sizes before comparing; confirm with a large sample before believing.

Why learn:

It’s the practical distillation of the pitfall—the discipline that separates a number from an opinion.

Key concepts:

Don't conclude too early; balance the sizes; confirm with a large sample; stay skeptical until the volume is there.

View Full Version
4.2 ~25 min

🛠️ Applicable Solutions

The corrected playbook and how to measure, inject, and reproduce properly—honesty over hype.

What it is:

Only two rules transfer with evidence: think before acting (85 vs 54) and close the loop with a test after editing (41 vs 2).

Why learn:

Focus matters more than volume: a two-rule playbook backed by evidence beats a ten-rule one without proof.

Key concepts:

Think-first (85 vs 54); test-after (41 vs 2); only what truly transfers.

What it is:

A hook SessionStart reads one .md separate (editable without changing settings), is fail-open and MERGES into the hooks array — does not overwrite.

{
  "hooks": {
    "SessionStart": [
      { "hooks": [
        { "type": "command",
          "command": "cat ~/.claude/fable-playbook.md 2>/dev/null || true" }
      ] }
    ]
  }
}
Why learn:

Separate the rule into a .md lets you iterate on the playbook without touching settings, and the || true ensures the session never breaks.

Key concepts:

SessionStart; .md editable; fail-open; merge (don't overwrite) the array.

What it is:

Balance the sample (cap by number of steps, e.g., ~950) and warn when the sizes differ too much between models.

Why learn:

That’s exactly what prevents the false +45pp: without balancing, the model with more data looks “better.”

Key concepts:

Cap by steps; warn about divergence; compare only when sizes are close.

What it is:

Download the RAW dump from Glint-Research/Fable-5-traces (are .jsonl from Claude Code, not a flattened chat) via the HF API — without the datasets — and run it in the fable_lib.

BASE=https://huggingface.co/datasets/Glint-Research/Fable-5-traces/resolve/main
curl -sL "$BASE/<arquivo>.jsonl" -o sessao.jsonl
python -m fable_lib.measure sessao.jsonl
Why learn:

It’s how to get a large reference set even without your own data—and preserve the event format the measurement tool needs.

Key concepts:

Raw dump; .jsonl of events; HF API; no datasets; run in fable_lib.

What it is:

The commands that reproduce the measurement (extract_corpus.py, compare_models.py, the HF meter) and save a DATED baseline somewhere durable—not in /tmp.

Why learn:

Without a durable, dated baseline, you have nothing to compare against later—and you lose proof that anything changed.

Key concepts:

extract_corpus / compare_models; HF meter; dated baseline; durable location (not /tmp).

What it is:

Presence ≠ content (reasoning is encrypted in the logs); model weights don’t transfer; the playbook must be iterated with new data.

Why learn:

Knowing the limits is what keeps the method honest—you copy the pace, not the model’s brain.

Key concepts:

Encrypted reasoning; weights don’t transfer; iterate with new data; honesty over hype.

View Full Version
4.3 ~25 min

🧰 Examples of usefulness

Where this is genuinely useful—choosing a model, diagnosing your own, onboarding, proving a change, and cutting through hype.

What it is:

Sonnet thinks only 10% and uses few tools (great for quick, mechanical, low-cost tasks); Opus and Fable for work that requires planning and closing the loop.

Why learn:

The yardstick becomes a routing criterion: you choose the model based on the nature of the task, with the numbers in hand.

Key concepts:

Sonnet 10% / few tools; Opus and Fable to plan+close the loop; model per task.

What it is:

Measure your own logs and see where you fail. E.g., if your Opus tests after editing only 2%, the biggest lever is closing the loop—not "thinking more".

Why learn:

Turns generic advice into a personal diagnosis: you fix the failure YOUR data shows.

Key concepts:

Measure your own logs; find the real failure; lever = where the number is lowest.

What it is:

Inject the good rhythm by default (hook SessionStart or CLAUDE.md) so every dev starts the session with think-before + test-after.

Why learn:

Standardizes good behavior without relying on each person to remember—the playbook becomes culture, not folklore.

Key concepts:

Default via hook/CLAUDE.md; think first + test afterward; don't rely on individual memory.

What it is:

Date the baseline BEFORE, apply the playbook, measure AFTER with a sufficient sample — turns “I think it improved” into a number.

Why learn:

It's the difference between impressions and evidence. Watch for dilution: isolate the new sessions so the effect doesn't disappear in the average.

Key concepts:

Before/after; dated baseline; sufficient sample; isolate new sessions (anti-dilution).

What it is:

The same method measures Codex and open-source models—you just need logs in event format. The field model is the key.

Why learn:

The yardstick isn’t vendor-specific: wherever there are events with model, you measure.

Key concepts:

Codex and open source; event format; field model as a key; method-agnostic.

What it is:

Use the yardstick to require a large sample size before believing any “model X is better than Y” claim.

Why learn:

That’s exactly how we debunked the "+45pp": requiring a sample is what separates marketing from measurement.

Key concepts:

Require a large sample; be skeptical of claims; that’s how the +45pp fell apart.

View Full Version