🧪 The three tests, in order
The Fable vs. Opus delta didn’t come from a single measurement—it came from a three-part journey. First, a handful of local sessions. Then an attempt to balance the sample size. Finally, an open dataset with thousands of steps. Each round produced a different number — and that difference is precisely what teaches you not to trust the first result.
7 local sessions
The first sample: ~950 steps from local Fable. Exciting, but tiny—and that’s where the trap began.
Equal sample (~950 each)
We tried to balance them by capping the number of steps. It helps—but if the other side is still tiny, it’s still noise.
4.892 steps from HF
Open dataset Fable-5-traces: 30 sessions. Large samples on both sides = a defensible number.
💡 Why measure three times?
Because the first number was too polished. When a result seems too good to be true, the first question isn’t "wow, how much?" — it’s "how many sessions did that come from?". Repeating it with more data is what separates a finding from an illusion.
1️⃣ Test 1 — the tiny sample
The first sample was 7 sessions / 950 steps from local Fable. The number came out round and exciting: 99% think in Fable versus 54% in Opus — a delta of +45pp. It looks like definitive proof. But 99% of a sample of 7 sessions is exactly the kind of number that should raise the alert, not balancing the books.
✓ What the sample seemed to show
- ✓Fable thinks before acting 99% of the time.
- ✓Opus thinks 54% — a chasm of +45pp.
- ✓“Done, it’s proven: Fable is much better.”
✗ What the sample of 7 concealed
- ✗7 sessions do not represent typical behavior.
- ✗99% almost never survives more data (dropped to 85%).
- ✗The +45pp was a sample-size artifact, not a fact.
🔎 The 99% pitfall
Such an extreme value (99%) with a small n is a classic sign of sample overfitting: few sessions, all similar, and the statistic "saturates." The right number isn't zero or 99% — it's the 85% that only shows up when you add thousands of steps to the count.
⚖️ Test 2 — equal sample
Skeptical of the +45pp, we tried the obvious: balance the sizes. We capped both sides at ~950 steps each. Surprising result — Opus jumped to 94%. But this Opus sample came from just 4 sessions. Capping the size helped remove volume bias, but traded one problem for another: 4 sessions is still noise, and the number became unstable.
📊 What we learned from test 2
- Balancing size is necessary — comparing 950 vs 42k steps will always distort the results.
- But that's not enough — if the balanced side has 4 sessions, you’ve only swapped one bias for variance.
- 94% from 4 sessions ≠ truth — it's the same error as test 1, now from the other side.
💡 Practical tip
Balancing the sample solves it a of the two threats (volume bias). The other one — small-sample variance — only disappears with more sessions. Balance the size e ensure there’s enough volume on both sides before drawing a conclusion.
📊 Test 3 — 4,892 HF steps
The answer came from the open dataset Glint-Research/Fable-5-traces: 4.892 steps
across 30 sessions from Fable, compared against Opus's broader baseline. Now “think before acting” has stabilized at
Fable 85% vs Opus 54% (+31pp). It’s not the euphoric +45pp or the inverted 94% — it’s the number that
survives at scale. A large sample on both sides is what makes the comparison defensible.
| Test | Sample | Sessions | % thinks (Fable vs Opus) | Reading |
|---|---|---|---|---|
| 1️⃣ local | ~950 steps | 7 sessions | 99% vs. 54% (+45pp) | misleading |
| ⚖️ equal | ~950 each | 4 sessions (Opus) | Opus jumps to 94% | unstable |
| 📊 HF | 4.892 steps | 30 sessions | 85% vs 54% (+31pp) | solid |
Fable-5-traces (HF)
4.892 / 262 turns
30
85 vs 54 = +31pp
🔀 Deception in TWO directions
This is the module’s central point. A small sample isn’t just wrong "on the high side" — it’s wrong in the two directions at the same time. In our case, it inflated the “think before acting” (showed 99%, the real figure is 85%) and at the same time hid a real advantage: the "test after editing," which the local version marked in 0%, in the large sample is 41%. In other words: the small sample overestimated one dimension and underestimated the other.
↑ Direction 1 — inflated
- ↑“Think before acting”: the local setup said 99%.
- ↑Large sample corrects to 85%.
- ↑The excess turned into a false +45pp over Opus.
↓ Direction 2 — hidden
- ↓“Test after editing”: the local setup marked 0%.
- ↓In the large sample, Fable tests 41% (Opus 2%).
- ↓A real advantage was hidden in the small sample.
💡 Remember the honesty of the measurement
The thinking comes encrypted in the logs — so we measure the presence of the reasoning, not the content. Even so, its presence is enough to show that the small sample was skewed in both directions: it inflated areas where there was little and erased areas where there was a lot.
📐 The golden rule
It all boils down to three commandments. Never conclude of a small sample. Balance the sizes before comparing (cap by number of steps). And confirm with a large sample before believing it. Following this standard is how the euphoric +45pp became a defensible +31pp — and how the hidden advantage of "testing" (0→41%) finally emerged.
✓ Do it this way
- ✓Balance the sample (e.g., ~950 steps per side).
- ✓Require volume on both sides, not just one.
- ✓Confirm with a large dataset before believing it.
✗ Never do this
- ✗Closing the account with just 7 sessions because the number looks good.
- ✗Compare 950 steps against 42k without balancing.
- ✗Believe the 99% (or the 94%) without repeating the test with more data.
🔎 The yardstick in one sentence
Before believing that “model X is better than Y,” ask: how many sessions this came from, and were both sides balanced? If the answer is “a few” or “no,” you don’t have a number yet—you have an impression.
🪤 Module Summary
Next:
Now that you know what misled you, let’s look at practical solutions—measure honestly, inject the focused rule, and reproduce everything.