PTENES
Skip to content
MODULE 4.1

🪤 The Sample Trap

We measure the same delta three times — and the number changed every time. Here you can see how a small sample misleads in the two directions: it inflated an advantage that didn’t exist (+45pp) and hid another that was real. The honest benchmark only appears with a large sample on both sides.

6
Topics
~25
Minutes
Advanced
Level
Conceptual
Type
7 7 sessions ~950 steps ~950 equal sample 4 sessions (Opus) 4.9k 4.892 steps (HF) 30 sessions +45pp misleading 94% Opus unstable +31pp solid / defensible
1

🧪 The three tests, in order

The Fable vs. Opus delta didn’t come from a single measurement—it came from a three-part journey. First, a handful of local sessions. Then an attempt to balance the sample size. Finally, an open dataset with thousands of steps. Each round produced a different number — and that difference is precisely what teaches you not to trust the first result.

1

7 local sessions

The first sample: ~950 steps from local Fable. Exciting, but tiny—and that’s where the trap began.

2

Equal sample (~950 each)

We tried to balance them by capping the number of steps. It helps—but if the other side is still tiny, it’s still noise.

3

4.892 steps from HF

Open dataset Fable-5-traces: 30 sessions. Large samples on both sides = a defensible number.

💡 Why measure three times?

Because the first number was too polished. When a result seems too good to be true, the first question isn’t "wow, how much?" — it’s "how many sessions did that come from?". Repeating it with more data is what separates a finding from an illusion.

2

1️⃣ Test 1 — the tiny sample

The first sample was 7 sessions / 950 steps from local Fable. The number came out round and exciting: 99% think in Fable versus 54% in Opus — a delta of +45pp. It looks like definitive proof. But 99% of a sample of 7 sessions is exactly the kind of number that should raise the alert, not balancing the books.

✓ What the sample seemed to show

  • ✓Fable thinks before acting 99% of the time.
  • ✓Opus thinks 54% — a chasm of +45pp.
  • ✓“Done, it’s proven: Fable is much better.”

✗ What the sample of 7 concealed

  • ✗7 sessions do not represent typical behavior.
  • ✗99% almost never survives more data (dropped to 85%).
  • ✗The +45pp was a sample-size artifact, not a fact.

🔎 The 99% pitfall

Such an extreme value (99%) with a small n is a classic sign of sample overfitting: few sessions, all similar, and the statistic "saturates." The right number isn't zero or 99% — it's the 85% that only shows up when you add thousands of steps to the count.

3

⚖️ Test 2 — equal sample

Skeptical of the +45pp, we tried the obvious: balance the sizes. We capped both sides at ~950 steps each. Surprising result — Opus jumped to 94%. But this Opus sample came from just 4 sessions. Capping the size helped remove volume bias, but traded one problem for another: 4 sessions is still noise, and the number became unstable.

📊 What we learned from test 2

  • Balancing size is necessary — comparing 950 vs 42k steps will always distort the results.
  • But that's not enough — if the balanced side has 4 sessions, you’ve only swapped one bias for variance.
  • 94% from 4 sessions ≠ truth — it's the same error as test 1, now from the other side.

💡 Practical tip

Balancing the sample solves it a of the two threats (volume bias). The other one — small-sample variance — only disappears with more sessions. Balance the size e ensure there’s enough volume on both sides before drawing a conclusion.

4

📊 Test 3 — 4,892 HF steps

The answer came from the open dataset Glint-Research/Fable-5-traces: 4.892 steps across 30 sessions from Fable, compared against Opus's broader baseline. Now “think before acting” has stabilized at Fable 85% vs Opus 54% (+31pp). It’s not the euphoric +45pp or the inverted 94% — it’s the number that survives at scale. A large sample on both sides is what makes the comparison defensible.

Test Sample Sessions % thinks (Fable vs Opus) Reading
1️⃣ local ~950 steps 7 sessions 99% vs. 54% (+45pp) misleading
⚖️ equal ~950 each 4 sessions (Opus) Opus jumps to 94% unstable
📊 HF 4.892 steps 30 sessions 85% vs 54% (+31pp) solid
Dataset

Fable-5-traces (HF)

Steps

4.892 / 262 turns

Sessions

30

Solid delta

85 vs 54 = +31pp

5

🔀 Deception in TWO directions

This is the module’s central point. A small sample isn’t just wrong "on the high side" — it’s wrong in the two directions at the same time. In our case, it inflated the “think before acting” (showed 99%, the real figure is 85%) and at the same time hid a real advantage: the "test after editing," which the local version marked in 0%, in the large sample is 41%. In other words: the small sample overestimated one dimension and underestimated the other.

what the small sample showed INFLATED the “thinking” 99% → 85% real is smaller HID the “test” 0% → 41% real is larger

↑ Direction 1 — inflated

  • ↑“Think before acting”: the local setup said 99%.
  • ↑Large sample corrects to 85%.
  • ↑The excess turned into a false +45pp over Opus.

↓ Direction 2 — hidden

  • ↓“Test after editing”: the local setup marked 0%.
  • ↓In the large sample, Fable tests 41% (Opus 2%).
  • ↓A real advantage was hidden in the small sample.

💡 Remember the honesty of the measurement

The thinking comes encrypted in the logs — so we measure the presence of the reasoning, not the content. Even so, its presence is enough to show that the small sample was skewed in both directions: it inflated areas where there was little and erased areas where there was a lot.

6

📐 The golden rule

It all boils down to three commandments. Never conclude of a small sample. Balance the sizes before comparing (cap by number of steps). And confirm with a large sample before believing it. Following this standard is how the euphoric +45pp became a defensible +31pp — and how the hidden advantage of "testing" (0→41%) finally emerged.

✓ Do it this way

  • ✓Balance the sample (e.g., ~950 steps per side).
  • ✓Require volume on both sides, not just one.
  • ✓Confirm with a large dataset before believing it.

✗ Never do this

  • ✗Closing the account with just 7 sessions because the number looks good.
  • ✗Compare 950 steps against 42k without balancing.
  • ✗Believe the 99% (or the 94%) without repeating the test with more data.

🔎 The yardstick in one sentence

Before believing that “model X is better than Y,” ask: how many sessions this came from, and were both sides balanced? If the answer is “a few” or “no,” you don’t have a number yet—you have an impression.

🪤 Module Summary

✓
Three tests, in order — 7 sessions → equal sample → 4,892 HF steps.
✓
Test 1 — the tiny sample — 99% vs. 54% (+45pp) from just 7 sessions. Misleading.
✓
Test 2 — same sample — Opus jumps to 94% in a 4-session subset. Unstable.
✓
Test 3 — 4,892 HF steps — Fable 85% vs. Opus 54% (+31pp). Solid.
✓
Misleading in TWO directions — inflated “thinking” (99→85) AND hid “testing” (0→41%).
✓
The golden rule — never draw conclusions from a small sample; balance it; confirm with a large sample.

Next:

Now that you know what misled you, let’s look at practical solutions—measure honestly, inject the focused rule, and reproduce everything.