PTENES
Skip to content
MODULE 4.3

🧰 Examples of usefulness

The question that matters: where is this actually useful? Here, the benchmark leaves the page and becomes a decision—choosing the right model for the task, diagnosing where YOU fail, standardizing the team’s rhythm, proving a change worked, measuring Codex and open source, and using measurement as an antidote to hype. Concrete cases, with canonical numbers.

6
Topics
~22
Minutes
Applied
Level
Cases
of use
the yardstick an honest measurement choose the right model diagnose your team onboarding prove before / after decisions with numbers
1

🧭 Choosing the right model for each task

The most direct use of the yardstick: match the model profile to the task. Sonnet 4.6 only thinks before acting 10% of the time and uses few tools (2,23 per turn) — perfect for what's quick, mechanical, and cheap. Opus 4.8, on the other hand (thinks 54%) and Fable 5 (thinks 85%, test after editing 41%) shine in work that requires plan and close the loop.

Model Profile (canonical numbers) Best use
Sonnet 4.6 thinks 10% · 2,23 tools/turn · lean quick, mechanical, inexpensive tasks
Opus 4.8 thinks 54% · test 2% · 7.96 tools/turn work that requires planning first
Fable 5 thinks 85% · test 41% · 10.45 tools/turn work that requires planning + testing (closing the loop)

💡 Practical example

Rename a field in 30 files? Sonnet—it’s mechanical, and spending Fable on it is wasteful. Refactor a module with a schema migration and test suite? Fable or Opus—you WANT it to plan and run the test. The guide tells you which profile fits; Fable isn’t "economical": it deliberately makes MORE tool calls per turn.

2

🩺 Diagnose YOUR model

The yardstick isn’t just for comparing famous models—it measures you. Run it on your own history and see where your session stumbles. If your Opus tests after editing only 2% of the time, the biggest lever isn’t “think more”: it’s close the loop. The number points to the real bottleneck, not the one you imagine.

›

Measure, don't guess

Run the meter on your ~/.claude/projects. Look at think-before, test-after, and tools/turn — yours, not the dataset’s.

›

Target the right lever

Testing 2% is the ceiling that holds you back the most. Injecting “run the test after editing” pays off much more than asking to “think harder.”

💡 Practical example

You think your agent “doesn’t think enough.” You measure it and find that it thinks 60% of the time—but only tests 3% of the time. The picture changes the prescription: the problem wasn’t reflection, it was not close the loop. Diagnosis before treatment.

3

👥 Team onboarding

A good rhythm shouldn't depend on every dev remembering. You inject the think-first + test-after by default — via hook SessionStart or a CLAUDE.md versioned in the repo — and every team session starts with the playbook enabled. The newcomer picks up the rhythm on day one, without training.

Channel

SessionStart hook or CLAUDE.md in the repo

Who wins

every developer, from day 1

What goes in

think-first + test-after by default

✓ With the playbook injected

  • ✓Standardized pace across all developers.
  • ✓Versioned in the repo: change it once, and it applies to everyone.
  • ✓It doesn't depend on anyone remembering the good habit.

✗ With nothing injected

  • ✗Each developer works at a different pace; quality is inconsistent.
  • ✗The good habit dies when someone forgets.
  • ✗Onboarding becomes a document nobody reads.
4

📈 Prove that a change worked

“I think it improved” isn’t evidence. The yardstick turns a hunch into a number: save one baseline dated BEFORE, apply the playbook, and measure AFTER with enough samples. The delta becomes evidence. Just watch out for the dilution: if you throw the new sessions into the same bucket as the old ones, the signal disappears — isolate new sessions.

Step What to do Caution
1. Before dated baseline in a durable location (not /tmp) large enough sample to be stable
2. Change apply the playbook (hook/skill/CLAUDE.md) change one thing at a time
3. After measure again and compare the delta isolate new sessions — don’t dilute them in the history

🔎 Practical example

June baseline: test-after at 5%. You add the rule to close the loop and run it for two weeks. Measure ONLY the new sessions: 38%. Now you have a number you can defend — not “it seems better,” but +33pp measured, in the right slice.

5

🌐 Applies beyond Fable/Opus

The method isn’t exclusive to the Fable/Opus pair. It works with Codex and open-source models — just have the logs in the event format (steps with tools, thinking, editing). The key is the field model: it's how you tell them apart and measure each one by the same standard.

# o que importa é o formato de eventos + o campo model
{ "model": "claude-opus-4-8", "type": "tool_use", "name": "Edit" }
{ "model": "gpt-5-codex",      "type": "thinking" }
{ "model": "qwen-2.5-coder",   "type": "tool_use", "name": "Bash" }

# mesma régua, agrupando por model:
python compare_models.py --a codex --b opus
Apply the

Codex, open source

Requirement

logs as events

The key

field model

The yardstick

the same as always

💡 Practical example

Want to compare a local model (Qwen Coder) with Codex in YOUR workflow? Run both on the same tasks, collect the events, group by model and measure think-first / test-after. The yardstick is vendor-agnostic — it only needs the format.

6

🛡️ Antidote to hype

The most valuable use of all: the yardstick is antidote to hype. Before believing any "model X is better than Y," require large sample. That’s exactly how we took down the +45pp: the number came from 7 sessions (99% vs. 54%) and, in a large HF sample (4,892 steps), became 85% vs 54% (+31pp) — solid, defensible, without illusion.

✗ What the hype does

  • ✗Concludes from 7 sessions: “+45pp, model X crushes it.”
  • ✗Compares samples of different sizes.
  • ✗Confuses the presence of “thinking” with content quality.

✓ What the standard requires

  • ✓Large sample from both sides before believing (4,892 steps).
  • ✓Balanced sizes—cap by the number of steps (~950).
  • ✓Measure presence, not content (the reasoning is encrypted).

⚠️ A Small Sample Misleads in BOTH Directions

It inflated “think” (99 → 85) AND hid “test after editing” (the local sample showed 0%; the real figure is 41%). That's why the benchmark isn't just for lowering inflated numbers — it's for finding out what the small sample had obscured.

🔎 Practical example

A post appears: “Model Z thinks 99% of the time!” You ask: how many sessions? Are both sides based on the same sample? If the answer is “7 sessions,” you already know—ask for the large sample before changing anything. The benchmark is your hype filter.

🧰 Module Summary

✓
Choosing the right model — Sonnet (10%, 2.23) for speed; Opus/Fable for planning + testing.
✓
Diagnose YOUR model — measure where YOU fail; testing at 2% calls for closing the loop, not “thinking more.”
✓
Team onboarding — inject the pace by default; every developer starts off well.
✓
Prove the change — set a dated baseline first, measure afterward, isolate new sessions.
✓
It works beyond Fable/Opus — Codex and open-source; the key is the field model.
✓
Antidote to hype — require a large sample; that’s how +45pp became a solid +31pp.

End of the learning path:

You’ve completed the full Fable Lite cycle: the logs are gold (T1), hands-on work extracts the delta (T2), the delta becomes an injected playbook (T3), and real proof measures without fooling you (T4). The yardstick isn’t a comparison trick—it’s your habit of demanding numbers before believing. Choose a model, diagnose yours, standardize the team, prove the change, and cut through the hype. The method is yours; run it on your history and iterate.