🧭 Choosing the right model for each task
The most direct use of the yardstick: match the model profile to the task. Sonnet 4.6 only thinks before acting 10% of the time and uses few tools (2,23 per turn) — perfect for what's quick, mechanical, and cheap. Opus 4.8, on the other hand (thinks 54%) and Fable 5 (thinks 85%, test after editing 41%) shine in work that requires plan and close the loop.
| Model | Profile (canonical numbers) | Best use |
|---|---|---|
| Sonnet 4.6 | thinks 10% · 2,23 tools/turn · lean | quick, mechanical, inexpensive tasks |
| Opus 4.8 | thinks 54% · test 2% · 7.96 tools/turn | work that requires planning first |
| Fable 5 | thinks 85% · test 41% · 10.45 tools/turn | work that requires planning + testing (closing the loop) |
💡 Practical example
Rename a field in 30 files? Sonnet—it’s mechanical, and spending Fable on it is wasteful. Refactor a module with a schema migration and test suite? Fable or Opus—you WANT it to plan and run the test. The guide tells you which profile fits; Fable isn’t "economical": it deliberately makes MORE tool calls per turn.
🩺 Diagnose YOUR model
The yardstick isn’t just for comparing famous models—it measures you. Run it on your own history and see where your session stumbles. If your Opus tests after editing only 2% of the time, the biggest lever isn’t “think more”: it’s close the loop. The number points to the real bottleneck, not the one you imagine.
Measure, don't guess
Run the meter on your ~/.claude/projects. Look at think-before, test-after, and tools/turn — yours, not the dataset’s.
Target the right lever
Testing 2% is the ceiling that holds you back the most. Injecting “run the test after editing” pays off much more than asking to “think harder.”
💡 Practical example
You think your agent “doesn’t think enough.” You measure it and find that it thinks 60% of the time—but only tests 3% of the time. The picture changes the prescription: the problem wasn’t reflection, it was not close the loop. Diagnosis before treatment.
👥 Team onboarding
A good rhythm shouldn't depend on every dev remembering. You inject the think-first + test-after by default —
via hook SessionStart or a CLAUDE.md versioned in the repo — and
every team session starts with the playbook enabled. The newcomer picks up the rhythm on day one, without training.
SessionStart hook or CLAUDE.md in the repo
every developer, from day 1
think-first + test-after by default
✓ With the playbook injected
- ✓Standardized pace across all developers.
- ✓Versioned in the repo: change it once, and it applies to everyone.
- ✓It doesn't depend on anyone remembering the good habit.
✗ With nothing injected
- ✗Each developer works at a different pace; quality is inconsistent.
- ✗The good habit dies when someone forgets.
- ✗Onboarding becomes a document nobody reads.
📈 Prove that a change worked
“I think it improved” isn’t evidence. The yardstick turns a hunch into a number: save one baseline dated BEFORE, apply the playbook, and measure AFTER with enough samples. The delta becomes evidence. Just watch out for the dilution: if you throw the new sessions into the same bucket as the old ones, the signal disappears — isolate new sessions.
| Step | What to do | Caution |
|---|---|---|
| 1. Before | dated baseline in a durable location (not /tmp) | large enough sample to be stable |
| 2. Change | apply the playbook (hook/skill/CLAUDE.md) | change one thing at a time |
| 3. After | measure again and compare the delta | isolate new sessions — don’t dilute them in the history |
🔎 Practical example
June baseline: test-after at 5%. You add the rule to close the loop and run it for two weeks. Measure ONLY the new sessions: 38%. Now you have a number you can defend — not “it seems better,” but +33pp measured, in the right slice.
🌐 Applies beyond Fable/Opus
The method isn’t exclusive to the Fable/Opus pair. It works with Codex and open-source models — just have the logs
in the event format (steps with tools, thinking, editing). The key is the field
model: it's how you tell them apart and measure each one by the same standard.
# o que importa é o formato de eventos + o campo model { "model": "claude-opus-4-8", "type": "tool_use", "name": "Edit" } { "model": "gpt-5-codex", "type": "thinking" } { "model": "qwen-2.5-coder", "type": "tool_use", "name": "Bash" } # mesma régua, agrupando por model: python compare_models.py --a codex --b opus
Codex, open source
logs as events
field model
the same as always
💡 Practical example
Want to compare a local model (Qwen Coder) with Codex in YOUR workflow? Run both on the same tasks, collect the events, group by model and measure think-first / test-after. The yardstick is vendor-agnostic — it only needs the format.
🛡️ Antidote to hype
The most valuable use of all: the yardstick is antidote to hype. Before believing any "model X is better than Y," require large sample. That’s exactly how we took down the +45pp: the number came from 7 sessions (99% vs. 54%) and, in a large HF sample (4,892 steps), became 85% vs 54% (+31pp) — solid, defensible, without illusion.
✗ What the hype does
- ✗Concludes from 7 sessions: “+45pp, model X crushes it.”
- ✗Compares samples of different sizes.
- ✗Confuses the presence of “thinking” with content quality.
✓ What the standard requires
- ✓Large sample from both sides before believing (4,892 steps).
- ✓Balanced sizes—cap by the number of steps (~950).
- ✓Measure presence, not content (the reasoning is encrypted).
⚠️ A Small Sample Misleads in BOTH Directions
It inflated “think” (99 → 85) AND hid “test after editing” (the local sample showed 0%; the real figure is 41%). That's why the benchmark isn't just for lowering inflated numbers — it's for finding out what the small sample had obscured.
🔎 Practical example
A post appears: “Model Z thinks 99% of the time!” You ask: how many sessions? Are both sides based on the same sample? If the answer is “7 sessions,” you already know—ask for the large sample before changing anything. The benchmark is your hype filter.
🧰 Module Summary
model.End of the learning path:
You’ve completed the full Fable Lite cycle: the logs are gold (T1), hands-on work extracts the delta (T2), the delta becomes an injected playbook (T3), and real proof measures without fooling you (T4). The yardstick isn’t a comparison trick—it’s your habit of demanding numbers before believing. Choose a model, diagnose yours, standardize the team, prove the change, and cut through the hype. The method is yours; run it on your history and iterate.