Establish a reference configuration
What it is
The walk it down starts with a configuration that can accomplish the task. Record model, available effort, tools, skill version, and the cases used. This is the reference against which lighter alternatives will be compared.
Why learn
Without a baseline you don’t know whether the drop in quality came from the configuration, the data, or the instruction. The names and results cited in the video are examples from that experiment, not a guaranteed price or performance hierarchy.
Key concepts
Experiment E01
Skill: version 1.0.0
Cases: C01 to C08
Model: identifier available in your account
Effort: value supported by the interface
Record: approval, time, consumption, and rework.
✓ Do it like this
Use names and options that are actually available in the environment.
✗ Avoid this mistake
Fix course pricing or menus that may change.
Practice before revealing
What information is missing in “model B was faster” ?
View commented response
Cases, number of runs, acceptance criteria, tools, effort, and time variation. A comparative statement needs those conditions.
Compare one variable at a time
What it is
Keep cases, skill, and rubric the same when changing the model. After you pick a candidate, test another effort level if the interface offers that control. Don’t assume all models accept the same options.
Why learn
Changing the model, prompt, and tool all at once prevents attributing the difference to a single cause. A simple, controlled comparison is worth more than a big spreadsheet full of incompatible conditions.
Key concepts
Round 1: model A vs B, equivalent effort when possible.
Round 2: best candidate, high vs medium effort.
Round 3: validate the choice on reserved cases.
If options aren’t equivalent, record the difference.
From concept to action
- Model A: identify the initial condition.
- Same cases: apply the described decision.
- Model B: check the effect in the example.
- Same rubric: record the output evidence.
✓ Do it like this
Keep inputs frozen during the comparison.
✗ Avoid this mistake
Improve the skill for only one candidate and call the test fair.
Practice before revealing
You changed the prompt to help B. What should you do before comparing?
View commented response
Run A again with the same skill version, or declare that they are two different systems. The important part is not attributing an instruction change to the model.
Choose cases that represent the real work
What it is
In the video’s visual study, one candidate was better on certain crops and worse on others. To find what works for your job, include examples of different natures: screen demonstrations, talk with few images, and content with charts.
Why learn
A single example favors coincidences. Repeat important cases and examine where the configuration fails, not just your average. A high acceptance rate can hide a serious error in a rare but essential type.
Key concepts
Report: normal, empty, missing field, invalid value.
Article: tutorial, interview, presentation with charts.
Repeat the cases with the highest risk and save raw results.
✓ Do it like this
Publish the sample size alongside the acceptance rate.
✗ Avoid this mistake
Generalize task performance across all skills.
Practice before revealing
A configuration passes in 9 out of 10 cases, but invents a value in the tenth. Can it win by average?
View commented response
Not inventing values is blocking. The acceptance policy needs to be applied before comparing costs; the average doesn’t compensate for a critical failure.
Calculate cost per accepted delivery
What it is
The price per call is only part of the cost. Include attempts and corrections. In the didactic example, A spends 12 units for 10 accepted deliveries; B spends 8 for 5. A costs 1.20 per accepted delivery and B costs 1.60.
Why learn
A nominally cheap configuration can consume more resources because of rework. Record human time separately or establish an explicit way to monetize it. Don’t mix API cost with a subscription without explaining the method.
Key concepts
A: 12 units / 10 accepted = 1.20 per acceptance
B: 8 units / 5 accepted = 1.60 per acceptance
If accepted = 0: no finite cost per acceptance exists.
Human review time: record in a separate column.
✓ Do it like this
Add up the cost of the attempts that failed.
✗ Avoid this mistake
Exclude retries to make the alternative look cheaper.
Practice before revealing
C costs 9 units and approves 6 outputs. What’s the cost per acceptance?
View commented response
1.50 units per accepted delivery. Only compare with A and B if the quality criterion and the set of cases are equivalent.
How much does each accepted delivery cost?
Include unsuccessful attempts in the cost. This simulation does not query service prices.
Reduce effort with a quality gate
What it is
After you find a suitable model, try reducing effort when that control exists. The decision depends on observed quality, time, and consumption. More effort is not a guarantee of better results for every task.
Why learn
A skill with clear rules and a calculation script may require less deliberation than a complex visual review. The quality gate prevents the economy from removing exactly the check that avoids errors.
Key concepts
Current configuration → passes the suite.
Lower effort → run the same suite.
Did it pass without relevant regressions? Compare time and cost.
Did it fail? Roll back to the reference or investigate the skill.
From concept to action
- Approved reference: identify the initial condition.
- Reduce effort: apply the described decision.
- Reevaluate: check the effect in the example.
- Adopt or roll back: record the output evidence.
✓ Do it like this
Save the previous configuration for rollback.
✗ Avoid this mistake
Skip the review to reduce time and call it optimization.
Practice before revealing
The lightweight configuration works in summary, but fails on visual cropping. What does that suggest?
View commented response
Possibly split the flow and use different configurations per stage, if the environment allows it. The decision needs a new integrated test; splitting stages also adds cost and complexity.
Record a decision that can be revisited
What it is
The result of the experiment must state which configuration was chosen, why, in which cases, and with what limits. Include a review date or events that justify repeating the comparison: a new tool, model, input type, or a contract change.
Why learn
Optimization is a situated decision. An update can change behavior or availability. Keeping the evidence lets you revisit the choice without starting from vague opinions or from remembering the best example.
Key concepts
Choice: configuration B for text summaries.
Base: passed on 12 cases, 3 rounds per case.
Limit: not evaluated for image editing.
Reevaluate: model change, rubric change, or input change.
✓ Do it like this
Differentiate measurement, estimation, and information not available.
✗ Avoid this mistake
State “the best model” without specifying the task and conditions.
Practice before revealing
What conclusion fits after a pilot with only two cases?
View commented response
“The candidate deserves a bigger evaluation.” Two cases might reveal a defect, but they don’t support a broad promise of reliability.
Check your understanding
Which number helps more to choose an economical configuration?
What you take from this module
Plan a fair comparison and compute cost per accepted delivery.
- Establish a reference configuration.
- Compare one variable at a time.
- Choose cases that represent real work.
- Calculate the cost per accepted delivery.
- Reduce effort with a quality gate.
- Record a decision that can be reviewed.
Next action: save the exercise in your learning lab and record what still needs review.