Theme

Font

Size

Text width

Line spacing

Controls accent

0 of 0 0%
MODULE 4.1

Compare quality, cost, and effort

Do the walk it down with equivalent cases and decisions supported by results.

At the end: Plan a fair comparison and calculate cost per accepted delivery.

6 topics
55 min estimated with hands-on practice
6 commented exercises
1 final check
1

Establish a reference configuration

Stable skill Fixed cases Configuration A Measured reference
First demonstrate that the task is executable; then look for a more economical configuration.

What it is

The walk it down starts with a configuration that can accomplish the task. Record model, available effort, tools, skill version, and the cases used. This is the reference against which lighter alternatives will be compared.

Why learn

Without a baseline you don’t know whether the drop in quality came from the configuration, the data, or the instruction. The names and results cited in the video are examples from that experiment, not a guaranteed price or performance hierarchy.

Key concepts

Baselineconfiguration that passed.
Versioninstruction frozen.
Environmentsame tools.
Recordobserved cost, time, and quality.
COMMENTED EXAMPLE · 4.1.1
Experiment E01
Skill: version 1.0.0
Cases: C01 to C08
Model: identifier available in your account
Effort: value supported by the interface
Record: approval, time, consumption, and rework.

✓ Do it like this

Use names and options that are actually available in the environment.

✗ Avoid this mistake

Fix course pricing or menus that may change.

Practice before revealing

What information is missing in “model B was faster” ?

View commented response

Cases, number of runs, acceptance criteria, tools, effort, and time variation. A comparative statement needs those conditions.

2

Compare one variable at a time

Model A Same cases Model B Same rubric
The comparison is fair when the candidates face the same relevant conditions.

What it is

Keep cases, skill, and rubric the same when changing the model. After you pick a candidate, test another effort level if the interface offers that control. Don’t assume all models accept the same options.

Why learn

Changing the model, prompt, and tool all at once prevents attributing the difference to a single cause. A simple, controlled comparison is worth more than a big spreadsheet full of incompatible conditions.

Key concepts

Independent variablewhat changes.
Controlswhat stays the same.
Repetitionsvariation between runs.
Orderalternate to reduce momentum effects.
COMMENTED EXAMPLE · 4.1.2
Round 1: model A vs B, equivalent effort when possible.
Round 2: best candidate, high vs medium effort.
Round 3: validate the choice on reserved cases.
If options aren’t equivalent, record the difference.

From concept to action

  1. Model A: identify the initial condition.
  2. Same cases: apply the described decision.
  3. Model B: check the effect in the example.
  4. Same rubric: record the output evidence.

✓ Do it like this

Keep inputs frozen during the comparison.

✗ Avoid this mistake

Improve the skill for only one candidate and call the test fair.

Practice before revealing

You changed the prompt to help B. What should you do before comparing?

View commented response

Run A again with the same skill version, or declare that they are two different systems. The important part is not attributing an instruction change to the model.

3

Choose cases that represent the real work

Screen Speech Charts Reserved cases
The sample must reflect what you receive, including difficult cases that don’t fit the average.

What it is

In the video’s visual study, one candidate was better on certain crops and worse on others. To find what works for your job, include examples of different natures: screen demonstrations, talk with few images, and content with charts.

Why learn

A single example favors coincidences. Repeat important cases and examine where the configuration fails, not just your average. A high acceptance rate can hide a serious error in a rare but essential type.

Key concepts

Diversityinput types.
Output instability repetition.
Stratify results by category.
Reserved case test outside the tuning.
COMMENTED EXAMPLE · 4.1.3
Report: normal, empty, missing field, invalid value.
Article: tutorial, interview, presentation with charts.
Repeat the cases with the highest risk and save raw results.

✓ Do it like this

Publish the sample size alongside the acceptance rate.

✗ Avoid this mistake

Generalize task performance across all skills.

Practice before revealing

A configuration passes in 9 out of 10 cases, but invents a value in the tenth. Can it win by average?

View commented response

Not inventing values is blocking. The acceptance policy needs to be applied before comparing costs; the average doesn’t compensate for a critical failure.

4

Calculate cost per accepted delivery

Attempts Accumulated cost Accepted deliveries Effective cost
The numbers are synthetic and meant for learning the calculation; they are not model prices.

What it is

The price per call is only part of the cost. Include attempts and corrections. In the didactic example, A spends 12 units for 10 accepted deliveries; B spends 8 for 5. A costs 1.20 per accepted delivery and B costs 1.60.

Why learn

A nominally cheap configuration can consume more resources because of rework. Record human time separately or establish an explicit way to monetize it. Don’t mix API cost with a subscription without explaining the method.

Key concepts

Total cost of all relevant attempts.
Accepted deliveries passed the contract.
Effective cost total divided by acceptances.
Zero acceptances means effective cost is not zero.
COMMENTED EXAMPLE · 4.1.4
A: 12 units / 10 accepted = 1.20 per acceptance
B: 8 units / 5 accepted = 1.60 per acceptance
If accepted = 0: no finite cost per acceptance exists.
Human review time: record in a separate column.

✓ Do it like this

Add up the cost of the attempts that failed.

✗ Avoid this mistake

Exclude retries to make the alternative look cheaper.

Practice before revealing

C costs 9 units and approves 6 outputs. What’s the cost per acceptance?

View commented response

1.50 units per accepted delivery. Only compare with A and B if the quality criterion and the set of cases are equivalent.

SIMULATOR · DIDACTIC NUMBERS

How much does each accepted delivery cost?

1.20 units per acceptance

Include unsuccessful attempts in the cost. This simulation does not query service prices.

5

Reduce effort with a quality gate

Approved reference Reduce effort Reevaluate Adopt or roll back
Only step down one more rung after checking the effect of the previous rung.

What it is

After you find a suitable model, try reducing effort when that control exists. The decision depends on observed quality, time, and consumption. More effort is not a guarantee of better results for every task.

Why learn

A skill with clear rules and a calculation script may require less deliberation than a complex visual review. The quality gate prevents the economy from removing exactly the check that avoids errors.

Key concepts

Candidate configuration to try.
Quality gate immutable minimum criteria.
Acceptance demonstrated in the cases.
Rollback return to the reference if it fails.
COMMENTED EXAMPLE · 4.1.5
Current configuration → passes the suite.
Lower effort → run the same suite.
Did it pass without relevant regressions? Compare time and cost.
Did it fail? Roll back to the reference or investigate the skill.

From concept to action

  1. Approved reference: identify the initial condition.
  2. Reduce effort: apply the described decision.
  3. Reevaluate: check the effect in the example.
  4. Adopt or roll back: record the output evidence.

✓ Do it like this

Save the previous configuration for rollback.

✗ Avoid this mistake

Skip the review to reduce time and call it optimization.

Practice before revealing

The lightweight configuration works in summary, but fails on visual cropping. What does that suggest?

View commented response

Possibly split the flow and use different configurations per stage, if the environment allows it. The decision needs a new integrated test; splitting stages also adds cost and complexity.

6

Record a decision that can be revisited

Results Choice Limitations Reevaluate when change
A good decision stays understandable when someone else needs to review it.

What it is

The result of the experiment must state which configuration was chosen, why, in which cases, and with what limits. Include a review date or events that justify repeating the comparison: a new tool, model, input type, or a contract change.

Why learn

Optimization is a situated decision. An update can change behavior or availability. Keeping the evidence lets you revisit the choice without starting from vague opinions or from remembering the best example.

Key concepts

Decision adopted configuration.
Scope covered tasks.
Evidence results and artifacts.
Review trigger relevant change.
COMMENTED EXAMPLE · 4.1.6
Choice: configuration B for text summaries.
Base: passed on 12 cases, 3 rounds per case.
Limit: not evaluated for image editing.
Reevaluate: model change, rubric change, or input change.

✓ Do it like this

Differentiate measurement, estimation, and information not available.

✗ Avoid this mistake

State “the best model” without specifying the task and conditions.

Practice before revealing

What conclusion fits after a pilot with only two cases?

View commented response

“The candidate deserves a bigger evaluation.” Two cases might reveal a defect, but they don’t support a broad promise of reliability.

CHECK WITHOUT BLOCKING

Check your understanding

Which number helps more to choose an economical configuration?

What you take from this module

Plan a fair comparison and compute cost per accepted delivery.

  • Establish a reference configuration.
  • Compare one variable at a time.
  • Choose cases that represent real work.
  • Calculate the cost per accepted delivery.
  • Reduce effort with a quality gate.
  • Record a decision that can be reviewed.

Next action: save the exercise in your learning lab and record what still needs review.

Module reference: the provided transcript and course sources and technical notes.