MODULE 10 · THREE LESSONS, SIX STEPS
Measure true quality
Compare methods without confusing tuned examples with independent evidence.
Adjust reading and appearance
Build a human reference
What is it?
An evaluation starts with defining what is correct. If labels change from person to person, the metric may be measuring process disagreement. Create a labeling guide, review examples, and document the cases that don’t fit well.
Separate development, calibration, and testing. Use development to improve questions; calibration to choose thresholds; testing to verify the frozen policy. Messages from the same conversation can’t be in different partitions, because that makes it easier to recognize content already seen. When there’s time variation, also evaluate later data.
The lab’s 24 fictional tickets are for learning how to run and read an evaluation. They are not a representative sample of a company and do not prove suitability for production. An operational dataset needs to cover the classes, ambiguities, and changes in real traffic, with human review and data protection.
Why learn
Compare methods without confusing tuned examples with independent evidence.
Key concepts
Use the example below to distinguish the available data, the judgment requested, and what still needs evidence.
Apply: Build human reference
Your turn
Can you adjust the question after looking at the final test error and publish the new rate in the same test as independent?
Check commented answer
No. The test was used for the adjustment. You need to reserve another set for a new independent evaluation.
Compare alternatives
What is it?
Use the same task and the same data to compare rules, Jev, a generative model with structured output, and a hybrid flow. Record configurations and versions. Differences in input or criteria make the comparison hard to interpret.
Accuracy tells you the overall proportion of correct answers, but it can hide small classes. The confusion matrix shows where each class was routed. Macro-F1 gives the same weight to classes, but it also needs to be read with the counts. Pay special attention to the error that has the highest operational cost.
Measure end-to-end latency and the cost of all attempts. The lexical baseline result for this project is reproducible by the CLI: it reports known correct and incorrect outcomes on the fictional dataset. Don’t compare this local rule time with network latency as if they were the same thing. The real Jev experiment can only be reported after execution with access to the service.
Practice with the current resources
The complementary evaluator packages.quality requires a reference for every question. Choice measures accuracy; Noul measures accuracy with cut 0.5 and Brier; Score measures absolute error and normalized error by the scale. The cut used in the metric does not define the action policy. Contract errors reduce coverage and accuracy; MAE and Brier use only valid responses, so they must be read together with coverage.
Why learn
Compare methods without confusing tuned examples with independent evidence.
Key concepts
Use the example below to distinguish the available data, the judgment requested, and what still needs evidence.
Apply: Compare alternatives
Your turn
A very common class has 99% accuracy and another important one has 40%. Is the global average enough?
Check commented answer
No. Report metrics by class, sample support, the confusion matrix, and the cost of errors. The average can hide the important class.
Choose and freeze thresholds
What is it?
As you increase a threshold, usually fewer responses go to the automatic suggestion. This may reduce some errors, but it increases review and can let through the more confident errors. We need to measure two things together: quality on the accepted cases and coverage, the fraction of cases the policy accepts.
Choose the threshold on the calibration set and freeze it before testing. Record whether it uses the class probability, confidence, or both. Don’t copy a threshold from one primitive to another, one language to another, or from one version to another without evaluation.
Measurement uncertainty matters. If ten examples passed without error, we still don’t have proof that the error rate is small. The tighter the target, the more data you usually need. When the sample is insufficient, the correct conclusion may be to collect more examples instead of declaring success.
Deepen the 1.2.0 version
The CLI experiment records dataset, template, policy, model, failures, and provenance. Replay mode lets you compare external responses, including LLMs, but it does not authenticate the stated source. A model-only comparison should have equivalent input; a flow comparison measures the entire task.
Group predictions by confidence ranges to investigate where the errors show up, while preserving the size of each group and class. Too few examples per range leads to fragile conclusions. A change in criteria that fixes a known case must be re-run on an independent set; it’s not enough to demonstrate the corrected case.
Why learn
Compare methods without confusing tuned examples with independent evidence.
Key concepts
Use the example below to distinguish the available data, the judgment requested, and what still needs evidence.
Apply: Choose and freeze thresholds
Your turn
Why does this interaction with invented values not prove that 0.95 is the right threshold?
Check commented answer
Because it doesn’t measure real responses against independent labels. It only shows how a policy reacts to numbers.
Module wrap-up
- Retrieve the chosen decision from the start of the course.
- Compare your answer with the examples from this module.
- Record a change to the criteria and the test needed to accept it.
Quick check
You corrected the criteria until you got every known example right. How do you measure the improvement?
Practice and continuity
Open the labs and answer keys · Project visual lab
# In the jev repo: offline demo, no API
python3 -m pacotes.executar reunioes
python3 -m pacotes.qualidade reunioesThese outputs use a fictional fixture. To test your data, use the human reference script and explicitly enable real mode.