PTENES
MODULE 4.4

🧪 Pilot, proof of concept, and success metrics

Define success before you start, distinguish a PoC from a pilot, set the scope that teaches you the most, keep a human in the loop — and have the courage to kill what didn’t work.

6
Topics
~45
Minutes
Plan
Level
Validation
Type

The pilot is the the project’s most valuable experiment — and the riskiest to get wrong. Getting it wrong in a pilot is cheap. Getting it wrong in production is expensive. The difference lies in how you structure the pilot: with success criteria defined upfront, minimal scope, and a human in the loop.

🎯 Define criteria first + baseline 🔬 Pilot minimum scope human in the loop go/ no-go ✅ GO scale · adjust the roadmap 🛑 NO-GO kill · pivot · document killing a pilot that didn’t work is process success

Pilot flow—the success criterion comes first, not after seeing the results.

1

🎯 Define success before you start

Defining the success criterion after seeing the results guarantees confirmation bias. The the question the pilot answers needs to be written down before any execution.

✓ Criteria defined in advance

  • ✓"The model classifies with >85% accuracy across 200 real cases"
  • ✓"Average triage time drops from 4h to <60min"
  • ✓Evaluated over 4 weeks using production data

✗ Criteria defined afterward

  • ✗"We think it turned out well"
  • ✗"The team loved using it"
  • ✗Without a number—no defensible decision

📝 Success criteria template

Main question: Can we classify support intent with sufficient accuracy?
Metric: Classification accuracy vs. human review
Minimum threshold: ≥ 82% accuracy
Volume: 500 real production tickets
Timeline: 4 weeks of operation
Comparison: Vs. baseline: 100% manual review, 4h/day
2

🔬 PoC vs. pilot vs. production

Three stages, three goals, three different datasets. Mixing up the stages is where wrong expectations begin—and those expectations later derail the project.

PoC

Proof of Concept — “is this possible?”

Lab data, with no real integration and no real users. It answers whether the technology can solve the problem under ideal conditions.

Duration: 1-2 weeks. Result: technically feasible / infeasible.

Pilot

Pilot—"Does this work here?"

Real data, real users, but controlled scope and volume. Answers whether the solution works in this specific operational context.

Duration: 4-8 weeks. Result: evidence for a go/no-go decision.

Prod.

Production — "scale with confidence"

Full volume, defined SLA, production monitoring. The pilot validated that it works — now scale it.

Prerequisite: pilot with go approval. SLA and runbook documented.

💡 Practical tip

When presenting the PoC to the client, make it clear: “This demonstrates what the technology can do — but we haven’t tested it with your real data and users yet. The pilot does that.” This manages expectations before they turn into disappointments.

3

🔭 Minimum scope that teaches you something

The ideal pilot is the smallest experiment that answers the most important question. Not the most comprehensive, not the most impressive — the smartest.

✓ Minimum scope that teaches

  • ✓One ticket category (not all 12)
  • ✓One pilot department (not the whole company)
  • ✓200 real cases (not 50.000)
  • ✓Manual integration OK (automation comes later)

✗ Inflated scope

  • ✗Build the entire integration before validating the model
  • ✗Cover every use case in the pilot
  • ✗A 6-month pilot with no decision date

💡 Scope test question

Before planning the pilot, ask: “What is the most important question?” Then: “What is the smallest experiment that answers it?” Anything that doesn’t answer that question is out of scope for the pilot.

4

👤 Human in the loop + failure plan

During the pilot, every AI output is reviewed by a human. This isn’t inefficiency—it’s the mechanism that catches errors before the end customer does and feeds the data back for improvement.

🔄 Human-in-the-loop structure

🤖

AI processes

produces a confident outcome

👁️

Human reviews

approves, corrects, or rejects

📊

Recorded data

errors become data for improvement

🚨 Failure plan—required

The failure plan defines what happens when the AI makes a mistake or becomes unavailable:

  • →Who reverses: a person identified by name, not by job title
  • →How long: Reversal SLA (e.g., within 2h)
  • →How the operation continues: documented temporary manual process
  • →When to trigger: clear criterion (e.g., error rate >10% for 1h)
5

📊 Measure vs. Baseline

The before-and-after comparison is the only argument that turns perception into evidence. The the metric must be defined using the same definition as the baseline — otherwise, we're comparing different things.

📐 Rules for a valid comparison

Same metric

If the baseline measured "average ticket triage time," the pilot measures the same thing—not "total triage time for the day."

Equivalent period

Measure across comparable weeks (same volume, same seasonality). Avoid comparing a Monday baseline with a Friday pilot.

Attributable variation

Isolate what changed because of AI from what changed for other reasons (different team, different volume, different training).

📊 What to document in the pilot report

• Preregistered metric: [value defined in advance]
• Pilot result: [measured value]
• Comparison with baseline: [% improvement or decline]
• Confounding factors identified: [list]
• Confidence in the result: [high/medium/low and why]
6

🚦 Go/no-go criteria: scale or kill

Go/no-go is the most important decision in the process. Were the criteria met? Go: scale up with what you learned. Weren’t they? No-go: kill it or pivot— no blame, just data and clear next steps.

🟢 GO — criteria met

Next steps when deciding GO:

  • • Update the roadmap with what the pilot taught us
  • • Document the reusable capabilities created
  • • Define an SLA and runbook for production
  • • Share the results with the sponsor, including the data
  • • Plan Wave 2 with expanded scope

🔴 NO-GO — criteria not met

Next steps when deciding NO-GO:

  • • Document why it didn’t work (data? technology? process?)
  • • Assess whether a viable pivot exists (smaller scope? different level?)
  • • Formally cancel the project if there’s no viable pivot
  • • Record lessons learned for upcoming pilots
  • • Tell the sponsor: “killing it was the right decision”

💡 Why killing it is success

A pilot that ends in a no-go and is canceled has delivered exactly what it should: the answer “don’t do this” before spending 10× more on production. That protects the budget, preserves trust, and frees the team to pursue the next, more promising use case. A consultant who can close a project that doesn’t work is more valuable than one who lets it continue out of inertia.

🎒 Module summary

✓
Success criterion defined beforehand, not afterward — define what "worked" means before looking at the data.
✓
PoC ≠ Pilot ≠ Production — three stages with distinct objectives and data.
✓
Minimum scope that teaches — the smallest experiment that answers the main question.
✓
Human in the loop + failure plan — catches errors before they reach the end customer and keeps operations running if it goes down.
✓
Data-driven go/no-go, not intuition — killing what doesn’t work is a process success.

Next track:

T5 — Consulting Delivery & Business