PTENES
TRACK 4

🗺️ Execution plans

From opportunity to an actionable plan: how to build the business case, sequence the roadmap in waves, choose the right pyramid level, and validate with a pilot before scaling. Always start at the base.

🔍 Opportunity ROI + total cost 🧾 Business Case 1 page, defensible 🪜 The right level foundation → AI → agent only if it pays off 🧪 Pilot go/no-go → scale crawl walk run roadmap in waves
4
Modules
24
Topics
~3h
Duration
Plan
Level

Learning path map

Detailed content

4.1~45 min

💰 From Opportunity to Business Case: ROI and Total Cost

How to determine whether an AI initiative really pays off: ROI components, total cost of ownership, baseline, and the one-page business case decision-makers understand.

What it is:

ROI = gain − (build cost + operating cost), adjusted for risk. Gain is what the problem costs today or what the solution generates; build is what you’ll spend to create it; operations is the ongoing cost after launch; risk is the discount factor based on the chance it won’t work.

Why learn:

Without this structure, estimates like "save X hours" are unsupported, and the decision-maker can’t approve or compare options. The formula gives technical and business teams a common language.

Key concepts:

Risk-adjusted ROI; opportunity cost; discount factor based on probability of success.

What it is:

The total cost of an AI solution includes: LLM APIs (billed per token), cloud infrastructure, maintenance costs when the model changes or the data drifts, and human oversight to review outputs. Projects that ignore these recurring costs often end up in the red after launch.

Why learn:

The calculation doesn’t end at deployment. Knowing how to estimate operating costs changes the decision between building, buying a ready-made SaaS, or simply not using AI.

Key concepts:

TCO (total cost of ownership); cost per API call; model drift; cost of human oversight in the loop.

What it is:

A baseline is a measurement of the current process before any changes: how many hours it takes, the error rate, the cost per unit, and the flow's NPS. Without this number, any improvement is perception, not evidence.

Why learn:

The baseline makes it possible to compare before and after objectively. Projects that skip this step can’t demonstrate results—and lose the next contract even if they delivered real value.

Key concepts:

Measure before changing; process metric vs. outcome metric; document the baseline as a deliverable.

What it is:

Payback is the time it takes to recover the initial investment. A solution that costs R$50k to build and saves R$5k/month has a 10-month payback period. The horizon defines how long the calculation remains valid — AI changes quickly, and solutions have a useful life.

Why learn:

AI projects with a payback period over 24 months rarely make it there intact—the context changes first. Knowing how to calculate payback helps you size the scope correctly and discard cases that don’t add up.

Key concepts:

Simple and discounted payback; validity horizon; accelerated depreciation in AI.

What it is:

Model metrics (accuracy, F1, latency) measure technical quality; business metrics (cost per transaction, NPS, revenue generated, errors avoided) measure outcomes. The decision-maker buys the latter—the former is an internal argument for the technical team.

Why learn:

Presenting only model metrics to a CFO is the most common mistake in AI projects that succeed technically but are rejected commercially. Translating the results is part of the consultant’s job.

Key concepts:

Outcome vs. output; business KPI; proxy metrics; the decision-maker buys results, not accuracy.

What it is:

A one-page business case includes: the problem to solve, the proposed solution (in plain language), estimated total investment, expected benefits and when it will pay off, key risks, and how to mitigate them. It fits on one page because the decision-maker can read it in two minutes.

Why learn:

A document that won’t fit on one page usually isn’t clear enough—the length constraint forces the consultant to take a position and cut the analytical fluff.

Key concepts:

Problem-solution-value-risk structure; brevity as a sign of clarity; approval as the document’s goal.

View Full
4.2~45 min

🌊 Wave-based roadmap: sequence and prioritize

Prove value before scaling: how to structure the roadmap around crawl/walk/run, respect dependencies, sequence by the best value-risk-learning tradeoff, and replan based on what the pilot taught you.

What it is:

Crawl is the first wave: the smallest possible scope, maximum learning, controlled risk. Walk expands coverage and automates more — based on the confidence gained from what crawl proved. Run scales what walk validated. Skipping waves is the leading cause of costly rollbacks in AI projects.

Why learn:

The wave isn’t just about pace—it’s a method for reducing risk. Each wave should deliver a concrete business result before moving forward, not just a technical deliverable.

Key concepts:

Incremental scope; results in waves; don’t advance without validation; inexpensive rollback.

What it is:

A dependency is anything that must be in place for a use case to work: clean, accessible data, system integrations, team capabilities, permissions, and governance. Building before resolving dependencies is building on sand.

Why learn:

The biggest waste in AI projects is building the model before you have the data. Mapping dependencies at the start of the roadmap prevents rework and keeps the schedule on track.

Key concepts:

Dependency tree; foundation before feature; data as a prerequisite; integrations as risk.

What it is:

For each use case, assess three axes: value (what is the impact if it works), risk (how likely it is to go wrong and the cost of failure), and learning (what this case teaches that will unlock others). The best candidate for the first wave has high learning value, visible impact, and manageable risk.

Why learn:

Without explicit criteria, the first project is chosen by whoever is most enthusiastic or most visible—not necessarily the most strategic. Criteria define what comes first and make the decision defensible.

Key concepts:

Value × risk matrix; learning as a sequencing criterion; defensible and reversible decision.

What it is:

Reusable capabilities are building blocks that, once created for one use case, become available for others: a clean data pipeline, an entity extraction module, a CRM integration. Planning the roadmap around reuse reduces the marginal cost of subsequent waves.

Why learn:

A consultant who thinks about the platform before the feature avoids starting from scratch with every new use case. That’s the difference between a project portfolio and a pile of silos.

Key concepts:

Build once, use many times; platform vs. point-to-point; decreasing marginal cost across waves.

What it is:

The visual roadmap places each use case on a timeline with clearly separated waves, validation milestones (go/no-go), visible dependencies, and an identified owner for each block. It doesn’t need to be an expensive tool: a simple table or diagram does the job.

Why learn:

Without shared visibility, each person has a different roadmap in mind. The visual artifact aligns expectations, makes reviews easier, and serves as an anchor for prioritization discussions.

Key concepts:

Alignment artifact; milestone ≠ delivery; one owner per wave; public visibility.

What it is:

The pilot always teaches you something the plan didn’t anticipate: worse-than-expected data, team resistance, more complex integration, or a use case with better results than expected. Replanning based on what the pilot revealed is the difference between a living roadmap and a document no one consults.

Why learn:

Projects that don’t update the roadmap after the first wave treat the plan as sacred—and miss the chance to apply the project’s most valuable lessons.

Key concepts:

Formal post-wave review; lessons learned inform the next plan; the plan is a hypothesis, not a decree.

View Full
4.3~45 min

🪜 Choosing the pyramid level for each case

Apply the deterministic/AI/agent pyramid to the specific case: when each level is the right answer, how to assess cost × risk × time at each level, and why to document the decision.

What it is:

Deterministic is enough when the rules are clear and stable, the volume doesn’t justify AI, errors are costly (100% accuracy required), or the data doesn’t exist yet. Rules, automations, filters, integrations, and SaaS solve most cases — at a drastically lower cost.

Why learn:

The impulse to go straight to AI overlooks the foundation, which solves problems more cheaply, quickly, and with fewer ways to fail. Training yourself to check the foundation first is the pyramid’s main goal.

Key concepts:

An explicit rule is better than an implicit model; predictable > intelligent when possible; baseline as the default.

What it is:

The middle of the pyramid (AI in one step of the workflow) is indicated when there is natural language, images, probabilistic predictions, or ambiguity that rules cannot cover at a reasonable cost. The workflow remains deterministic in the steps before and after — AI is used surgically only where rules don’t work.

Why learn:

Using AI for one step keeps the rest of the workflow reliable and auditable while solving what rules can’t. It’s the most common and balanced pattern for AI projects in production.

Key concepts:

Surgical AI; deterministic wrapper; output validation; fallback to a human.

What it is:

Agents are suited when the problem requires multiple chained, autonomous decisions; available tools (APIs, searches, code execution) are necessary; and the cost of an error per action is tolerable. They have the largest failure surface, highest cost, and greatest need for oversight.

Why learn:

Knowing when an agent is justified prevents using the most expensive and risky option when an AI workflow for one step would solve the problem. Agents are rare in typical consulting use cases — and should be.

Key concepts:

Autonomy requires oversight; tools expand risk; explicitly justify the top of the pyramid.

What it is:

Each level of the pyramid has a different profile: deterministic has low fixed cost and low risk; AI in one step has recurring API costs and medium risk; agents have high cost, high latency, and high risk. Presenting this comparison to the client with real estimates is part of the business case.

Why learn:

Making the trade-off explicit keeps the client from choosing the wrong level because they think “more AI = better results.” The decision should be informed, not driven by enthusiasm.

Key concepts:

Comparison table by level; fixed vs. variable costs; increasing operational risk; latency as a cost.

What it is:

The hybrid pattern keeps business logic in deterministic code (predictable, testable, inexpensive) and uses AI only where rules can’t solve the problem: classifying intent in a free-text field, extracting an entity from text, summarizing a long document. The result goes back into the deterministic workflow.

Why learn:

The hybrid combines the reliability of the foundation with the power of AI where it truly adds value. It’s the most recommended pattern for consulting in production—and what distinguishes a mature solution from a prototype.

Key concepts:

Deterministic as the container; AI as a replaceable module; testability preserved; explicit fallback.

What it is:

For each use case, explicitly record which level was chosen, why the levels below it won't work, and what criteria would justify moving up a level in the future. This record protects the consultant, aligns the team, and serves as a reference when the context changes.

Why learn:

Decisions without records are repeatedly revisited. Minimal documentation of the level choice creates project memory and allows for an informed review when the pilot returns with real data.

Key concepts:

Simplified ADR (architecture decision record); review criteria; decision as a revisable hypothesis.

View Full
4.4~45 min

🧪 Pilot, proof of concept, and success metrics

Define success before you start, distinguish a PoC from a pilot, set the smallest scope that teaches you something, keep a human in the loop, and decide whether to scale or kill it using go/no-go criteria.

What it is:

Before writing a line of code or configuring any tool, pilot success must be defined in numbers: which metric, what minimum acceptable value, and what timeline. Without that, the pilot ends with “that was nice” — which isn’t data you can use to decide anything.

Why learn:

Defining success after seeing the results guarantees confirmation bias. Setting the criteria in advance protects the objectivity of the evaluation and makes the go/no-go decision defensible to all parties.

Key concepts:

Pre-registration of the criteria; metrics vs. intuition; what will go in the results slides before you see the data.

What it is:

A PoC (proof of concept) answers "Is this technically possible?" with minimal effort—often without real data. A pilot answers "Does this work in this specific context?" using real data and real users, but with a limited scope. Production is when you scale with confidence. Skipping steps is the main reason for costly rollbacks.

Why learn:

Mixing up the stages creates false expectations: treating a PoC as a pilot leads to decisions based on lab data. Knowing how to distinguish them sets the right tone for conversations with the client.

Key concepts:

PoC = technical feasibility; pilot = operational validity; production = validated scale; each stage has a distinct deliverable.

What it is:

The pilot’s minimum scope is the smallest experiment that answers the project’s most important question. If the question is "Can the AI classify intent with 80% accuracy?", the pilot tests only that, with the minimum volume, without building the full integration before knowing the answer.

Why learn:

Scope creep is the main reason pilots take months and end without a clear answer. The smallest scope that teaches you the most is always the target—not the most complete, but the smartest.

Key concepts:

MLP (minimum learning pilot); main pilot question; don't build what doesn't answer the question.

What it is:

Human in the loop means every AI output has a human review point during the pilot. The failure plan defines what happens when the AI makes a mistake: who rolls it back, how quickly, and how operations continue without the solution. Pilots without a failure plan assume the system will never make a mistake.

Why learn:

The pilot is when uncertainty is highest. Keeping a human in the loop lets you catch errors before they reach the end customer and provides data for improvement. The failure plan ensures an error doesn’t interrupt the client’s operations.

Key concepts:

Human-in-the-loop; operational fallback; review SLA; supervision cost in the pilot.

What it is:

During the pilot, measure the same metric recorded in the baseline, using the same definition over an equivalent period. The before-and-after comparison is the only argument that turns perception into evidence and justifies the decision to scale or abandon the project.

Why learn:

Pilots that don't compare against a baseline produce stories, not data. Rigorous comparison is what separates an informed decision to scale from a decision driven by enthusiasm.

Key concepts:

Same metric definition before and after; comparable period; attributable change; evidence vs. story.

What it is:

Go/no-go is the formal decision at the end of the pilot: were the criteria met? If so, scale up—with what you learned and with adjustments to the roadmap. If not, kill it or pivot—without guilt, because pilots exist precisely to answer that question. Killing a pilot that didn’t work is a process success, not a project failure.

Why learn:

Without a formal go/no-go criterion, every pilot becomes production by inertia. The ability to kill a project that doesn’t work is what keeps the portfolio healthy and customer trust intact.

Key concepts:

Pre-registered criterion; stop without guilt; pivot vs. scale; explicit next steps in either case.

View Full
← Back to start Next track: Delivery →