PTENES
Skip to content
MODULE 4.6 · TEACHING MODE

🎬 Checkpoints & smooth review

You delegated to AI (all of Track 4). Now: how do you reviews without becoming a bottleneck? Let’s see what the review actually gives you, how to carefully get the human out of the way (and who reviews the agent that says "everything looks good"), and the trick that makes review enjoyable — the agent records a video of itself and narrates. Each term is explained as it comes up.

6
Topics
~45
Minutes
T4
Advanced
End
From the track
Progress: 0% 0 of 6

📖 Living glossary (read first — come back whenever you need to)

Here, you review the work the agent did on its own (AFK, GitHub Actions, queues). These are the new terms in this module—learn them:

Review — the moment when you look at what the agent produced and approve, adjust, or reject it. It's the "quality control" before the change enters the project.
Checkpoint — a stopping point where the work waits human approval to continue. Each checkpoint gives you confidence, but also turns you into a bottleneck.
PR (Pull Request) — the "bundle of changes" the agent opens on GitHub to propose code changes. It's where you review before accepting (doing the merge).
Merge — merge the PR change into the main code. “Auto-merge” means the change goes in by itself, without you pressing the button.
Observability — your ability to see inside how the system (and the agent) is working. Without it, you delegate blindly.
TTS (Text-to-Speech) — "text to speech": software that reads text aloud. The agent uses it to narrate A video of the change itself.
Walkthrough — a “guided walkthrough”: a short video in which the agent moves through the change showing that it works.
AFK — Away From Keyboard ("away from the keyboard"): the agent works on its own while you're away. The review is how you reconnect with that work.
1

🔍 What the review gives you

🧠 Imagine it this way: you’re a newspaper editor. Reading the reporter's copy before publication isn't just about "catching errors" — it's also about helping you understand how that reporter thinks, where they tend to stumble, and how to improve the next assignment. The review has two functions, not one.

When the agent works alone—in AFK, in PRs, in task queues—one question remains: what is the review? Matt Pocock gives a two-part answer. The first function is obvious: the review is a gate for dangerous things — leaked secrets, security holes, changes that could, for example, expose Claude Code’s own source code. Here, the review is a protective barrier, and it matters.

But there’s a second function almost everyone misses—and Pocock insists on it: the review gives you insight in your own system. When you watch the agent work, you discover how your harness (the plumbing that produces the code) behaves in practice. It’s observability: you improve the harness over time, and to do that you need to see what's happening. In his words: "we're reviewing not just the code, but the system that produces the code." O common mistake is treating review only as bug hunting and, anxious to move faster, cutting everything — you save time and go blind. Without observability, you delegate blindly and never tune the engine that produces everything.

REVIEW you look at what the agent did 1 · GATE blocks the dangerous stuff security · secrets leak source code 2 · WINDOW insight in the system observability you improve the harness Don’t cut the review just to save time: you’d miss the window.

Review isn’t just bug hunting: it’s also your window into understanding and fine-tuning the system that produces the code.

Conceptual illustration: a futuristic review dashboard with a security gate on one side and an observability window on the other

⚠️ Common beginner mistake

Trying to “automate everything” and turn off review on day 1. You may gain speed—but you become blind to your own system and don’t know how to improve it. Start by reviewing a lot; gradually remove checkpoints as you gain confidence.

In one sentence: the review is a gate (it blocks dangerous things) and a window (it shows you the system)—don't throw the window away.

Going deeper (optional): why "leak Claude Code's source code"?

When the agent runs inside your environment, it has access to sensitive files—API keys, production secrets, and, in some setups, parts of the harness itself. A poorly reviewed PR could accidentally paste a secret into a public file or open a vulnerability. That’s why the security gate is the checkpoint you cut last: the cost of a mistake here is too high.

2

➡️ Push the checkpoint to the right

🧠 Imagine it this way: A factory conveyor belt. Before, an inspector stopped EVERY part to check it. Expensive and slow. The solution wasn’t to remove the inspector — it was to move it to the end of the conveyor, where it checks the finished product. The piece flows; control happens as close to the output as possible.

Imagine the timeline of a task: telemetry detects a bug → it becomes a issue → the agent investigates → proposes the fix → opens the PR → gets into the code (merge). At each of these arrows, you could add a checkpoint human. Pocock’s question is: where does it deliver more? The answer is the key phrase of the learning path: "push human-in-the-loop checkpoints further toward the final output" — push the human approval points farther and farther to the right, toward the final output.

Why? Because a checkpoint in the start holds up everything else: nothing moves without you. The later the checkpoint, the more work the AI has already done on its own by the time it reaches you — and the more each minute of yours is worth. The bottleneck stops being you. The common mistake is the opposite: people who review every micro-step (every command, every file) and, without realizing it, rebuild the manual work they wanted to delegate. Remember the queue from module 4.4: you’re the king who prioritizes and approves the result—not the soldier checking every move.

bug issue fix PR merge ✗ too early ✓ checkpoint here The farther to the right the checkpoint, the more the AI can do on its own and the less you become the bottleneck.

Quick recall: "push the checkpoint to the right" means…

In one sentence: move your checkpoints to the right (near the output) so you’re no longer the bottleneck—while still approving the result.

3

✂️ When to remove the human

🧠 Imagine it this way: changing the coffee brand in the company break room doesn't need to go through the CEO. Changing the layoff policy does. The secret is matching the approval level to the size of the risk — not checking everything the same way.

The goal, Pocock says, is to remove checkpoints wherever possible — carefully. Not every change deserves the same scrutiny. His example is clear: a PR that’s just a refactoring internal — reorganizes the code, but doesn't change the behavior visible — a natural candidate for having the checkpoint removed. The AI itself can evaluate and say: "this doesn't need human review; it's just internal cleanup."

The practical rule is to classify by blast radius. Low risk (a refactor that doesn’t change behavior, a text tweak, a safe dependency update) → a good candidate for auto-merge. High risk (anything that touches security, user data, money, or changes behavior the customer can see) → keep the human involved. The common mistake is binary: either review everything (a bottleneck) or nothing (dangerous). The way forward is gradient: start by reviewing a lot, then loosen the rules for low-risk items as you gain confidence in the system. This connects to module 4.5 (telemetry → issue → fix → auto-merge tag or ping the human): the “auto-merge or human” decision is exactly this risk judgment.

Illustration: two PR conveyor belts, one low-risk belt going straight to automatic merge and another high-risk belt stopping at a human checkpoint

✓ Let it run (auto-merge)

  • • Internal refactor (doesn’t change behavior).
  • • Text edits, formatting, comments.
  • • Dependency bump tested by CI.
  • • Change covered by passing tests.

✗ Keep the human involved

  • • Security, authentication, permissions.
  • • User data or payments.
  • • Visible behavior change.
  • • Anything irreversible.
Agent PR (coming from AFK) classifies by risk LOW RISK → auto-merge refactor that doesn't change behavior HIGH RISK → human checkpoint security · data · $ · behavior approved closer to the final output → A gradient, not an on/off switch: let low-risk tasks through, hold back high-risk ones — and move the checkpoint to the right.

Match the approval level to the radius of impact: low risk flows on its own, high risk reaches a human down the line.

In one sentence: remove the checkpoint where risk is low (refactoring without changing behavior); keep it where risk is high.

4

🛡️ Who reviews the agent

🧠 Imagine it this way: A student who grades their own test and says, “I got a 10.” Maybe they did. But you can’t trust them blindly — you need to sample a few tests to calibrate whether their “10” matches a “real 10.”

Here's the paradox Pocock raises in a sharp phrase: "who reviews the AI that says it's fine?" — who reviews the AI that says everything is fine? If you let the AI itself decide that a PR “doesn’t need review,” you’ve created a judge that judges its own work. You can’t trust that blindly. The solution isn’t to turn off self-evaluation—it’s to audit a sample. You still need to check some of the PRs the agent declared okay, and keep improving that judgment over time.

It’s the same idea from topic 1, now in action: "we're reviewing not just the code, but the SYSTEM that produces the code." You're not just checking that specific PR—you're calibrating the judgment system of the agent. Every time you catch an "it’s fine" that actually wasn’t, you adjust the skill, prompt, or auto-merge rule to get it right next time. The common mistake is the leap of faith: giving auto-merge total without ever sampling. Without sampling, you can't know whether the agent is calibrated — you only find out after a serious error has slipped through.

PRs that the AI said “looks good” ? human audit sample some PRs calibrates the system

In one sentence: the AI may say "looks good," but you still sample some of these PRs — to review the system, not just the code.

5

🎥 Video walkthrough + TTS

🧠 Imagine it this way: instead of getting an architecture blueprint full of numbers to decipher, you get a 40-second video: someone walks through the rooms of the house and narrates, “here’s the kitchen; notice that the sink has moved.” In seconds, you understand—without reading anything.

This is Pocock’s favorite trick for making review enjoyable, and it applies to any change to front end. The idea: instead of handing you a PR with 30 changed files for you to read line by line, the AI itself records a video walkthrough — it moves through the change using the actual application (clicks the buttons, navigates the new screens, shows the before and after). Then it calls a TTS e narrates over it, explaining what changed. The PR now has a video of it working attached.

Why is this so powerful? Because your brain processes a narrated video much faster than a diff of code. You check the behavior (does it work? does it look good? is it what I asked for?) in seconds, instead of inferring its behavior from the code. Pocock sums up the spirit of all this: "optimize for human review, make it faster — we're scratching the surface." Optimize for human review and make it faster—we’re only scratching the surface of what’s possible. The common mistake is thinking this is fluff: in practice, it’s what keeps you able to approve many PRs a day without burning out.

Illustration: a Pull Request with an attached video player showing the agent navigating the interface and narrating the change
1 · run thechange 2 · records thescreen video 3 · TTS narrateson top PR with video narrated, working You watch 40s and approve—instead of reading 30 files of diffs.

🔬 Worked example: the same PR, two reviews

Task delivered by AI in AFK mode: "redesign the checkout screen (front end)". Same change, two ways to review it:

Review without video

You open the PR, see "27 files changed," read the CSS and JSX diff, trying to imagine what the screen looked like. 25 minutes later, you still aren’t sure whether the payment button works. It’s exhausting, and you approve on a hunch.

Review with walkthrough + TTS

The PR has a 45s video: the AI fills the cart, opens the new checkout, clicks “pay,” and narrates, “the button is now fixed to the bottom of the screen on mobile.” You can see that it works and approve it with confidence — in less than 1 minute.

Result: the same AI work, but your review time plummeted and confidence rose—because you evaluated behavior, not code.

In one sentence: the AI records a video of itself using the change and narrates it with TTS — you review behavior in seconds, not code in minutes.

Going deeper (optional): how does the agent record this video in practice?

Usually through a browser automation tool (such as Playwright) that the agent controls: it starts the application, opens a "headless" browser, runs a scripted series of clicks through the affected screens, and records the screen as a video. In parallel, it generates a narration script based on what changed in the PR and passes that text through TTS to produce the audio track. Finally, it combines the video and audio (e.g., with ffmpeg) and attaches them to the PR. Pocock calls this "fluid review" and says we’re only "scratching the surface" — there’s much more you can do.

6

⚡ Faster review with AI

🧠 Imagine it this way: instead of the teacher grading 20 identical tests one by one, they read them all and write a handout: "these were the 3 most common mistakes — study this." You improve the whole class at once, and yourself as a teacher.

The last piece is optimizing the human review itself — because, as Pocock says, "GitHub wasn't built for the agentic era". The tool was designed for a few large PRs from humans, not dozens of small PRs from agents. The smart move: instead of reviewing 20 small fixes One at a time, ask the AI to generate a "teach skill"-style HTML (remember track 3?) that distills the common patterns of those bugs—what repeated, what the root cause was, what to learn. You review the summary, not the 20 pages.

Notice what this does: the goal is "optimize so you can improve yourself and the system". You don’t just approve faster — you learn the pattern behind the errors and adjust the harness so they don’t happen again. That closes the loop on Topic 1: review the system, not just the code. The common mistake is treating review as a mindless, repetitive task; here it becomes a tool for continuous improvement. Use the diagnosis below to decide how to speed up your review for any batch of PRs — copy and paste it when the backlog fills up:

review-fluido.txt
REVIEW FLUIDO — como acelerar o review dos PRs do agente

1) CLASSIFIQUE por risco antes de abrir o diff:
   [ ] baixo risco (refactor sem mudar comportamento, texto, dep) -> auto-merge
   [ ] alto risco (seguranca, dados, $, comportamento visivel) -> humano

2) EMPURRE o checkpoint pra direita:
   [ ] o humano aprova perto da SAIDA, nao em cada micro-passo

3) PECA o walkthrough em qualquer mudanca de front-end:
   "Grave um video navegando pela mudanca e narre com TTS o que mudou;
    anexe ao PR." -> reviso comportamento em segundos, nao diff em minutos

4) AGREGUE lotes em vez de revisar 1 a 1:
   "Em vez de 20 fixes soltos, gere um HTML tipo teach skill com os
    padroes comuns desses bugs (causa raiz + o que aprender)."

5) AUDITE a auto-avaliacao da IA:
   [ ] amostre alguns PRs que a IA disse "ta ok"  (quem revisa a IA?)
   [ ] ajuste skill/prompt/regra de merge -> revise o SISTEMA, nao so o codigo

In one sentence: use AI to aggregate and narrate what needs reviewing—that way you approve faster and also improve the system that produces the code.

Quick recall: to review 20 small fixes from the agent without burning out, Pocock suggests…

🧾 Module Summary

✓
Review = gate + window — blocks the dangerous stuff AND gives you system observability. Don’t cut the window just for speed.
✓
Move the checkpoint to the right — approve near the final output so you stop being the bottleneck.
✓
Remove the human where the risk is low — refactor without changing behavior → auto-merge; security/data/$ → human.
✓
Who reviews the agent? — sample a few PRs the AI said were “okay”; calibrate the system, not just the code.
✓
Walkthrough + TTS — the AI records a video of the change and narrates it; you review the behavior in seconds.

You completed Track 4!

Next stop: Track 5 — Ready-to-Copy Solutions. Everything you saw (grill-me, AFK setup, Actions, crons, smooth review) becomes a ready-made recipe you can paste into your project today.