📖 Living glossary (read first — come back whenever you need to)
Here, you review the work the agent did on its own (AFK, GitHub Actions, queues). These are the new terms in this module—learn them:
🔍 What the review gives you
🧠 Imagine it this way: you’re a newspaper editor. Reading the reporter's copy before publication isn't just about "catching errors" — it's also about helping you understand how that reporter thinks, where they tend to stumble, and how to improve the next assignment. The review has two functions, not one.
When the agent works alone—in AFK, in PRs, in task queues—one question remains: what is the review? Matt Pocock gives a two-part answer. The first function is obvious: the review is a gate for dangerous things — leaked secrets, security holes, changes that could, for example, expose Claude Code’s own source code. Here, the review is a protective barrier, and it matters.
But there’s a second function almost everyone misses—and Pocock insists on it: the review gives you insight in your own system. When you watch the agent work, you discover how your harness (the plumbing that produces the code) behaves in practice. It’s observability: you improve the harness over time, and to do that you need to see what's happening. In his words: "we're reviewing not just the code, but the system that produces the code." O common mistake is treating review only as bug hunting and, anxious to move faster, cutting everything — you save time and go blind. Without observability, you delegate blindly and never tune the engine that produces everything.
Review isn’t just bug hunting: it’s also your window into understanding and fine-tuning the system that produces the code.
⚠️ Common beginner mistake
Trying to “automate everything” and turn off review on day 1. You may gain speed—but you become blind to your own system and don’t know how to improve it. Start by reviewing a lot; gradually remove checkpoints as you gain confidence.
In one sentence: the review is a gate (it blocks dangerous things) and a window (it shows you the system)—don't throw the window away.
Going deeper (optional): why "leak Claude Code's source code"?
When the agent runs inside your environment, it has access to sensitive files—API keys, production secrets, and, in some setups, parts of the harness itself. A poorly reviewed PR could accidentally paste a secret into a public file or open a vulnerability. That’s why the security gate is the checkpoint you cut last: the cost of a mistake here is too high.
➡️ Push the checkpoint to the right
🧠 Imagine it this way: A factory conveyor belt. Before, an inspector stopped EVERY part to check it. Expensive and slow. The solution wasn’t to remove the inspector — it was to move it to the end of the conveyor, where it checks the finished product. The piece flows; control happens as close to the output as possible.
Imagine the timeline of a task: telemetry detects a bug → it becomes a issue → the agent investigates → proposes the fix → opens the PR → gets into the code (merge). At each of these arrows, you could add a checkpoint human. Pocock’s question is: where does it deliver more? The answer is the key phrase of the learning path: "push human-in-the-loop checkpoints further toward the final output" — push the human approval points farther and farther to the right, toward the final output.
Why? Because a checkpoint in the start holds up everything else: nothing moves without you. The later the checkpoint, the more work the AI has already done on its own by the time it reaches you — and the more each minute of yours is worth. The bottleneck stops being you. The common mistake is the opposite: people who review every micro-step (every command, every file) and, without realizing it, rebuild the manual work they wanted to delegate. Remember the queue from module 4.4: you’re the king who prioritizes and approves the result—not the soldier checking every move.
Quick recall: "push the checkpoint to the right" means…
In one sentence: move your checkpoints to the right (near the output) so you’re no longer the bottleneck—while still approving the result.
✂️ When to remove the human
🧠 Imagine it this way: changing the coffee brand in the company break room doesn't need to go through the CEO. Changing the layoff policy does. The secret is matching the approval level to the size of the risk — not checking everything the same way.
The goal, Pocock says, is to remove checkpoints wherever possible — carefully. Not every change deserves the same scrutiny. His example is clear: a PR that’s just a refactoring internal — reorganizes the code, but doesn't change the behavior visible — a natural candidate for having the checkpoint removed. The AI itself can evaluate and say: "this doesn't need human review; it's just internal cleanup."
The practical rule is to classify by blast radius. Low risk (a refactor that doesn’t change behavior, a text tweak, a safe dependency update) → a good candidate for auto-merge. High risk (anything that touches security, user data, money, or changes behavior the customer can see) → keep the human involved. The common mistake is binary: either review everything (a bottleneck) or nothing (dangerous). The way forward is gradient: start by reviewing a lot, then loosen the rules for low-risk items as you gain confidence in the system. This connects to module 4.5 (telemetry → issue → fix → auto-merge tag or ping the human): the “auto-merge or human” decision is exactly this risk judgment.
✓ Let it run (auto-merge)
- • Internal refactor (doesn’t change behavior).
- • Text edits, formatting, comments.
- • Dependency bump tested by CI.
- • Change covered by passing tests.
✗ Keep the human involved
- • Security, authentication, permissions.
- • User data or payments.
- • Visible behavior change.
- • Anything irreversible.
Match the approval level to the radius of impact: low risk flows on its own, high risk reaches a human down the line.
In one sentence: remove the checkpoint where risk is low (refactoring without changing behavior); keep it where risk is high.
🛡️ Who reviews the agent
🧠 Imagine it this way: A student who grades their own test and says, “I got a 10.” Maybe they did. But you can’t trust them blindly — you need to sample a few tests to calibrate whether their “10” matches a “real 10.”
Here's the paradox Pocock raises in a sharp phrase: "who reviews the AI that says it's fine?" — who reviews the AI that says everything is fine? If you let the AI itself decide that a PR “doesn’t need review,” you’ve created a judge that judges its own work. You can’t trust that blindly. The solution isn’t to turn off self-evaluation—it’s to audit a sample. You still need to check some of the PRs the agent declared okay, and keep improving that judgment over time.
It’s the same idea from topic 1, now in action: "we're reviewing not just the code, but the SYSTEM that produces the code." You're not just checking that specific PR—you're calibrating the judgment system of the agent. Every time you catch an "it’s fine" that actually wasn’t, you adjust the skill, prompt, or auto-merge rule to get it right next time. The common mistake is the leap of faith: giving auto-merge total without ever sampling. Without sampling, you can't know whether the agent is calibrated — you only find out after a serious error has slipped through.
In one sentence: the AI may say "looks good," but you still sample some of these PRs — to review the system, not just the code.
🎥 Video walkthrough + TTS
🧠 Imagine it this way: instead of getting an architecture blueprint full of numbers to decipher, you get a 40-second video: someone walks through the rooms of the house and narrates, “here’s the kitchen; notice that the sink has moved.” In seconds, you understand—without reading anything.
This is Pocock’s favorite trick for making review enjoyable, and it applies to any change to front end. The idea: instead of handing you a PR with 30 changed files for you to read line by line, the AI itself records a video walkthrough — it moves through the change using the actual application (clicks the buttons, navigates the new screens, shows the before and after). Then it calls a TTS e narrates over it, explaining what changed. The PR now has a video of it working attached.
Why is this so powerful? Because your brain processes a narrated video much faster than a diff of code. You check the behavior (does it work? does it look good? is it what I asked for?) in seconds, instead of inferring its behavior from the code. Pocock sums up the spirit of all this: "optimize for human review, make it faster — we're scratching the surface." Optimize for human review and make it faster—we’re only scratching the surface of what’s possible. The common mistake is thinking this is fluff: in practice, it’s what keeps you able to approve many PRs a day without burning out.
🔬 Worked example: the same PR, two reviews
Task delivered by AI in AFK mode: "redesign the checkout screen (front end)". Same change, two ways to review it:
Review without video
You open the PR, see "27 files changed," read the CSS and JSX diff, trying to imagine what the screen looked like. 25 minutes later, you still aren’t sure whether the payment button works. It’s exhausting, and you approve on a hunch.
Review with walkthrough + TTS
The PR has a 45s video: the AI fills the cart, opens the new checkout, clicks “pay,” and narrates, “the button is now fixed to the bottom of the screen on mobile.” You can see that it works and approve it with confidence — in less than 1 minute.
Result: the same AI work, but your review time plummeted and confidence rose—because you evaluated behavior, not code.
In one sentence: the AI records a video of itself using the change and narrates it with TTS — you review behavior in seconds, not code in minutes.
Going deeper (optional): how does the agent record this video in practice?
Usually through a browser automation tool (such as Playwright) that the agent controls: it starts the application, opens a "headless" browser, runs a scripted series of clicks through the affected screens, and records the screen as a video. In parallel, it generates a narration script based on what changed in the PR and passes that text through TTS to produce the audio track. Finally, it combines the video and audio (e.g., with ffmpeg) and attaches them to the PR. Pocock calls this "fluid review" and says we’re only "scratching the surface" — there’s much more you can do.
⚡ Faster review with AI
🧠 Imagine it this way: instead of the teacher grading 20 identical tests one by one, they read them all and write a handout: "these were the 3 most common mistakes — study this." You improve the whole class at once, and yourself as a teacher.
The last piece is optimizing the human review itself — because, as Pocock says, "GitHub wasn't built for the agentic era". The tool was designed for a few large PRs from humans, not dozens of small PRs from agents. The smart move: instead of reviewing 20 small fixes One at a time, ask the AI to generate a "teach skill"-style HTML (remember track 3?) that distills the common patterns of those bugs—what repeated, what the root cause was, what to learn. You review the summary, not the 20 pages.
Notice what this does: the goal is "optimize so you can improve yourself and the system". You don’t just approve faster — you learn the pattern behind the errors and adjust the harness so they don’t happen again. That closes the loop on Topic 1: review the system, not just the code. The common mistake is treating review as a mindless, repetitive task; here it becomes a tool for continuous improvement. Use the diagnosis below to decide how to speed up your review for any batch of PRs — copy and paste it when the backlog fills up:
REVIEW FLUIDO — como acelerar o review dos PRs do agente
1) CLASSIFIQUE por risco antes de abrir o diff:
[ ] baixo risco (refactor sem mudar comportamento, texto, dep) -> auto-merge
[ ] alto risco (seguranca, dados, $, comportamento visivel) -> humano
2) EMPURRE o checkpoint pra direita:
[ ] o humano aprova perto da SAIDA, nao em cada micro-passo
3) PECA o walkthrough em qualquer mudanca de front-end:
"Grave um video navegando pela mudanca e narre com TTS o que mudou;
anexe ao PR." -> reviso comportamento em segundos, nao diff em minutos
4) AGREGUE lotes em vez de revisar 1 a 1:
"Em vez de 20 fixes soltos, gere um HTML tipo teach skill com os
padroes comuns desses bugs (causa raiz + o que aprender)."
5) AUDITE a auto-avaliacao da IA:
[ ] amostre alguns PRs que a IA disse "ta ok" (quem revisa a IA?)
[ ] ajuste skill/prompt/regra de merge -> revise o SISTEMA, nao so o codigo
In one sentence: use AI to aggregate and narrate what needs reviewing—that way you approve faster and also improve the system that produces the code.
Quick recall: to review 20 small fixes from the agent without burning out, Pocock suggests…
🧾 Module Summary
You completed Track 4!
Next stop: Track 5 — Ready-to-Copy Solutions. Everything you saw (grill-me, AFK setup, Actions, crons, smooth review) becomes a ready-made recipe you can paste into your project today.