📖 Living glossary (read first — come back whenever you need to)
This learning path has already established the core vocabulary (model, agent, skill, AFK, sandbox, queue). Here are the terms new from this module — remember these:
💡 You don't need the expensive model
🧠 Imagine it this way: a mediocre detective in a well-lit neighborhood finds more crimes than a brilliant detective blindfolded in a dark room. It's not just talent — it's being in the right place, with the right flashlight. Finding a deep bug is the same: the place and the light (the harness) matter just as much as the model's "IQ".
There’s a tempting myth: “if I use the smartest (and most expensive) model, it’ll find those deep bugs that no other model can find.” In the conversation, it was said that one model found deep bugs and failures of security that other models hadn’t seen. The easy reaction is: “See? You need the top model.” Pocock disagrees at the root. In his words: "you could find these bugs with cheaper models if you looked in the right places and gave it the right prompt/harness". In other words, what the “other models” lacked wasn’t intelligence — it was direction.
This connects directly to what you saw in Track 1 (Token Economy): a harness well, you let a “dumber” model do the same work. Here, the idea is the same, applied to bug discovery. Instead of paying a lot for a genius model to run once, you build a system that puts a simple model in the right place, repeatedly. The reason matters: an expensive model is a cost that grows with every call; a well-targeted system is an asset that pays off every day. The common mistake is confusing “it found a hard bug” with “I always need the harder model”—when what actually found the bug was where e how you looked.
The secret isn’t the model’s brain—it’s where you point the flashlight (the harness).
⚠️ Common beginner mistake
Conclude, "I need the most expensive model to find serious bugs." Almost always, the failure was within reach of a cheaper model — all that was missing was point it to the right place, with the right prompt. Pay for the system, not the brain.
In one sentence: you find deep bugs with the right place + the right prompt—not with the most expensive model.
Going deeper (optional): why is the cheap model enough here?
Finding a bug doesn't require reasoning about the entire project at once — it requires looking at a small, well-defined piece with a clear question ("does this function validate the input?"). When you narrow the scope and give the right prompt, the problem becomes easy enough for a simple model. An expensive model shines on large, ambiguous tasks; sliced-up review is the opposite — so you save money without losing quality.
⏰ Security review cron
🧠 Imagine it this way: A guard making the rounds of a building. They don’t check all 30 floors every night — they’d do a poor job. They check one floor per night, calmly, and covered the whole building in a month, several times over. The security review cron is that watchman.
Pocock’s concrete recipe: run a cron journal of security review that checks, every day, a new part of the repository, using a relatively simple model. The keyword is rotating: you don’t try to scan the entire project at once (expensive, shallow, and blowing up the context window). You break it into pieces. Today, the authentication folder; tomorrow, payments; then uploads. In a few weeks, the entire repo has been audited—and the cycle starts again, catching what’s changed.
Why scheduled and not “when I remember”? Because the lesson Pocock draws from an incident is direct: "what did you learn? That there are security issues in your code — you should have something that runs and checks for more in the future." Security can’t depend on your memory on a busy day. The reason of the simple model: a sliced review is a small, well-scoped task—exactly where a cheap model is enough (topic 1). The common mistake is setting up a cron job that reviews the entire repo at once: it gets expensive, creates noise, and nobody reads the huge report. Small, daily, rotating, actionable.
Below, the copyable artifact: a line of crontab that runs a headless security review agent every day at 6 a.m., pointed at “the day’s slice.” Add it to your crontab (crontab -e) and adjust the repo path:
# Edite com: crontab -e
# Roda TODO DIA às 06:00 — checa uma fatia rotativa do repo (dia-do-mês % nº de pastas).
0 6 * * * cd /caminho/do/repo && DIRS=(auth payments uploads api db) && TARGET=${DIRS[$(( $(date +\%d) \% ${#DIRS[@]} ))]} && claude -p "Faça um SECURITY REVIEW so da pasta ./$TARGET. Liste cada falha (arquivo:linha), severidade (alta/media/baixa) e a correcao sugerida. Se nada grave, responda 'OK'. Saida em markdown." --model claude-haiku-4-5 --allowedTools "Read,Grep,Glob" >> ~/security-reviews/$(date +\%F)-$TARGET.md 2>&1
Quick recall: why does the security review cron check only one slice per day?
In one sentence: A scheduled guard who reviews one slice per day covers the entire building without costing a fortune.
📡 Telemetry → issue → fix
🧠 Imagine it this way: your home’s smoke alarm doesn’t just beep—it automatically calls the fire department, which arrives already knowing which room. Telemetry → issue → fix means setting up this pipeline: a production error triggers a chain that ends (almost) on its own with a fix.
Here, self-improvement stops being proactive (the cron) and becomes reactive, but automated. Pocock’s example: a bug in telemetry appears in a tool like the Sentry. Instead of reading, copying, and manually opening an issue, the system automatically creates an issue and the brand with a label (label) like agent:explore. That label triggers an agent—the exact queue pattern from module 4.4.
The agent then returns structured data: can it be fixed now, or does it need a human? If it’s straightforward, it implements the fix and opens the PR, and the flow moves on to review—eventually with a auto-merge (runs on its own) or a ping for you to decide. Pocock puts it all together in this phrase: "push the human-in-the-loop checkpoints as close as possible to the final result." In other words: the machine handles the whole process; you step in near the finish line, where your attention goes further. The common mistake is letting the pipeline auto-merge everything — with no human gate on changes that affect behavior, you trade manual work for silent risk.
🔬 Worked example: from a Sentry error to a fix
End-to-end scenario — one TypeError intermittent at checkout:
- Telemetry: Sentry groups 312 occurrences of
Cannot read 'total' of undefinedincheckout.ts. - Automatic issue: A webhook creates the issue
#812with the stack trace and applies the labelagent:explore. - Agent explores: triggered by the label, returns structured data— cause: empty cart not handled; difficulty: trivial; can fix AFK: yes.
- Fix + PR: the agent adds the guard
cart?.total ?? 0+ one test, open the PR#813linking the issue. - Checkpoint at the end: if it changes behavior, it does NOT auto-merge — it pings you, and you review it in 30s and approve.
You only showed up at step 5—the checkpoint was pushed closer to the final result.
In one sentence: a production error automatically triggers the issue→agent→fix chain; you only make decisions near the finish line.
🔒 The root cause of the bug
🧠 Imagine it this way: your bike gets stolen. You buy another bike. It gets stolen again. You buy another. The right question isn’t “which bike to buy” — it’s "why is it stealable? Where's the lock?". Fixing the bug is buying another bike. Finding the root cause is installing the lock.
This is the module’s turning point, and Pocock sums it up with an image that sticks: "if someone keeps stealing your bike, maybe get a lock." Fixing the bug that appeared is treating the symptom. The question that truly improves the system is about the root cause (root cause): why did this bug exist for so long without you noticing? What was missing from your process — a test, a review, a lint check — that let this slip through until it became an incident?
Notice the difference in level. "I found a security bug and fixed it" stops at the bike. "There are security flaws in my code, so I'm going to put together something that runs and checks for more in the future" installs the lock—it’s literally Topic 2’s cron job born from an incident. Each bug becomes more than just a fix; it becomes a question: Which system guardrail let this through? The answer becomes a test suite, automated review, or refactor. The common mistake is stopping happily at the fix: you close the issue, feel like you solved it — and the same kind of bug comes back through another door, because the lock was never bought.
Going deeper (optional): the "5 whys"
A classic technique for getting to the root cause is to ask “why?” five times. Checkout bug → why? Empty cart not handled → why? Nobody tested this case → why? There was no empty-cart test → why? The suite doesn't cover edge cases → why? We don't have a cron/checklist that checks for these cases. By the fifth “why,” you’re no longer looking at the bug — you’re looking at the system that let it through. That’s where you buy the padlock.
In one sentence: don't buy another bike—ask why it keeps getting stolen and install the lock.
♻️ Self-improvement loops
🧠 Imagine it this way: A gym where every workout is recorded, and the record adjusts the next workout. You don’t get stronger “for free” — you get stronger because the training system learns from every session. Self-improvement loops are that idea applied to your code.
Putting it all together: a self-improving system is made up of several loops that reinforce each other. Pocock names three concrete ingredients: test suites (that catch regressions before they reach the user), human reviews (that provide insight into the system, not just the code) and refactor (which reduces the chance of the next bug appearing). Each incident feeds these loops: it becomes a test, a review checklist, or a cleaner part of the code.
The critical detail—and this is where it connects to module 4.4 (Loops × Queues)—is that “self-improving” no means "leave an infinite loop running on its own and disappear." It means setting up triggers (cron, telemetry, label) that put tasks in a queue, and feedback loops that make each pass better than the last. You’re still in charge: you prioritize, review the risky parts, and adjust the triggers. The common mistake is imagining a 100% autonomous system that fixes itself forever without supervision—that doesn’t exist; what does exist is a system that does the repetitive work and calls you in when it matters.
In one sentence: detect → fix → prevent → learn, and each cycle leaves the system stronger than the last.
🛠️ Review the system
🧠 Imagine it this way: a factory doesn’t inspect only the products that go out—it inspects the assembly line that produces them. If the line is crooked, every product comes out crooked. Reviewing only the code is checking the products; reviewing the system is straightening the line.
The thesis that closes the module is Pocock’s most important statement here: "review not just the code — also review the SYSTEM that produces the code." In the agentic era, the agent writes most of the code. So your review focus moves up a level: less "is this line correct?" and more "the process that generated this line trustworthy?” The prompts, skills, crons, review gates — this is the plumbing that, when tuned, keeps the next bug from ever being born.
This also answers a dangerous question that comes up in the next module (4.6): "who reviews the AI that says everything is okay?" Response: you, reviewing the system. That’s why it’s worth checking some PRs from time to time that the agent assured you are okay—not out of paranoid distrust, but to improve the system over time. Below is a practical summary in the form of a copyable checklist — run it whenever a bug slips through, to turn the fix into a system improvement:
Bug escapou? Antes de fechar a issue, COMPRE O CADEADO: [ ] CAUSA RAIZ — por que esse bug existiu tanto tempo sem ninguém notar? [ ] TESTE — adicionei um teste que pegaria esse caso na próxima vez? [ ] DETECÇÃO — meu cron de security/review olharia essa parte? Se não, ajusto o rodízio. [ ] PROCESSO — que gate (review, lint, label) deixou isso passar? Reforço ele. [ ] SISTEMA — revisei o código OU o sistema que produz o código? (era pra ser os dois) Se você só corrigiu o bug, comprou outra bike. Instale o cadeado.
Quick recall: what does "review the system that produces the code" mean?
In one sentence: review not just the code, but the assembly line that produces it — that’s where the next bug dies before it’s born.
🧾 Module Summary
Next module:
4.6 — Checkpoints & smooth review: how to make review painless, push the checkpoint to the right, and record video+TTS of the agent at work.