PTENES
Skip to content
MODULE 4.5 · "TEACHING MODE"

♻️ Self-improving systems

Finding deep bugs isn't magic from an expensive model. Matt Pocock shows how to build a system that takes care of itself: one cron daily security review that looks at a new part of the code each day, telemetry that becomes an issue that becomes a fix—and, above all, the right question: "if they keep stealing your bike, get a lock". Here you learn to review not just the code, but the system that produces the code.

6
Topics
~40
Minutes
T1-T4
Prerequisite
Practice
Type
Progress: 0% 0 of 6

📖 Living glossary (read first — come back whenever you need to)

This learning path has already established the core vocabulary (model, agent, skill, AFK, sandbox, queue). Here are the terms new from this module — remember these:

Self-improving system — a setup that, on its own and on a recurring basis, finds problems in your code and helps you fix them and prevent them from happening again. The code doesn’t “fix itself”; the process a loop that keeps getting better.
Cron — a scheduler: “run this command every day at 6 a.m.” The “crontab line” is a line of text that describes when e what running. It’s the heart of tasks that happen without you pressing anything.
Security review — security review: inspect the code for vulnerabilities (exposed passwords, injection, incorrect permissions) before that someone might abuse them.
Telemetry — the data your application sends about how it’s running (errors, slowdowns, exceptions). Tools like Sentry collect this in production.
Issue — a “ticket” on record (in GitHub, for example): a task or bug noted down for someone—human or agent—to resolve.
Root cause — the underlying reason for the problem, not the symptom. Not “bug X happened,” but “why was it there so long without anyone noticing?”
Test suite — the project's set of automated tests. Each test checks that one part works; running the suite tells you when something broke.
Rotating — that changes targets with each run. A rotating review cron checks one folder today, another tomorrow—covering the entire repo over time.
1

💡 You don't need the expensive model

🧠 Imagine it this way: a mediocre detective in a well-lit neighborhood finds more crimes than a brilliant detective blindfolded in a dark room. It's not just talent — it's being in the right place, with the right flashlight. Finding a deep bug is the same: the place and the light (the harness) matter just as much as the model's "IQ".

There’s a tempting myth: “if I use the smartest (and most expensive) model, it’ll find those deep bugs that no other model can find.” In the conversation, it was said that one model found deep bugs and failures of security that other models hadn’t seen. The easy reaction is: “See? You need the top model.” Pocock disagrees at the root. In his words: "you could find these bugs with cheaper models if you looked in the right places and gave it the right prompt/harness". In other words, what the “other models” lacked wasn’t intelligence — it was direction.

This connects directly to what you saw in Track 1 (Token Economy): a harness well, you let a “dumber” model do the same work. Here, the idea is the same, applied to bug discovery. Instead of paying a lot for a genius model to run once, you build a system that puts a simple model in the right place, repeatedly. The reason matters: an expensive model is a cost that grows with every call; a well-targeted system is an asset that pays off every day. The common mistake is confusing “it found a hard bug” with “I always need the harder model”—when what actually found the bug was where e how you looked.

EXPENSIVE model, without direction $$$ genius look in the dark → find 1 CHEAP model + harness $ simple + flashlight the right place → find 5

The secret isn’t the model’s brain—it’s where you point the flashlight (the harness).

Conceptual illustration: a flashlight illuminating hidden code, with a simple model finding flaws that an expensive model in the dark cannot see

⚠️ Common beginner mistake

Conclude, "I need the most expensive model to find serious bugs." Almost always, the failure was within reach of a cheaper model — all that was missing was point it to the right place, with the right prompt. Pay for the system, not the brain.

In one sentence: you find deep bugs with the right place + the right prompt—not with the most expensive model.

Going deeper (optional): why is the cheap model enough here?

Finding a bug doesn't require reasoning about the entire project at once — it requires looking at a small, well-defined piece with a clear question ("does this function validate the input?"). When you narrow the scope and give the right prompt, the problem becomes easy enough for a simple model. An expensive model shines on large, ambiguous tasks; sliced-up review is the opposite — so you save money without losing quality.

2

⏰ Security review cron

🧠 Imagine it this way: A guard making the rounds of a building. They don’t check all 30 floors every night — they’d do a poor job. They check one floor per night, calmly, and covered the whole building in a month, several times over. The security review cron is that watchman.

Pocock’s concrete recipe: run a cron journal of security review that checks, every day, a new part of the repository, using a relatively simple model. The keyword is rotating: you don’t try to scan the entire project at once (expensive, shallow, and blowing up the context window). You break it into pieces. Today, the authentication folder; tomorrow, payments; then uploads. In a few weeks, the entire repo has been audited—and the cycle starts again, catching what’s changed.

Why scheduled and not “when I remember”? Because the lesson Pocock draws from an incident is direct: "what did you learn? That there are security issues in your code — you should have something that runs and checks for more in the future." Security can’t depend on your memory on a busy day. The reason of the simple model: a sliced review is a small, well-scoped task—exactly where a cheap model is enough (topic 1). The common mistake is setting up a cron job that reviews the entire repo at once: it gets expensive, creates noise, and nobody reads the huge report. Small, daily, rotating, actionable.

daily cron sec · /auth have · /payments Wed · /uploads entire repo covered one slice per day, cycle restarts

Below, the copyable artifact: a line of crontab that runs a headless security review agent every day at 6 a.m., pointed at “the day’s slice.” Add it to your crontab (crontab -e) and adjust the repo path:

crontab · daily-security-review
# Edite com:  crontab -e
# Roda TODO DIA às 06:00 — checa uma fatia rotativa do repo (dia-do-mês % nº de pastas).
0 6 * * * cd /caminho/do/repo && DIRS=(auth payments uploads api db) && TARGET=${DIRS[$(( $(date +\%d) \% ${#DIRS[@]} ))]} && claude -p "Faça um SECURITY REVIEW so da pasta ./$TARGET. Liste cada falha (arquivo:linha), severidade (alta/media/baixa) e a correcao sugerida. Se nada grave, responda 'OK'. Saida em markdown." --model claude-haiku-4-5 --allowedTools "Read,Grep,Glob" >> ~/security-reviews/$(date +\%F)-$TARGET.md 2>&1

Quick recall: why does the security review cron check only one slice per day?

In one sentence: A scheduled guard who reviews one slice per day covers the entire building without costing a fortune.

3

📡 Telemetry → issue → fix

🧠 Imagine it this way: your home’s smoke alarm doesn’t just beep—it automatically calls the fire department, which arrives already knowing which room. Telemetry → issue → fix means setting up this pipeline: a production error triggers a chain that ends (almost) on its own with a fix.

Here, self-improvement stops being proactive (the cron) and becomes reactive, but automated. Pocock’s example: a bug in telemetry appears in a tool like the Sentry. Instead of reading, copying, and manually opening an issue, the system automatically creates an issue and the brand with a label (label) like agent:explore. That label triggers an agent—the exact queue pattern from module 4.4.

The agent then returns structured data: can it be fixed now, or does it need a human? If it’s straightforward, it implements the fix and opens the PR, and the flow moves on to review—eventually with a auto-merge (runs on its own) or a ping for you to decide. Pocock puts it all together in this phrase: "push the human-in-the-loop checkpoints as close as possible to the final result." In other words: the machine handles the whole process; you step in near the finish line, where your attention goes further. The common mistake is letting the pipeline auto-merge everything — with no human gate on changes that affect behavior, you trade manual work for silent risk.

telemetry(Sentry) issuelabel: explore agentstructured data PR / fix auto-merge pings you
Illustration: an automation conveyor carrying a telemetry error signal to a fix, with a human decision point near the end

🔬 Worked example: from a Sentry error to a fix

End-to-end scenario — one TypeError intermittent at checkout:

  1. Telemetry: Sentry groups 312 occurrences of Cannot read 'total' of undefined in checkout.ts.
  2. Automatic issue: A webhook creates the issue #812 with the stack trace and applies the label agent:explore.
  3. Agent explores: triggered by the label, returns structured data— cause: empty cart not handled; difficulty: trivial; can fix AFK: yes.
  4. Fix + PR: the agent adds the guard cart?.total ?? 0 + one test, open the PR #813 linking the issue.
  5. Checkpoint at the end: if it changes behavior, it does NOT auto-merge — it pings you, and you review it in 30s and approve.

You only showed up at step 5—the checkpoint was pushed closer to the final result.

In one sentence: a production error automatically triggers the issue→agent→fix chain; you only make decisions near the finish line.

4

🔒 The root cause of the bug

🧠 Imagine it this way: your bike gets stolen. You buy another bike. It gets stolen again. You buy another. The right question isn’t “which bike to buy” — it’s "why is it stealable? Where's the lock?". Fixing the bug is buying another bike. Finding the root cause is installing the lock.

This is the module’s turning point, and Pocock sums it up with an image that sticks: "if someone keeps stealing your bike, maybe get a lock." Fixing the bug that appeared is treating the symptom. The question that truly improves the system is about the root cause (root cause): why did this bug exist for so long without you noticing? What was missing from your process — a test, a review, a lint check — that let this slip through until it became an incident?

Notice the difference in level. "I found a security bug and fixed it" stops at the bike. "There are security flaws in my code, so I'm going to put together something that runs and checks for more in the future" installs the lock—it’s literally Topic 2’s cron job born from an incident. Each bug becomes more than just a fix; it becomes a question: Which system guardrail let this through? The answer becomes a test suite, automated review, or refactor. The common mistake is stopping happily at the fix: you close the issue, feel like you solved it — and the same kind of bug comes back through another door, because the lock was never bought.

"did someone steal the bike?" 😵 Treat the symptom buy a bike → it gets stolen → buy a bike… fix the bug → it comes back through another door the cycle never stops 🔒 Address the root cause "why did it exist for so long?" install the lock: test + cron + review the system gets stronger
Illustration: a bicycle secured with a sturdy, glowing lock, symbolizing addressing the root cause instead of the symptom
Going deeper (optional): the "5 whys"

A classic technique for getting to the root cause is to ask “why?” five times. Checkout bug → why? Empty cart not handled → why? Nobody tested this case → why? There was no empty-cart test → why? The suite doesn't cover edge cases → why? We don't have a cron/checklist that checks for these cases. By the fifth “why,” you’re no longer looking at the bug — you’re looking at the system that let it through. That’s where you buy the padlock.

In one sentence: don't buy another bike—ask why it keeps getting stolen and install the lock.

5

♻️ Self-improvement loops

🧠 Imagine it this way: A gym where every workout is recorded, and the record adjusts the next workout. You don’t get stronger “for free” — you get stronger because the training system learns from every session. Self-improvement loops are that idea applied to your code.

Putting it all together: a self-improving system is made up of several loops that reinforce each other. Pocock names three concrete ingredients: test suites (that catch regressions before they reach the user), human reviews (that provide insight into the system, not just the code) and refactor (which reduces the chance of the next bug appearing). Each incident feeds these loops: it becomes a test, a review checklist, or a cleaner part of the code.

The critical detail—and this is where it connects to module 4.4 (Loops × Queues)—is that “self-improving” no means "leave an infinite loop running on its own and disappear." It means setting up triggers (cron, telemetry, label) that put tasks in a queue, and feedback loops that make each pass better than the last. You’re still in charge: you prioritize, review the risky parts, and adjust the triggers. The common mistake is imagining a 100% autonomous system that fixes itself forever without supervision—that doesn’t exist; what does exist is a system that does the repetitive work and calls you in when it matters.

1 · detect 2 · fix 3 · prevent 4 · learn → review the system each cycle makes the next detection stronger

In one sentence: detect → fix → prevent → learn, and each cycle leaves the system stronger than the last.

6

🛠️ Review the system

🧠 Imagine it this way: a factory doesn’t inspect only the products that go out—it inspects the assembly line that produces them. If the line is crooked, every product comes out crooked. Reviewing only the code is checking the products; reviewing the system is straightening the line.

The thesis that closes the module is Pocock’s most important statement here: "review not just the code — also review the SYSTEM that produces the code." In the agentic era, the agent writes most of the code. So your review focus moves up a level: less "is this line correct?" and more "the process that generated this line trustworthy?” The prompts, skills, crons, review gates — this is the plumbing that, when tuned, keeps the next bug from ever being born.

This also answers a dangerous question that comes up in the next module (4.6): "who reviews the AI that says everything is okay?" Response: you, reviewing the system. That’s why it’s worth checking some PRs from time to time that the agent assured you are okay—not out of paranoid distrust, but to improve the system over time. Below is a practical summary in the form of a copyable checklist — run it whenever a bug slips through, to turn the fix into a system improvement:

revisar-o-sistema.txt
Bug escapou? Antes de fechar a issue, COMPRE O CADEADO:
[ ] CAUSA RAIZ — por que esse bug existiu tanto tempo sem ninguém notar?
[ ] TESTE — adicionei um teste que pegaria esse caso na próxima vez?
[ ] DETECÇÃO — meu cron de security/review olharia essa parte? Se não, ajusto o rodízio.
[ ] PROCESSO — que gate (review, lint, label) deixou isso passar? Reforço ele.
[ ] SISTEMA — revisei o código OU o sistema que produz o código? (era pra ser os dois)
Se você só corrigiu o bug, comprou outra bike. Instale o cadeado.

Quick recall: what does "review the system that produces the code" mean?

In one sentence: review not just the code, but the assembly line that produces it — that’s where the next bug dies before it’s born.

🧾 Module Summary

✓
It’s not the expensive model — you find deep bugs with the right place + the right prompt/harness; a cheap model is enough.
✓
Rotating daily cron — one slice of the repo per day covers everything, cheaply and with actionable results.
✓
Telemetry → issue → fix — a production error triggers the chain; the human checkpoint is near the end.
✓
Buy the lock — address the root cause, not the symptom; and review the SYSTEM that produces the code.

Next module:

4.6 — Checkpoints & smooth review: how to make review painless, push the checkpoint to the right, and record video+TTS of the agent at work.