📖 Living glossary (read first — come back whenever you need to)
This module is full of automation terms. We define the new in plain language here—the fundamentals (model, agent, harness, skill) were already covered in Track 1:
.github/workflows/.tsc in TypeScript) that catches type errors before running. It's the cheap "passed quality control" check.0 9 * * *).🦴 Action skeleton
🧠 Imagine it this way: A factory with a conveyor belt. Every time a new part arrives (a PR), the belt starts automatically: it picks up the part, runs it through quality control, calls the inspector (the agent), and sticks on a label with the assessment. You don’t start the belt by hand — it reacts to the trigger. A GitHub Action is that conveyor belt, but for code.
Pocock describes a "agent review action": one GitHub Action that “runs on a PR, checks out the code, runs a type check, runs the review agent (a local prompt), and responds ‘looks good.’” The secret to every Action is its skeleton: it has a trigger (o on: — what connects to the pipeline), one or more jobs (what runs), and inside each job, a sequence of steps (the steps). The skeleton is always the same—only what goes inside changes.
Why does this unlock so much? Because “run agents on GitHub Actions is unreasonably effective"—you can parallelize as much as you want without worrying about local machine resources. The pipeline runs in GitHub's cloud: your laptop can be turned off. The local prompt (a review instruction file that lives in the repo itself) is the inspector's brain—versioned alongside the code, just like a skill. Common mistake: thinking you need your own server running 24/7. You don't: the Action only wakes up when triggered, does the work, and shuts down — you pay only for the minutes used.
Trigger → job → steps. This skeleton is universal; the content of the steps is what you swap out.
Below is the Real review action YAML — the first of this module’s two copyable pieces. Save it as .github/workflows/review.yml. Notice the structure: on: (trigger), jobs:, and the steps: in the right order. Copy and adjust only the command for your type check:
name: agent-review
on:
pull_request:
types: [opened, synchronize, reopened]
jobs:
review:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- name: Checkout PR
uses: actions/checkout@v4
- name: Type check
run: npx tsc --noEmit # guard rail barato e determinístico
- name: Review agent
run: npx claude -p "$(cat .github/review-prompt.md)" --output review.md
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- name: Comment on PR
run: gh pr comment "$PR" --body-file review.md
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PR: ${{ github.event.pull_request.number }}
In one sentence: An Action is a conveyor belt in the cloud — a trigger starts it, steps run, and the rest of your machine doesn’t even need to be on.
Going deeper (optional): why does the review prompt live in the repo?
If the review’s "local prompt" lives inside the repository (e.g., .github/review-prompt.md), it’s versioned along with the code: improve it in a PR, it’s recorded in the history, and everyone on the team sees the same instruction. It’s the same principle as human DRY (module 3-6): "I’ve done this review 100 times → it becomes a procedure, I share it with the team, everyone reviews the same way." The Action is just the pipeline; the prompt is the skill it runs.
🔧 checkout → type check → review
🧠 Imagine it this way: Before the (expensive) specialist examines the patient, the nurse has already taken their blood pressure and temperature (quick and cheap). If the temperature is already through the roof, there’s no need for the specialist. The type check is the cheap triage; the review agent is the specialist—and you only call the specialist once the basics are covered.
These three steps are the backbone of the review action, in the order Pocock cites: "do checkout, type check, run the review agent.” The order matters for efficiency: first you download the code, then run the check deterministic and inexpensive (the type check catches obvious errors in seconds), and only then spend tokens on the agent. If the type check fails, it’s often not even worth waking the agent—you already know the PR is broken.
This is pure "token economy" (module 1-5): the agent is the most expensive step, so surround it with guardrails cheap. Another reason for the order: the agent reviews better when the code already compiles. Common mistake: throw everything at the agent at once (“review and fix the types and run the tests”). No — separate the steps. Deterministic tasks (type check, lint, tests) do what’s deterministic; the agent does what requires judgment. Each in its own lane, with the costly one last.
Quick recall: why does the type check comes before the review agent?
In one sentence: start with the cheap, deterministic checks (checkout, type check); use the expensive agent only last—and only if the basics pass.
💬 Comment on the PR
🧠 Imagine it this way: the factory inspector doesn’t keep the assessment in their head—they stick a visible note on the part: “ok” or “check point X.” Whoever comes later reads the note. The PR comment is that note: the agent’s result stays where the team looks, not lost in a log.
The final step of the review action is comment on your own PR — Pocock says the agent "replies 'looks good'" right there. Why comment on the PR instead of in an email or chat? Because the PR is where the decision happens: the comment stays attached to the diff, next to the discussion, and GitHub already notifies whoever is following it. It's the natural output of AFK in the cloud—the agent works alone and makes the assessment visible, how to make the bot github-actions that posts a “Triage” with a TL;DR and difficulty level (you saw this in module 4-4, Loops × Queues).
Here's the second value of the review (module 4-6): the comment isn't just for gate ("this is dangerous, hold on") — it gives you observability. You read what the agent found and see your own plumbing. Common mistake: letting the agent approve or reject in silence. If the assessment doesn’t appear in the PR, you lose the chance to improve the harness over time—and fall into the trap of “who reviews the AI that says everything is fine?” The visible comment is what keeps you in command.
🔬 Worked example: a PR goes through the entire pipeline
A dev opens PR #128 ("adds login endpoint"). The review action triggers:
- checkout — the Action downloads the PR branch in the cloud (your machine doesn't even need to be on).
- type check —
tscruns in 12s. Passes. (If it failed, the agent wouldn’t even be called.) - review agent — the agent reads the diff with the local review prompt. It finds an unvalidated input.
- comment on the PR — posts: “✅ type check ok; ⚠️ validate input in /login”. The dev sees it right away.
Result: review in ~2 min, without you opening your laptop, and the assessment is recorded in the right place.
In one sentence: the agent’s assessment goes into the PR—where the team looks—to serve as a gate and give you observability into your system.
🗓️ Rotating daily cron
🧠 Imagine it this way: "if someone keeps stealing your bike, maybe get a lock." Instead of just fixing a security problem when it blows up, you set up a guard who makes a round every day — inspecting a different street in the neighborhood each day. In a month, the whole neighborhood has been reviewed without anyone getting tired.
The second automation is the module’s gem. Pocock says the Fable model found deep bugs and security flaws others missed—but emphasizes: isn't model magic. "You'd find these bugs with cheaper models if you looked in the right places and gave them the right prompt/harness." The recipe: a daily security review cron job that “each day it checks a new part of the repo with a relatively simple model.” The cron is the scheduler; “rotating” is the trick—you don’t review the entire repo every day (expensive and slow); you review one slice per day, cycling through the repository.
Why does rotation work so well? Because it spreads the cost over time and eventually covers everything. A repo with 30 folders? One folder a day → full coverage in 30 days, using a cheap model. And the real benefit isn’t just finding the bug: it’s the lesson. “What did you learn? That there are security issues in your code—you should have something that runs and checks for more in the future.” That’s the heart of self-improving systems (module 4-5): you don’t just review the code; you review the system that produces the code. Common mistake: running the security review on the entire repo at once — it gets expensive, takes a while, and you end up turning it off. Rotation is what makes the habit sustainable.
The line of crontab is the technical heart of this—and the module’s second copyable piece. It has five fields—minute, hour, day of the month, month, day of the week—and the * means "all." The line below is the actual one you paste into the schedule: a GitHub Action to run every day at 9h UTC. Copy:
# campos do cron: minuto hora dia-do-mês mês dia-da-semana
# "0 9 * * *" = todo dia às 09:00 UTC
on:
schedule:
- cron: "0 9 * * *" # security review diário, fatia rotativa do repo
In one sentence: a daily cron job reviews a rotating slice of the repo with a cheap model — buy the lock before the bike gets stolen again.
Going deeper (optional): how do you decide WHICH slice to check today?
Two simple approaches. (1) By calendar: use the day of the month to index a list of folders (day 1 → folder 0, day 2 → folder 1…). (2) By priority: keep a queue of the "hottest" folders (most recent commits, or ones that touch auth/payment) and rotate through them first. The important thing is that the slice be small and different every day — this way, the cheaper model can handle it and the cost stays flat. Pocock also suggests finding the root cause: why did this bug live so long without you noticing? The answer becomes the next system improvement.
🧱 Structure explore data
🧠 Imagine it this way: the detective returns from the investigation. If they tell you the story in a rambling paragraph, you have to read it all to decide. But if they hand you a completed form — "Serious? Yes. Can it fix it on its own? Yes. Confidence? 90%" — you decide in 2 seconds, and even a robot can read it. The secret to AFK is having the agent deliver the form, not the story.
Pocock describes the complete workflow of explore: “telemetry bug (Sentry) → create issue → label explore → the agent returns structured data (can it be fixed now, or does it need a human?) → implement → review → tag auto-merge or pings a human." The key point is the structured output: instead of the agent returning a wall of text, it returns fixed fields—severity, whether it can self-correct, and confidence. Why? Because a “self-correctable: yes/no” field is something you another program can read and use it to decide the next step on its own. Loose text, no.
It’s what connects the whole queue (module 4-4): the issue gets labels (agent:explore → agent:in-progress), the agent does the AFK triage and returns structured data, and that data determines the routing—fix it now, or push the human-in-the-loop checkpoint further along. "Push human-in-the-loop checkpoints further toward the final output." Common mistake: letting the agent respond in prose and trying to use a regex to guess whether it approved. Fragile. Ask for JSON with defined fields so the automation can chain steps without ambiguity.
{
"issue": 795,
"severity": "high",
"auto_fixable": true,
"confidence": 0.9,
"root_cause": "input não validado em /login",
"suggested_label": "agent:implement",
"needs_human": false
}
In one sentence: make explore return fields (JSON), not prose — that way, automation can read, decide, and chain things together on its own.
⚠️ Auto-merge with caution
🧠 Imagine it this way: An automatic gate that opens by itself is wonderful — until the day it opens for the wrong car. The solution isn’t to lock the gate forever; it’s to set clear conditions (only opens for a registered plate, during the day, with the sensor working). Auto-merge is that gate: great with gates, dangerous left unchecked.
At the end of the workflow, the agent “tag auto-merge or pings a human". Auto-merge is the pinnacle of AFK: the PR enters the main codebase without anyone clicking. But it’s also the most dangerous place—which is why the topic is called “with caution.” Pocock emphasizes the right goal: “push human-in-the-loop checkpoints further toward the final output" — push the approval point further along, removing checkpoints wherever possible, carefully. It’s not about removing humans from everything; it’s about removing them from where they add no value.
Where can you safely auto-merge? When “the PR is just an internal refactor; it doesn’t change behavior—the AI says ‘no review needed.’” But then comes Pocock’s golden question: "who reviews the AI that says it's fine?" (who reviews the AI that says everything is fine?). The answer: you need sample some PRs the agent approved on its own, to calibrate trust over time. “We're reviewing not just the code, but the SYSTEM that produces the code.” Common mistake: enable auto-merge for everything on day 1. Start narrow (only the fields auto_fixable: true + confidence ≥ 0.9 + green type check), audit it for a few days, and only then expand. Caution here is what keeps the system reliable.
Quick retrieval: what is Pocock's right stance on auto-merge?
In one sentence: auto-merge is an automatic gate—great with clear gates, and you still sample-check "who reviews the AI that says everything is fine?".
🧾 Module Summary
Next module:
5.6 — Smooth Review + Where to Find It: video walkthrough with TTS, optimize the human review, and where to get ready-made skills (mattpocock/skills, aihero.dev/skills).