Learning path map
🕹️ Human-in-the-loop × AFK
Get out of the loop
📦 Parallelize & sandboxes
Isolated agents
⚙️ GitHub Actions + agents
AFK in the cloud
🧮 Loops × Queues
Queue, not loop
♻️ Self-improving systems
Buy the lock
🎬 Checkpoints & smooth review
Pain-free review
Detailed content
🕹️ Human-in-the-loop × AFK
Get out of the loop: stop approving every step and let the agent run on its own (AFK).
The default mode: you approve every step and every agent action.
It’s safe, but it becomes a bottleneck — you’re the speed limit.
Manual approval; security × throughput.
Let the agent run on its own while you’re away from the keyboard.
It’s the productivity leap: the work happens without you.
AFK = autonomous agent; no babysitter.
The moment you trust the harness enough to step out of the loop.
Unlocks parallelism: several agents working at the same time.
Trust in the harness → autonomy.
Risky or ambiguous tasks still need you in the loop.
Blindly going AFK on the wrong task is like handing a junior the car keys.
Risk × reversibility determines the mode.
Configure permissions so the agent can act without asking for approval at every step.
Every manual confirmation is friction that kills AFK.
Allowlist; auto-approve with limits.
AFK + parallelism = several “yous” working on different fronts.
It’s the method’s real scaling lever.
Fleet of agents; you as the manager.
📦 Parallelize & sandboxes
Isolated agents: run several in parallel without one breaking another's environment.
Run multiple tasks with different agents at the same time.
Multiplies throughput without multiplying your time.
Parallelism = throughput; you become the orchestrator.
An AFK agent with full access to your machine can cause damage.
AFK without isolation is a recipe for disaster.
Blast radius; isolate before releasing.
The sandbox tool Matt uses to run isolated agents.
It’s the practical shortcut to safe AFK.
Ready-to-use sandbox; isolation for each task.
Containers as a sandbox for each agent.
Industry-standard isolation, easy to discard and recreate.
Ephemeral container; reproducible environment.
Managed cloud sandboxes for running agents without using your machine.
Takes the agent off your laptop and completely frees you up.
Sandbox as a service; without bogging down your local machine.
Coordinate a fleet of agents in separate sandboxes.
It’s where parallelism + isolation become real scale.
Isolated fleet; each one in its own box.
⚙️ GitHub Actions + agents
AFK in the cloud: agents that run in CI, open PRs, and never slow down your machine.
Run agents inside GitHub Actions as a CI step.
CI is a free AFK sandbox you already have.
Agent as a job; ephemeral environment.
An Action that automatically reviews every PR with an agent.
Consistent review on every PR, without you remembering to ask.
Review on push; feedback on the PR.
Applying a label to an issue/PR triggers the agent to act.
Turns into a "send the agent" button inside GitHub.
Trigger by label; on demand.
The agent delivers the work as a PR ready for you to review.
A PR is the natural checkpoint between AFK and human review.
Reviewable output; nothing goes straight to main.
All the work runs on GitHub runners, not your laptop.
You’re free while the agent works in the cloud.
Remote compute; laptop unlocked.
Build your own agent Action from scratch (covered in Track 5).
Your Action adapts exactly to your workflow.
Minimal YAML; checkout → agent → PR.
🧮 Loops × Queues
Queue, not loop: why a task queue beats an agent’s infinite loop.
Geoff Huntley’s “Ralph loop”: run the agent in a loop until it solves the problem.
It’s the starting point — and where many people get stuck.
Brute-force loop; repeats until "done".
A blind loop repeats without prioritizing or scoping — wasting tokens.
Repeating isn’t the same as organizing the work.
Loop without triage = costly and erratic.
Instead of a loop, a queue of scoped tasks for the agent to consume.
A queue provides order, priority, and a clean stop.
Task queue; consume in order.
Break down and prioritize tasks with a clear scope before queuing them.
A well-scoped task is one the agent can finish on its own.
Triage; clear scope for each item.
You are the king who gives orders; the agents are the subjects who carry them out.
Set the mindset: you’re in command, not doing the execution.
The king delegates; subjects work through the queue.
Multiple agents (nodes) pulling from the same queue in parallel.
Combine a queue with parallelism for maximum throughput.
Workers pulling from the queue; horizontal scaling.
♻️ Self-improving systems
Buy the lock: systems that detect problems and fix themselves.
A well-built system doesn’t require the most expensive model to sustain itself.
"Buy the lock": invest in the system, not the expensive part.
Affordable, robust systems > premium model.
A cron that periodically runs a security review agent.
Security becomes an automatic routine, not a one-off effort.
Daily cron; recurring scan.
Telemetry detects the problem, opens an issue, and triggers the fix.
Closes the loop from observation to fix without you in the middle.
Detect → open an issue → fix.
The agent looks for the root cause, not just the symptom of the bug.
Root cause analysis prevents the same problem from coming back.
Root cause > band-aid.
The system learns from every failure and adjusts itself.
Self-improvement compounds: the system gets better on its own over time.
Feedback loop; compounding improvement.
Periodically review the self-improving system itself.
Unsupervised self-improvement can drift in the wrong direction.
Audit the self-adjustment; keep a human at the next level up.
🎬 Checkpoints & smooth review
Pain-free review: push the checkpoint to the right and review the agent without friction.
Review is the point where you check and correct the agent’s direction.
It’s your safety net in the AFK world.
Checkpoint = quality control.
Move the review point to the end, giving the agent more autonomy.
The farther to the right the checkpoint, the more AFK you can be.
Move the checkpoint as confidence increases.
Identify the checkpoints where the human no longer adds value.
Removing the right human speeds things up; removing the wrong one breaks things.
Remove the redundant checkpoint; keep the critical one.
Use one agent to review another’s work before you do.
Filters out most errors before the human checkpoint.
Reviewer agent; double-check.
The agent records a video walkthrough of what it did, with TTS narration.
You review by watching instead of reading diff by diff.
Narrated walkthrough; review by video.
Use AI to summarize, highlight, and speed up your human review.
Fast review keeps AFK flowing without becoming a bottleneck.
AI summarizes the diff; painless review.