Track map
🧭 Why skills exist
The 4 problems that kill projects with agents
🔥 Grilling: alignment
Stop guessing what you want
📖 Ubiquitous Language
1 word instead of 20
✅ Code That Works
Red, green, refactor — with the agent alongside
🏗️ Healthy Architecture
Get out of the mud ball without refactoring everything
🔄 Complete workflow
From idea to merged PR
Detailed content
🧭 Why skills exist
The 4 problems that kill projects with agents—and how skills tackle each one.
A structured package (SKILL.md + scripts + refs) that the agent loads on demand to perform a specific task.
Without skills, you repeat context every session. With skills, knowledge is versioned, testable, and reusable.
SKILL.md, frontmatter, triggers, references, scripts, single scope.
Gap between what you asked for and what the agent understood — creates silent rework.
It’s the #1 cause of “the agent doesn’t work.” Recognizing it early prevents hours of writing the wrong code.
Ambiguity, implicit assumptions, lack of examples, vague scope.
Too many words, repetition, and noise in the prompt degrade the agent’s accuracy.
Compact ubiquitous language captures intent. 1 well-defined term is worth 20 lines of explanation.
Token budget, noise vs. signal, shared terminology.
Code generated without live tests: it appears to work but fails in edge cases.
TDD with an agent turns “I hope it works” into “prove that it works.”
Red/green/refactor, vertical slice, regression test.
Codebase with unclear boundaries—any change breaks things far away.
Agents amplify poor coupling. Architecture skills combat entropy.
Coupling, boundaries, ADRs, progressive improvement.
Matt Pocock’s skill system — grill, ubiqua, tdd, architecture — chained in the workflow.
Understanding the map before the details speeds adoption and prevents ad hoc use.
Composition, usage order, slash commands, agentic OS.
🔥 Grilling: alignment with the agent
Stop guessing what you want — let the agent ask first.
A technique where the agent asks pointed questions before coding—uncovering hidden requirements.
Prevents weeks of rework. 30min of grilling saves 5h of writing the wrong code.
Adversarial questioning, assumptions, scope, acceptance criteria.
/grill-me attacks the raw idea; /grill-with-docs compares the plan against the project’s CONTEXT.md and ADRs.
Using the wrong mode wastes the grilling. Each has its own trigger.
Raw idea, mature plan, reference document, ADR.
Triggers: ambiguous feature, architectural decision, team conflict, pre-PRD.
Knowing WHEN grilling adds value helps you avoid using it where it isn’t needed.
Trigger heuristics, cost/benefit, timing.
Annotated transcript of a /grill-me session turning a vague idea into an actionable PRD.
See in practice what a good question looks like versus an evasive answer.
Drill-down, counterexample, making assumptions explicit.
Shallow grilling, rhetorical questions, an agent that always agrees — signs of a wasted session.
Detecting early prevents the illusion of alignment.
Sycophancy, leading questions, anchoring.
A completed grilling session becomes a PRD via /to-prd — with no loss of context.
Connects alignment with execution. The end of “I did the grilling and lost everything.”
PRD, acceptance criteria, vertical slices, handoff.
📖 Ubiquitous Language: CONTEXT.md + ADRs
1 word instead of 20 — terminology shared between you and the agent.
Shared vocabulary across the domain, code, and agent — same term, same meaning.
Reduces mental translation. "Materialization cascade" is worth a paragraph of explanation.
DDD, bounded context, living glossary, domain concepts.
Root document with domain, concepts, boundaries, and terms — read by the agent in every context.
Without CONTEXT.md, the agent relearns your domain every session.
Required sections, glossary, boundaries, examples.
A short, immutable record of each architectural decision—context, options, choice, and consequences.
Prevents redoing an old decision without knowing it. The agent respects documented constraints.
Decision, context, alternatives, status, consequence.
A real case from Matt—the term was coined once in CONTEXT.md and reused in dozens of prompts.
See the concrete gain from compression through vocabulary.
Semantic compression, term reuse, dense intent.
Routine of updating CONTEXT.md and ADRs in every relevant PR — don’t let them turn into archaeology.
An outdated document is worse than no document.
Docs-as-code, drift, PR gates.
Metrics comparing before and after adopting ubiquitous language — tokens, session time, bugs.
Justify adoption to the team with numbers, not opinions.
Token spend, lead time, rework, dev NPS.
✅ Code That Works: TDD + Diagnose
Red, green, refactor — with the agent alongside. And when something breaks, /diagnose.
TDD used as a contract between you and the agent—test before code.
Without a test first, the agent "hallucinates" behavior. With a test, behavior is verifiable.
Executable contract, feedback loop, objective criteria.
Slash command that orchestrates the TDD cycle with the agent — generates a test, runs it, implements, refactors.
Standardizes quality. Every feature starts with a passing test.
Cycle, gates, refactoring automation.
Deliver minimum end-to-end value at a time — UI, API, database — instead of horizontal layers.
Enables real TDD. Each slice can be tested on its own.
Slice, walking skeleton, MVP, testable increment.
A skill that guides systematic bug investigation — reproduce, minimize, hypothesize, test.
An intermittent bug without a method becomes an infinite loop. /diagnose enforces discipline.
Hypothesis, observation, isolation, falsification.
Required sequence: reproduce the bug, reduce the case, formulate a hypothesis before touching the code.
Skipping a step leads to a correction that masks the bug.
MCVE, bisection, falsification.
Test written from the minimized repro—ensures the bug doesn’t come back.
Without a regression test, you pay for the same bug 3x a year.
Capture, fixture, mutation testing.
🏗️ Healthy Architecture
Get out of the mud ball without refactoring everything — agent-guided progressive improvement.
Signs: changing 1 line breaks 3 distant features, nobody knows where something lives, fear of making changes.
Diagnosis comes before treatment. If you don’t name the symptoms, you’ll treat the wrong symptom.
Coupling, cohesion, boundaries, shotgun surgery.
A skill that scans the code, identifies hotspots, and suggests prioritized improvements.
Replaces "big bang refactoring" with surgical interventions.
Hotspot, prioritization, safe change.
A skill that abstracts from the current file to the whole system — breaks tunnel vision.
A good local decision can be bad globally. Zooming out forces a systems perspective.
Macros vs. micro, systems thinking, cascading impact.
Areas of the code where investing in modeling yields disproportionate returns.
Time is finite. Going deeper in the wrong place = polishing the sidelines.
Core domain, modeling ROI, leverage.
Use the Track 1.3 documents to guide where and how to improve — the agent follows documented constraints.
Refactoring without a north star becomes personal taste. A documented, auditable north star.
Constraints, compliance, traceability.
Codebase walkthrough from “messy” to “healthy” through small steps guided by skills.
Inspires confidence: you don't need to stop everything to fix it.
Strangler fig, boy scout rule, short branches.
🔄 Complete workflow
From idea to merged PR — all the skills in the track, in order of use.
Command that installs all of Matt's skills in your Claude Code and prepares the environment.
Everything starts here. Without setup, the rest of the track won’t work.
Bootstrap, dependencies, verification.
Workflow: idea → /grill-me → /to-prd. Output is a PRD ready for slicing.
Connects modules 1.2 and 1.6 — without this bridge, grilling becomes a lost conversation.
PRD, acceptance criteria, documentation handoff.
/to-prd structures the document; /to-issues breaks it down into testable issues for execution.
A well-scoped issue = an unambiguous sprint.
Slicing, INVEST, acceptance criteria.
A skill that decides which role (PM, dev, reviewer) the agent takes on based on the issue’s state.
Prevents an agent from acting as a developer when it should be reviewing, and vice versa.
State machine, persona, transition gates.
Each triaged issue becomes a /tdd cycle—test, code, refactor—until green.
It’s where theory becomes a merge. Without TDD here, quality falls apart.
Definition of done, CI gates, complete slice.
Final step: automated + human review, update CONTEXT.md/ADRs, merge.
Without a docs + review closeout, the cycle "leaks" and the architecture starts degrading again.
Review gates, doc update, merge policy.