🤔 Why TDD with an agent
An agent is a terrible verifier of its own work. It reads the code it just generated and sees that "looks right", and declares victory. Without a test run, “looks right” is just an opinion—and the agent’s opinion is biased by the code it wrote itself.
TDD solves this by turning things around: the test comes first, fails first, and only then is the code written to make the test pass. The agent can’t fool itself: either the test passes or it doesn’t.
🎯 The core principle
Agents need objective signs of progress. Without this, they go in circles, “fixing” things that already worked and breaking what was good.
- • Red test = work to do
- • Green test = you can stop
- • No test = no stopping criterion
💡 Practical tip
Before asking the agent to implement anything new, ask: "What test proves this works?" If you can't describe the test, the agent won't be able to verify the result either.
🔄 /tdd: the red-green-refactor loop
The command /tdd puts the agent in a disciplined three-step loop. Each step has a single goal — don't skip, combine, or optimize prematurely.
Red — write the failing test
Time: 2-5 min
Define the expected behavior as a test. Run the test and watch it fail. If it passes right away, either the test is wrong or the feature already exists. That “watch it fail” is proof that the test measures something real.
Green — minimum code to pass
Time: 5-15 min
Write the smallest amount of code possible to make the test pass. Hardcoding is allowed. Ugly is allowed. Slow is allowed. The only goal is to see green — you’ll improve it in the next step.
Refactor — clean up while keeping everything green
Time: 5-10 min
Now, remove duplication, improve names, and extract functions. Run the tests after each change. If they turn red, undo it. The test is the safety net that makes refactoring safe.
💻 Real /tdd dialogue
# você /tdd implementa função slugify(texto) — minúsculas, espaço vira "-", remove acentos, remove caracteres especiais # agente 🔴 Red: escrevi 4 testes em test_slugify.py - "Olá Mundo" → "ola-mundo" - " Espaço extra " → "espaco-extra" - "Já!!" → "ja" - "" → "" rodando pytest... 4 failed (módulo slugify não existe) ✓ esperado 🟢 Green: criando slugify.py com unicodedata + re.sub rodando pytest... 4 passed ✓ ♻️ Refactor: extraindo etapas em funções nomeadas _strip_accents(s), _to_kebab(s) rodando pytest após cada mudança... 4 passed ✓ pronto.
🍰 Vertical slices: one complete slice at a time
When a feature touches several layers (UI, API, DB), there are two ways to tackle it: horizontally (builds the entire DB, then the entire API, then the entire UI) or vertically (one feature end to end, then the next). With an agent, vertical slices always win.
Vertical slice = you can run it and see it working at the end of each slice. Horizontal = you only discover that everything is wrong when you put it all together at the end.
✓ Vertical (DO)
- ✓ Feature "create comment": migration + endpoint + form, all in one slice
- ✓ End-to-end test runs at the end of each slice
- ✓ Demoable: you can show it working after each slice
- ✓ Quick feedback: integration tested with every feature
✗ Horizontal (AVOID)
- ✗ "First I create all the tables, then all the endpoints"
- ✗ Nothing runs end to end through to the end of the project
- ✗ Big bang integration: everything breaks together at the end
- ✗ Agent loses context between layers and becomes inconsistent
📐 Vertical vs. Horizontal — diagram
VERTICAL (recomendado) HORIZONTAL (evitar) ┌──────┬──────┬──────┐ ┌──────────────────┐ │ UI │ UI │ UI │ │ UI │ ← slice 3 ├──────┼──────┼──────┤ ├──────────────────┤ │ API │ API │ API │ │ API │ ← slice 2 ├──────┼──────┼──────┤ ├──────────────────┤ │ DB │ DB │ DB │ │ DB │ ← slice 1 └──────┴──────┴──────┘ └──────────────────┘ feat1 feat2 feat3 tudo de uma vez cada slice roda só roda no fim de tudo end-to-end (e geralmente quebra)
💡 Practical rule
If the agent finishes a slice and you can't open the app and use that feature, the slice isn't done. Vertical = demoable at the end.
🔬 /diagnose: for difficult bugs
TDD works well for new work. But bugs in existing code—especially intermittent ones, ones that only happen in production, ones no one understands—call for a different approach. The /diagnose is a six-step loop designed for this.
Reproduce
Reproduce the bug reliably in the smallest possible scenario. Without a reproduction, any “fix” is a guess.
Minimize
Reduce the case to its skeleton: fewer files, less data, fewer steps. The smaller the case, the more obvious the cause.
Hypothesize
List possible hypotheses. Don't jump straight to the first one that seems promising — write 3-5 and rank them by probability.
Instrument
Add logs, print statements, breakpoints that prove or disprove each hypothesis. Don't fix anything yet — just measure.
Fix
Now, with the cause confirmed, write the fix. Keep it minimal and focused, without taking the opportunity to "fix this too."
Regression test
Turn the reproduction from step 1 into an automated test. If the bug comes back someday, CI will catch it before the user does.
📊 Why this loop instead of "just fixing things as you go"
A difficult bug is difficult because your intuition has already failed at least once (otherwise, you would have fixed it already). Jumping to “fix” before “reproduce + instrument” means betting your intuition will be right this time — it usually isn’t, and the agent reinforces the mistake by writing code that looks like it solves the problem but doesn’t.
🎯 Reproduce before fixing
The most costly debugging mistake is skipping the reproduction step. The agent reads the stack trace, sees a familiar name, and starts writing a fix. Sometimes it works—but when it doesn’t, no one notices, because no one set up the scenario to verify it.
✓ Reproduce first
- ✓ "Before touching the code, I'll write a minimal test that fails the same way as the bug"
- ✓ Confirm that the test fails for the right reason (same error, same stack)
- ✓ Only then address the cause — and measure the result with the same test
- ✓ When the test passes, you HAVE PROOF that you fixed it
✗ Anti-pattern: jumping to the fix
- ✗ "Oh, I know what that is, I'll fix it" — without reproducing it
- ✗ An agent that reads the error and immediately starts changing 5 files
- ✗ "I think that was it" — no way to prove it
- ✗ The bug comes back a week later because the “fix” doesn’t address the real case
⚠️ Warning sign
If the agent says "I'll adjust this here" and starts editing code without showing a reproduction, stop it. Ask first: "How do you reproduce the bug?" If it can't answer, it's guessing.
🛡️ Regression tests: every fixed bug becomes a test
A bug that came back is the most frustrating of all — because it has already been fixed once. The rule is simple: every fixed bug becomes a regression test. If you don’t have a test for the case, it will come back.
The best time to write this test is right after you’ve fixed the bug—the scenario is fresh, you know exactly what was failing, and the reproduction already exists (you used it in /diagnose).
💻 Example: regression test
# Bug reportado em produção: "comentário com emoji 🎉 quebrava a API" # Causa encontrada via /diagnose: encoding latin-1 num middleware antigo # Fix aplicado: trocar pra utf-8 # Teste de regressão (commit junto com o fix): def test_comentario_com_emoji_nao_quebra_api(): """Regression: bug #482 — emoji no body causava UnicodeDecodeError.""" payload = {"texto": "Show de bola 🎉🚀"} response = client.post("/api/comentarios", json=payload) assert response.status_code == 201 assert response.json()["texto"] == "Show de bola 🎉🚀" # Anote no docstring: número do bug, causa raiz. Quando o teste falhar # de novo em 6 meses, quem ler vai entender o porquê de existir.
💡 If there's no test, the bug comes back
It's not pessimism, it's statistics. Future refactors, dependency changes, someone cleaning up “old code nobody uses” — any of these can bring the bug back. The test is the automatic reminder: "this case here needs to keep working".
📊 Docstring convention
- • Prefix:
Regression:make it explicit that the test exists because of a bug - • Reference: issue/ticket/PR number to track context
- • Root cause: one line — so whoever edits the test later understands
- • DO NOT delete: a regression test is an asset of the project
🔀 Combining /tdd + /diagnose
The two loops don't compete—they work together. /tdd is for new work when you know what you want. /diagnose is for understanding why something existing isn't doing what it should. Complex bugs almost always involve both.
🎬 Workflow: intermittent bug in production
Imagine: a bug report says "sometimes the payment is duplicated." You can’t reproduce it locally. What should you do?
- 1. /diagnose — reproduce: can’t. Go back a step: instrument production (transaction ID, timestamp, request ID logs).
- 2. /diagnose — hypothesise: client retry? race condition in the worker? duplicate gateway webhook?
- 3. /diagnose — instrument + minimize: logs show that two workers pick up the same job when the client retries in <1s. Found it: a race condition.
- 4. /tdd takes over from here — red: write a test that dispatches two workers simultaneously for the same payload. It fails (duplicates).
- 5. /tdd — green: implements an idempotency key in the payments table. Test passes.
- 6. /tdd — refactor + regression: cleans up the code and keeps the test as a permanent regression test. The bug won’t come back.
🧭 When to switch
- • Do you know what you want to build? → Go straight to /tdd
- • Have a bug and know how to reproduce it? → Use /diagnose to find the root cause, then /tdd for the fix
- • Have a bug and DON’T know how to reproduce it? → /diagnose starts with remote instrumentation before any fix
- • Refactor with existing tests? → It stays in the green-refactor stage of /tdd
- • Refactor WITHOUT tests? → Start with /tdd: cover the current behavior with tests; refactor only afterward
🏋️ Hands-on exercise
Pick an open bug in your project — any one, preferably one you've been avoiding because "it's weird." Run /diagnose and document each step. By the end, you'll have the bug fixed AND review material to understand how agents handle bugs.
📋 Exercise outline
- 1. Choose the bug. Paste into the chat: problem description, stack trace if available, minimal context.
- 2. Run
/diagnoseand follow the 6 steps WITHOUT skipping any. - 3. At each step, make a note in a
diagnose-log.md: what the agent did, what it discovered, what still wasn’t clear. - 4. When you get to the fix, BEFORE applying it, write the regression test (step 6 of /diagnose).
- 5. Run the test — it should fail. Apply the fix. Run it again — it should pass.
- 6. Commit the fix and test together, with a message that references the root cause.
💡 Success criteria
It's not “the bug was fixed.” It's “I can prove the bug was fixed, and I have a test that will sound the alarm if it comes back.” If you can't check both boxes, redo the exercise.
📚 Module Summary
Next Module:
1.5 — Next steps in the Fundamentals track