PTENES
MODULE 1.4

✅ Code That Works: TDD + Diagnose

Red, green, refactor — with the agent alongside. How to use TDD and the /diagnose loop to make sure the code actually works instead of just "looking right" to the agent.

9
Topics
45
Minutes
Intermediate
Level
Practice
Type
1

🤔 Why TDD with an agent

An agent is a terrible verifier of its own work. It reads the code it just generated and sees that "looks right", and declares victory. Without a test run, “looks right” is just an opinion—and the agent’s opinion is biased by the code it wrote itself.

TDD solves this by turning things around: the test comes first, fails first, and only then is the code written to make the test pass. The agent can’t fool itself: either the test passes or it doesn’t.

🎯 The core principle

Agents need objective signs of progress. Without this, they go in circles, “fixing” things that already worked and breaking what was good.

  • • Red test = work to do
  • • Green test = you can stop
  • • No test = no stopping criterion

💡 Practical tip

Before asking the agent to implement anything new, ask: "What test proves this works?" If you can't describe the test, the agent won't be able to verify the result either.

Evidence
test running > "seems right"
Criterion
green = ready
Anti-loop
tests prevent infinite loops
Contract
the test is the specification
2

🔄 /tdd: the red-green-refactor loop

The command /tdd puts the agent in a disciplined three-step loop. Each step has a single goal — don't skip, combine, or optimize prematurely.

🔴

Red — write the failing test

Time: 2-5 min

Define the expected behavior as a test. Run the test and watch it fail. If it passes right away, either the test is wrong or the feature already exists. That “watch it fail” is proof that the test measures something real.

🟢

Green — minimum code to pass

Time: 5-15 min

Write the smallest amount of code possible to make the test pass. Hardcoding is allowed. Ugly is allowed. Slow is allowed. The only goal is to see green — you’ll improve it in the next step.

♻️

Refactor — clean up while keeping everything green

Time: 5-10 min

Now, remove duplication, improve names, and extract functions. Run the tests after each change. If they turn red, undo it. The test is the safety net that makes refactoring safe.

💻 Real /tdd dialogue

# você
/tdd implementa função slugify(texto) — minúsculas, espaço vira "-",
remove acentos, remove caracteres especiais

# agente
🔴 Red: escrevi 4 testes em test_slugify.py
  - "Olá Mundo" → "ola-mundo"
  - "  Espaço extra  " → "espaco-extra"
  - "Já!!" → "ja"
  - "" → ""
rodando pytest... 4 failed (módulo slugify não existe) ✓ esperado

🟢 Green: criando slugify.py com unicodedata + re.sub
rodando pytest... 4 passed ✓

♻️ Refactor: extraindo etapas em funções nomeadas
  _strip_accents(s), _to_kebab(s)
rodando pytest após cada mudança... 4 passed ✓
pronto.
Discipline
one step at a time
See it fail
the test has to fail first
Minimum
don’t optimize on green
Safety net
test-protected refactor
3

🍰 Vertical slices: one complete slice at a time

When a feature touches several layers (UI, API, DB), there are two ways to tackle it: horizontally (builds the entire DB, then the entire API, then the entire UI) or vertically (one feature end to end, then the next). With an agent, vertical slices always win.

Vertical slice = you can run it and see it working at the end of each slice. Horizontal = you only discover that everything is wrong when you put it all together at the end.

✓ Vertical (DO)

  • ✓ Feature "create comment": migration + endpoint + form, all in one slice
  • ✓ End-to-end test runs at the end of each slice
  • ✓ Demoable: you can show it working after each slice
  • ✓ Quick feedback: integration tested with every feature

✗ Horizontal (AVOID)

  • ✗ "First I create all the tables, then all the endpoints"
  • ✗ Nothing runs end to end through to the end of the project
  • ✗ Big bang integration: everything breaks together at the end
  • ✗ Agent loses context between layers and becomes inconsistent

📐 Vertical vs. Horizontal — diagram

VERTICAL (recomendado)              HORIZONTAL (evitar)
┌──────┬──────┬──────┐               ┌──────────────────┐
│  UI  │  UI  │  UI  │               │       UI         │  ← slice 3
├──────┼──────┼──────┤               ├──────────────────┤
│ API  │ API  │ API  │               │       API        │  ← slice 2
├──────┼──────┼──────┤               ├──────────────────┤
│  DB  │  DB  │  DB  │               │       DB         │  ← slice 1
└──────┴──────┴──────┘               └──────────────────┘
 feat1  feat2  feat3                   tudo de uma vez

cada slice roda           só roda no fim de tudo
end-to-end                (e geralmente quebra)

💡 Practical rule

If the agent finishes a slice and you can't open the app and use that feature, the slice isn't done. Vertical = demoable at the end.

Demoable
each demonstrable slice
End-to-end
test integration now
Feedback
finds out early what breaks
Context
agent focuses on one feature
4

🔬 /diagnose: for difficult bugs

TDD works well for new work. But bugs in existing code—especially intermittent ones, ones that only happen in production, ones no one understands—call for a different approach. The /diagnose is a six-step loop designed for this.

1

Reproduce

Reproduce the bug reliably in the smallest possible scenario. Without a reproduction, any “fix” is a guess.

2

Minimize

Reduce the case to its skeleton: fewer files, less data, fewer steps. The smaller the case, the more obvious the cause.

3

Hypothesize

List possible hypotheses. Don't jump straight to the first one that seems promising — write 3-5 and rank them by probability.

4

Instrument

Add logs, print statements, breakpoints that prove or disprove each hypothesis. Don't fix anything yet — just measure.

5

Fix

Now, with the cause confirmed, write the fix. Keep it minimal and focused, without taking the opportunity to "fix this too."

6

Regression test

Turn the reproduction from step 1 into an automated test. If the bug comes back someday, CI will catch it before the user does.

📊 Why this loop instead of "just fixing things as you go"

A difficult bug is difficult because your intuition has already failed at least once (otherwise, you would have fixed it already). Jumping to “fix” before “reproduce + instrument” means betting your intuition will be right this time — it usually isn’t, and the agent reinforces the mistake by writing code that looks like it solves the problem but doesn’t.

Repro
seeing it happen = the beginning
Minimizes
cut everything that doesn’t matter
Measure
prove it before fixing it
Lock
tests prevent regressions
5

🎯 Reproduce before fixing

The most costly debugging mistake is skipping the reproduction step. The agent reads the stack trace, sees a familiar name, and starts writing a fix. Sometimes it works—but when it doesn’t, no one notices, because no one set up the scenario to verify it.

✓ Reproduce first

  • ✓ "Before touching the code, I'll write a minimal test that fails the same way as the bug"
  • ✓ Confirm that the test fails for the right reason (same error, same stack)
  • ✓ Only then address the cause — and measure the result with the same test
  • ✓ When the test passes, you HAVE PROOF that you fixed it

✗ Anti-pattern: jumping to the fix

  • ✗ "Oh, I know what that is, I'll fix it" — without reproducing it
  • ✗ An agent that reads the error and immediately starts changing 5 files
  • ✗ "I think that was it" — no way to prove it
  • ✗ The bug comes back a week later because the “fix” doesn’t address the real case

⚠️ Warning sign

If the agent says "I'll adjust this here" and starts editing code without showing a reproduction, stop it. Ask first: "How do you reproduce the bug?" If it can't answer, it's guessing.

Test
repro shows that it exists
Measured
same test measures the fix
Anti-guessing
no repro = no certainty
Actionable
repro becomes a test later
6

🛡️ Regression tests: every fixed bug becomes a test

A bug that came back is the most frustrating of all — because it has already been fixed once. The rule is simple: every fixed bug becomes a regression test. If you don’t have a test for the case, it will come back.

The best time to write this test is right after you’ve fixed the bug—the scenario is fresh, you know exactly what was failing, and the reproduction already exists (you used it in /diagnose).

💻 Example: regression test

# Bug reportado em produção: "comentário com emoji 🎉 quebrava a API"
# Causa encontrada via /diagnose: encoding latin-1 num middleware antigo
# Fix aplicado: trocar pra utf-8
# Teste de regressão (commit junto com o fix):

def test_comentario_com_emoji_nao_quebra_api():
    """Regression: bug #482 — emoji no body causava UnicodeDecodeError."""
    payload = {"texto": "Show de bola 🎉🚀"}

    response = client.post("/api/comentarios", json=payload)

    assert response.status_code == 201
    assert response.json()["texto"] == "Show de bola 🎉🚀"

# Anote no docstring: número do bug, causa raiz. Quando o teste falhar
# de novo em 6 meses, quem ler vai entender o porquê de existir.

💡 If there's no test, the bug comes back

It's not pessimism, it's statistics. Future refactors, dependency changes, someone cleaning up “old code nobody uses” — any of these can bring the bug back. The test is the automatic reminder: "this case here needs to keep working".

📊 Docstring convention

  • • Prefix: Regression: make it explicit that the test exists because of a bug
  • • Reference: issue/ticket/PR number to track context
  • • Root cause: one line — so whoever edits the test later understands
  • • DO NOT delete: a regression test is an asset of the project
Lock
the bug doesn't come back without an alert
Memory
documents the root cause
Cheap
repro already exists; turn it into a test
Cumulative
suite becomes a shield
7

🔀 Combining /tdd + /diagnose

The two loops don't compete—they work together. /tdd is for new work when you know what you want. /diagnose is for understanding why something existing isn't doing what it should. Complex bugs almost always involve both.

🎬 Workflow: intermittent bug in production

Imagine: a bug report says "sometimes the payment is duplicated." You can’t reproduce it locally. What should you do?

  1. 1. /diagnose — reproduce: can’t. Go back a step: instrument production (transaction ID, timestamp, request ID logs).
  2. 2. /diagnose — hypothesise: client retry? race condition in the worker? duplicate gateway webhook?
  3. 3. /diagnose — instrument + minimize: logs show that two workers pick up the same job when the client retries in <1s. Found it: a race condition.
  4. 4. /tdd takes over from here — red: write a test that dispatches two workers simultaneously for the same payload. It fails (duplicates).
  5. 5. /tdd — green: implements an idempotency key in the payments table. Test passes.
  6. 6. /tdd — refactor + regression: cleans up the code and keeps the test as a permanent regression test. The bug won’t come back.

🧭 When to switch

  • • Do you know what you want to build? → Go straight to /tdd
  • • Have a bug and know how to reproduce it? → Use /diagnose to find the root cause, then /tdd for the fix
  • • Have a bug and DON’T know how to reproduce it? → /diagnose starts with remote instrumentation before any fix
  • • Refactor with existing tests? → It stays in the green-refactor stage of /tdd
  • • Refactor WITHOUT tests? → Start with /tdd: cover the current behavior with tests; refactor only afterward
Complementary
doesn’t choose one; use both
Handoff
diagnose hands off to tdd
Repro→Test
the bridge between the two
Definitive
fix + test = protection
8

🏋️ Hands-on exercise

Pick an open bug in your project — any one, preferably one you've been avoiding because "it's weird." Run /diagnose and document each step. By the end, you'll have the bug fixed AND review material to understand how agents handle bugs.

📋 Exercise outline

  1. 1. Choose the bug. Paste into the chat: problem description, stack trace if available, minimal context.
  2. 2. Run /diagnose and follow the 6 steps WITHOUT skipping any.
  3. 3. At each step, make a note in a diagnose-log.md: what the agent did, what it discovered, what still wasn’t clear.
  4. 4. When you get to the fix, BEFORE applying it, write the regression test (step 6 of /diagnose).
  5. 5. Run the test — it should fail. Apply the fix. Run it again — it should pass.
  6. 6. Commit the fix and test together, with a message that references the root cause.

💡 Success criteria

It's not “the bug was fixed.” It's “I can prove the bug was fixed, and I have a test that will sound the alarm if it comes back.” If you can't check both boxes, redo the exercise.

Concrete
real bug, not an example
Documented
log of each step
Verifiable
the test proves the fix
Reusable
log becomes a future reference

📚 Module Summary

✓
The agent needs objective evidence — "looks right" doesn't count; a running test does.
✓
/tdd: red → green → refactor — one step at a time; see the test fail before implementing.
✓
Vertical slices > horizontal slices — a complete end-to-end feature, demoable at the end of each slice.
✓
/diagnose: 6 steps for difficult bugs — reproduce, minimise, hypothesise, instrument, fix, regression-test.
✓
Reproduce before fixing — without a repro, a "fix" is a guess; with a repro, a test measures the fix.
✓
Fixed bug = regression test — if there's no test, the bug comes back. Statistics, not pessimism.
✓
/diagnose + /tdd complement each other — diagnose finds the cause; tdd implements the permanent fix.

Next Module:

1.5 — Next steps in the Fundamentals track