🪣 Symptoms of a ball of mud
"Big ball of mud" is the architectural term for code where nothing belongs: mixed responsibilities, circular dependencies, modules that grow by accretion. You recognize it by the smell: every simple change cascades across three unrelated files.
✓ Healthy code
- ✓Each module has a clear, nameable responsibility
- ✓Dependencies flow in one direction (high → low level)
- ✓Consistent vocabulary across code, tests, and docs
- ✓Duplication is an exception and is justified
- ✓You predict where a new feature will live
- ✓Changing one implementation doesn’t leak into 10 places
✗ Ball of mud
- ✗"Giant "helper" / "utils" / "core" modules with no focus
- ✗A imports B, which imports A (cycles)
- ✗The same concept has 3 names (user / customer / account)
- ✗Copy-paste with small changes in several places
- ✗Every new feature requires a "tour" to find out where to put it
- ✗Changing one rule requires hunting through 8 files
Typical smells (architectural code smells)
- • God modules: a file with 2,000 lines that "everyone imports"
- • Shotgun surgery: a simple change touches 15 files
- • Feature envy: methods that operate more on another class's data than on their own
- • Magic strings / numbers scattered without a central constant
- • Conditionals by type: giant if/elif checking "type === 'X'"
- • Wrappers that only pass things along: layers that don't make any decisions
🔍 /improve-codebase-architecture
The command /improve-codebase-architecture doesn’t ask you to "refactor everything."
It does the most useful thing: list deepening opportunities — points where
the code can become significantly healthier, ordered by impact. The analysis is informed by the
CONTEXT.md and the project's ADRs, so the suggestions respect
what has already been decided.
⚙️ What the command does
- •Reads CONTEXT.md to understand the project's domain and vocabulary
- •Reads the ADRs to learn which architectural decisions are in effect
- •Scans the code structure for smells that contradict the domain
- •Returns a prioritized list of opportunities, not a rewrite plan
- •Each item has: description, estimated impact, effort, and a starting point
Example output
terminal$ /improve-codebase-architecture # 5 deepening opportunities identified (sorted by impact) [1] HIGH src/billing/ — "materialization cascade" lives in 4 modules Impact: unifica lógica fragmentada, fecha conceito do CONTEXT.md Effort: 2-3 dias | Risk: medium Start: src/billing/cascade.py (extract; ADR-009 references this) [2] HIGH src/utils/helpers.py — 1,847 lines, 23 unrelated functions Impact: quebra god module; melhora descoberta Effort: 1-2 dias | Risk: low (mostly mechanical) Start: agrupar por consumer; mover; testar [3] MED "customer" vs "user" vs "account" usados intercambiavelmente Impact: alinha código ao glossário do CONTEXT.md (Sec. 2.3) Effort: 1 dia | Risk: low Start: renomear customer.* (canonical) e atualizar callsites [4] MED ciclo: api/orders → services/inventory → api/orders Impact: remove ciclo; clarifica direção de dependência Effort: meio dia | Risk: low Start: extract InventoryPort interface [5] LOW tests/ duplica fixtures de billing em 6 arquivos Impact: reduz manutenção de testes Effort: 2h | Risk: very low Start: conftest.py compartilhado por subdir
💡 Tip
Run this command before for planning the next cycle. It gives you an honest list of architectural debt so you can choose 1 item per sprint—not tackle everything at once.
🔭 /zoom-out — broader perspective
There’s a common failure pattern: the agent opens a file, starts optimizing a function, and after
20 minutes is rewriting isolated logic — without realizing that that function
shouldn't even exist in this module. /zoom-out is for
this: forcing the agent to stop and ask for higher-level context before continuing.
Initial zoom-in (detail)
The agent opens file X and starts editing
A narrow focus is necessary to make the immediate change, but it becomes a trap when the function, file, or even the entire module is in the wrong place. The agent "digs in" without asking whether the problem is bigger.
Zoom-out (overview)
/zoom-out — asks for architectural context
The command moves up a level: what is this module responsible for? What does CONTEXT.md say? Where does this logic should belong? The right question changes from "how can I improve this function" to "does this function belong here?". Often, the best edit is move or remove, don't optimize.
Informed zoom-in (detail with a map)
Back to the code, but with the big picture
Now local editing is guided by the architecture. You know what NOT to do (create new coupling, duplicate logic from another module) and what TO do (extract, move, align names with the glossary). Details matter, but they serve the structure.
💡 Use /zoom-out when…
- • You've been editing the same file for >15 minutes with no clear progress
- • A simple change is turning into 4 different files
- • The agent is "fixing" something that's actually a symptom of another problem
- • Before accepting a major suggestion that touches multiple modules
🌱 Deepening opportunities, not rewrites
Deepening is a term borrowed from Eric Evans (DDD): deepen the modeling where the domain matters most, without throwing away what already works. It's the opposite of the impulse to "rewrite from scratch"—which is expensive, rarely delivers what it promises, and discards all the learning embedded in the current code.
✗ Destructive rewrite
- ✗"Let's throw it out and redo it with the new architecture"
- ✗Parallel branch for 3–6 months without delivering
- ✗Misses edge cases the current code already handles
- ✗Big bang merge → mass silent regressions
- ✗Team gets stuck on the new work; support for the old work falls behind
- ✗Lessons learned (CONTEXT.md, ADRs) turn into junk
✓ Incremental deepening
- ✓"Let's improve 1 opportunity per sprint without stopping delivery"
- ✓Each PR is small, reviewable, and mergeable
- ✓Keeps edge cases while refining the form
- ✓Regressions show up early, in small changes
- ✓Team keeps delivering features alongside it
- ✓CONTEXT.md and ADRs evolve together, becoming a living reference
📖 Term: deepening
Deepening = deepen domain modeling incrementally. Instead of rewriting, you identify concepts that are "crooked" in the code (hidden in utils, scattered around, poorly named) and elevate them to first-class concepts—one extraction at a time.
Origin: Domain-Driven Design (Eric Evans, 2003) — "deepening the model".
🧭 CONTEXT.md as a guide
O CONTEXT.md isn’t just onboarding — it’s
the architectural standard of the project. The domain language it
describes guides where to refactor: modules that implement core concepts deserve more architectural attention
than generic infrastructure.
🎯 Principle: domain concept = architectural attention
Example: if CONTEXT.md describes “materialization cascade” as a key concept in the billing system —
the chain through which price changes propagate to active contracts — then the module that implements
that cascade can’t be hidden in utils/helpers.py
split into 4 separate functions.
The obvious deepening here is to extract billing/cascade.py with the cascade as a named entity,
aligned with the CONTEXT.md vocabulary. The command /improve-codebase-architecture see this
misalignment because it reads CONTEXT.md.
Questions that CONTEXT.md answers
- • What are the core concepts? → They deserve their own modules and consistent names
- • What is the canonical glossary? → “customer” vs. “user” is settled here, not in PR review
- • What are the boundaries? → where a feature belongs is determined by the bounded context
- • What invariants exist? → rules that MUST NOT be violated during refactoring
💡 Tip
If CONTEXT.md is vague or out of date, this is the first deepening opportunity. Without it, any architectural improvement will be local and lose coherence.
📜 ADRs as architectural memory
ADR = Architecture Decision Record. One short document for each significant architectural decision. Without ADRs, the next refactor silently undoes a deliberate choice—because nobody remembers why X was separated from Y.
ADR-009 — example
docs/adr/0009-separate-cascade.md# ADR-009: Separar materialization cascade de pricing ## Status Accepted — 2025-08-14 ## Contexto A "materialization cascade" propaga mudanças de preço para contratos ativos. Originalmente vivia dentro de pricing/calculator.py porque parecia "cálculo de preço". Na prática, são duas responsabilidades: - pricing/ → calcula o preço de UM item, sem estado - cascade/ → reage a mudanças, atualiza N contratos, com estado Misturadas, qualquer mudança em pricing arriscava efeito colateral em milhares de contratos ativos. ## Decisão Extrair cascade para billing/cascade.py como entidade nomeada e independente. pricing/ vira pure function. cascade/ orquestra. ## Consequências + pricing testável sem mock de DB + cascade tem ownership claro de invariantes (idempotência, ordem) + vocabulário do CONTEXT.md (Sec. 4) finalmente reflete o código - duas pastas onde antes havia uma (custo de descoberta) - imports adicionais entre os módulos (acoplamento explícito) ## Alternativas consideradas - Manter junto: rejeitado, acopla cálculo a propagação - Event bus: prematuro, sem necessidade de async hoje - Inline em cada caller: rejeitado, duplicação garantida
🧠 Why ADRs matter in deepening
- • Without an ADR, in 6 months someone will "consolidate" the cascade back into pricing because it seems redundant
- • The ADR is proof that the separation has already been considered and there's a reason
- •
/improve-codebase-architecturereads ADRs before suggesting anything — it won’t propose something that’s already been rejected - • Discontinued ADRs are useful too: they record what we learned ("we tried X; it didn't work because Y")
💡 Tip: short ADR > perfect ADR
A 1-page ADR beats a 10-page ADR nobody writes. Status, Context, Decision, Consequences—
four headings, short paragraphs. Number them sequentially (0001, 0002…) and don’t delete them: mark them as
Superseded by ADR-XXX when replaced.
📈 Example of progressive improvement
How a team gets out of a ball of mud without stopping the product. A 3-sprint narrative based on a real pattern: one opportunity per sprint, CONTEXT.md and ADRs updated every cycle.
Sprint 0 — Diagnosis
We ran /improve-codebase-architecture for the first time
The output lists 5 opportunities. The team discusses them for 30 minutes, chooses 3 to tackle over the next 3 sprints, in order of impact. The other 2 explicitly go into the backlog — not forgotten, not urgent.
Sprint 1 — Extract the cascade
Opportunity #1: scattered materialization cascade
Team extracts billing/cascade.py as a named module. Small PR, ~400 lines moved, no behavior changes. Write ADR-009 documenting the decision. Update CONTEXT.md (Sec. 4) with the canonical vocabulary. New features continue alongside it, without a parallel branch.
Sprint 2 — Break up the god module
Opportunity #2: utils/helpers.py with 1,847 lines
Mostly mechanical: group functions by consumer, move them into nearby modules. 3 “orphan” functions raise questions — what are they? The answer comes from CONTEXT.md (one is a cascade and should have gone in S1; the other two expose a new concept that becomes ADR-010). Bug avoided.
Sprint 3 — Align vocabulary
Opportunity #3: customer / user / account used interchangeably
Renaming guided by the CONTEXT.md glossary. It’s mechanical, but touches many files—feature flags don’t help here, so: small PRs by subdirectory, approved by that subdirectory’s code owner. In the end, the code speaks the same language as the docs and product teams.
Sprint 4 — Re-diagnosis
Run /improve-codebase-architecture again
New list: the 3 original opportunities are gone (they were resolved). The 2 that remained in the backlog are still there. But 2 new ones appeared because the domain evolved over the 3 sprints. Architecture improved visibly, no big bang, no stopping delivery.
📊 Result in 4 sprints
- •3 high-impact opportunities resolved, 2 new ADRs
- •CONTEXT.md updated 3 times, becoming a living reference
- •Zero pauses in product features
- •Bug avoided (orphan cascade in helpers) found by coincidence
- •Team trusts the process: the next round is already planned
🎯 Hands-on exercise
You don’t learn deepening by reading. You learn by running the command in your project and defending the choice of the first opportunity to tackle.
🛠️ Steps
-
1.
Confirm the basics: does your project have a CONTEXT.md? Does it have at least 1-2 ADRs? If not, start there (that's the pre-deepening).
-
2.
Run /improve-codebase-architecture in the actual project (not in a sandbox).
-
3.
Write down the first 3 suggestions with: description, impact, effort. Not all 5 — just the first 3.
-
4.
Which one do you tackle first? It may be #1 — but it may not be. Defend your choice in 3 sentences: why this one, why now, and what the risk is if you leave it.
-
5.
Bonus: draft the ADR before touching the code. If you can't write 1 page explaining why, maybe the suggestion isn't mature enough yet.
💡 Selection criterion
Good first choice: high impact, low-to-medium effort, low risk, touches the concept in CONTEXT.md. A bad first choice: high impact but high risk—save it for when the team already trusts the process.
⚠️ Do not
- ✗Tackle all 5 opportunities at once — you'll end up rewriting
- ✗Skip the ADR — in 6 months, no one remembers why
- ✗Refactor without a pre-existing test covering the path—write one first
✅ Module Summary
Next Module:
1.6 — continuing the skills journey.