PTENES
MODULE 5.3

πŸš€ RAG Architect

The skill that looks at your data and chooses among four RAG architectures and explains whyβ€” before you write a single line of code. Delivers an implementation plan and runnable starter code, not stubs.

6
Topics
50
Minutes
Advanced
Level
Practice
Type
1

🧭 Why bad RAG exists

Most RAGs fail for the same reason: treating all data the same. β€œChunk everything, embed it, retrieve the top-k.” That works for simple use casesβ€”a FAQ, a manual. But as soon as the data gets complex, the approach falls apart, and the system starts retrieving the wrong information.

🎯 The central thesis

There is no "the RAG." There are different architectures, and which one fits depends on the data:

  • β€’A catalog with SKU codes needs exact routing, not just vector search.
  • β€’Tickets, PDF manuals, and Notion docs need separate indexes, not all at once.
  • β€’Legal cases with citations need graph traversal, not similarity search.

⚠️ The symptom

"My RAG isn't retrieving the right information." Almost always, the cause isn't the model or the prompt β€” it's the architecture chosen before understanding the data. Choosing wrong here costs weeks of rework.

Chunk
Document excerpt.
Embedding
Meaning vector.
Retrieval
Search for what’s relevant.
Architecture
The decision that matters.
2

πŸ‘₯ Who it's for (and who decides)

It’s for anyone connecting data to an LLM. You don’t even need to know what RAG means. You describe what you have and what you want to ask; the skill figures out which architecture fits, explains why, and delivers the code. That’s the consulting approach: the client talks business, the skill talks engineering.

πŸ” The signals the skill reads in your data

  • Text length distribution β€” short, uniform records vs. long docs define the chunking strategy.
  • Relationships β€” cross-references and mentions between files: high interconnection β†’ Graph RAG.
  • Cardinality β€” ID/category columns vs. free text: only free text benefits from embeddings.
  • Scale β€” small (<1K docs), medium (1K–100K), large (100K+) affects the choice of vector database.
  • Consultation pattern β€” do the questions mix semantic intent with exact lookups (SKU, IDs)? That signals hybrid search.

βœ“ How to describe the data well

  • βœ“"3,000 support tickets, inconsistent formatting."
  • βœ“"Catalog with SKU and description, questions mix code and text."
  • βœ“"I'm in Supabase and want to ask in natural language."

βœ— What gets in the way of the decision

  • βœ—"I want a RAG" β€” without saying anything about the data.
  • βœ—Request Graph RAG because it sounds advanced.
  • βœ—Hiding that the data is in four places.
3

πŸ—‚οΈ The Four Architectures

The heart of the skill: four architectures, each with a clear trigger. Knowing all four and when each one fits is what separates a useful recommendation from a guess.

Describe the data Naive Β· simple data Advanced Β· hybrid + SKU Modular Β· multi-source Graph Β· relationships Plan + code

1 Β· Naive RAG

simple, uniform data

When: well-structured data, a single source, uniform document types, direct queries.

How it works: chunk β†’ embed β†’ vector DB β†’ top-k β†’ generate response.

Fits into: Company FAQ, product documentation, a single manual.

2 Β· Advanced RAG

semantic + exact codes

When: data that mixes semantic content with exact identifiers (SKU, IDs, ticket numbers); queries blend natural language with structured filters.

How it works: Naive + hybrid search (vector + BM25), query analysis (detect exact patterns), reranking, and direct routing to look up exact fields.

Performance tip: cache deterministic lookups β€” an SKU query always returns the same result.

3 Β· Modular / Agentic RAG

multi-source with a router

When: multiple sources, each with its own structure, update frequency, and characteristics; sometimes the query needs to be broken down.

How it works: orchestration layer with query router, indexes separated by source (not a single index with a filter), parallel retrieval, fusion with source diversity, and reranking.

Fits into: knowledge base spanning Notion, Confluence, Slack, and a SQL database.

4 Β· Graph RAG

relationships between entities

When: highly relational data, where the relationships between entities matter as much as the content; questions involve traversals ("who worked with whom?", "all cases that mention X").

How it works: knowledge graph (entities + relationships) on top of vector searchβ€”entity extraction, graph construction (Neo4j or Apache AGE in Postgres), and a query classifier that routes to graph, vector, or both.

Fits into: legal databases with citations, organizational data, biomedical literature.

🧩 Practical tip: don't over-engineer

Match complexity to the problem. Don’t recommend Graph RAG for a simple FAQ. A good recommendation also explains why the alternatives were ruled outβ€”not just why the winner was chosen.

4

βœ‚οΈ Implementation decisions

Once you choose the architecture, the skill delivers a concrete plan. What separates a good plan from a generic one is justify the magic numbers β€” chunk size, top-k β€” with reasoning the user can use to adjust them later.

Decision Default How to justify
Chunk size800–1500 tokensBased on the text's length, not arbitrarily.
Embeddingtext-embedding-3-smallDimension trade-off if relevant.
Vector DBSupabase pgvectorNo new infrastructure if it’s already in Supabase.
top-k candidates~20 before rerankingSufficient diversity, latency <200ms.

πŸ“Š Architecture-based recovery method

  • Naive: top-k similarity search.
  • Advanced: hybrid search + reranking + query preprocessing + exact routing.
  • Modular: source-based retrieval with routing logic, parallel execution, and diversity-aware fusion.
  • Graph: query classification β†’ graph traversal / vector search / both β†’ merge.

🧩 Stack awareness

If the user mentions their stack (Supabase, Vercel, Python, n8n), the recommendation has to use it. If you’re on Supabase, you get pgvector, not β€œset up a Qdrant.” If you’re using Python, you get a Python scaffold. Generic recommendations that ignore your existing infrastructure are actively useless.

5

πŸ“¦ Runnable Scaffold, Not a Stub

The difference between a useful skill and a boilerplate generator: code that runs on the first try. The skill generates a mini-project with complete SQL migrations and a README that guides you through setup from start to finish.

generated scaffold /rag-setup
/rag-setup
β”œβ”€β”€ ingest.ts     # chunk, embed, upsert no vector DB
β”œβ”€β”€ query.ts      # pipeline de recuperacao + geracao
β”œβ”€β”€ config.ts     # modelo, conexao, chunk size + SQL migrations
└── README.md     # o que rodar e em que ordem

βœ“ Real scaffold

  • βœ“Complete SQL in config.ts β€” copy and paste into the Supabase editor and it works.
  • βœ“Each source has working ingestion and queries, not placeholder comments.
  • βœ“Official SDKs (openai, @anthropic-ai/sdk), not fetch() raw.
  • βœ“README with npm install exact and test command.

βœ— What doesn't count as a deliverable

  • βœ—// implemente a ingestΓ£o aqui as the function body.
  • βœ—"See the recommendation for SQL" instead of the actual SQL.
  • βœ—External API (Notion) with no real endpoint, header, or parsing.
  • βœ—README with no dependencies or environment variables.

πŸ“Š Golden rule

A user following the README from top to bottom should end up with a working system. The only values to fill in are marked with // TODO: β€” everything else is real, runnable code.

6

πŸ› οΈ Packaging your architect

This is the most advanced example in the track: a skill that engineering judgment, not just text generation. The SKILL.md guides five phases β€” understand the data, recommend, plan, scaffold, and (optionally) build a companion.

1

Understanding the data

Capture the signals that matter (size, relationships, cardinality, scale, query pattern) β€” concise, without turning it into an essay.

2

Recommend

Which of the 4 architectures to use, why it fits, what would go wrong with something simpler, and the honest trade-offs.

3

Plan & generate scaffold

Chunking/embedding/DB/retrieval plan + the runnable mini-project with migrations and README.

4

Build companion (optional)

"Want me to help build and test it?" β€” connect the pipeline to the real data, run test queries, tune chunk/top-k.

SKILL.md β€” rag-architect (skeleton) front matter + phases
---
name: rag-architect
description: |
  Projeta o RAG ideal para os dados do usuario. Analisa estrutura,
  formato, relacoes e escala para recomendar a arquitetura certa e
  entregar plano + codigo inicial. Dispara em qualquer pedido sobre
  RAG, busca vetorial, semantica, "conversar com meus dados", chunking
  ou qualidade de retrieval β€” mesmo sem dizer "RAG".
---

# RAG Architect

Voce e um arquiteto de sistemas RAG.

## Fase 1 β€” Entender o dado
Pergunte: que dado, quanto, que perguntas, qual stack
(default: Next.js + Supabase + Claude API).

## Fase 2 β€” Recomendar 1 de 4
Naive | Advanced | Modular/Agentic | Graph
=> qual, por que, o que daria errado simples, trade-offs.

## Fase 3 β€” Plano
chunk + overlap (justificado), embedding, vector DB, retrieval.

## Fase 4 β€” Scaffold rodavel
/rag-setup: ingest.ts, query.ts, config.ts (com SQL!), README.
SEM stubs. SDKs oficiais. Stack do usuario sempre.

## Fase 5 β€” Build companion (opcional)
"Quer construir e testar?" => wire + tune.

⚑ Prompt to trigger the skill

I have 3,000 support tickets in Postgres (Supabase), inconsistent formatting, and questions that mix problem descriptions with ticket numbers. What RAG architecture should I use, and give me the starter code.

πŸ“ Module Summary

βœ“
Data-driven architecture β€” treating all data the same is the #1 cause of incorrect RAG.
βœ“
Describe the data, not the RAG β€” the skill reads the signals (size, relationships, cardinality, scale, query) and decides.
βœ“
Four architectures β€” Naive, Advanced, Modular/Agentic, and Graph, each with a clear trigger.
βœ“
Justified numbers + stack awareness β€” chunking and top-k with reasoning; always use the user's infrastructure.
βœ“
Runnable scaffold β€” complete migrations and a guide README, no stubs.

🎯 Hands-on exercises

  1. For 3 datasets (FAQ; catalog with SKU; legal database with citations), choose the architecture and justify why you ruled out the others.
  2. Justify a chunk size and a top-k for long docs versus short records β€” explain the reasoning, not just the number.
  3. Sketch the config.ts with SQL migrations for a Naive RAG in pgvector (table + index + match function).
  4. Create a runnable SKILL.md called rag-architect that guides you through the 5 phases, recommends 1 of 4 architectures based on a data description, and generates the scaffold /rag-setup. Test it with the prompt from topic 6.

End of Track 5:

Next: Track 6 β€” Advanced Skills Architecture