π§ Why bad RAG exists
Most RAGs fail for the same reason: treating all data the same. βChunk everything, embed it, retrieve the top-k.β That works for simple use casesβa FAQ, a manual. But as soon as the data gets complex, the approach falls apart, and the system starts retrieving the wrong information.
π― The central thesis
There is no "the RAG." There are different architectures, and which one fits depends on the data:
- β’A catalog with SKU codes needs exact routing, not just vector search.
- β’Tickets, PDF manuals, and Notion docs need separate indexes, not all at once.
- β’Legal cases with citations need graph traversal, not similarity search.
β οΈ The symptom
"My RAG isn't retrieving the right information." Almost always, the cause isn't the model or the prompt β it's the architecture chosen before understanding the data. Choosing wrong here costs weeks of rework.
π₯ Who it's for (and who decides)
Itβs for anyone connecting data to an LLM. You donβt even need to know what RAG means. You describe what you have and what you want to ask; the skill figures out which architecture fits, explains why, and delivers the code. Thatβs the consulting approach: the client talks business, the skill talks engineering.
π The signals the skill reads in your data
- Text length distribution β short, uniform records vs. long docs define the chunking strategy.
- Relationships β cross-references and mentions between files: high interconnection β Graph RAG.
- Cardinality β ID/category columns vs. free text: only free text benefits from embeddings.
- Scale β small (<1K docs), medium (1Kβ100K), large (100K+) affects the choice of vector database.
- Consultation pattern β do the questions mix semantic intent with exact lookups (SKU, IDs)? That signals hybrid search.
β How to describe the data well
- β"3,000 support tickets, inconsistent formatting."
- β"Catalog with SKU and description, questions mix code and text."
- β"I'm in Supabase and want to ask in natural language."
β What gets in the way of the decision
- β"I want a RAG" β without saying anything about the data.
- βRequest Graph RAG because it sounds advanced.
- βHiding that the data is in four places.
ποΈ The Four Architectures
The heart of the skill: four architectures, each with a clear trigger. Knowing all four and when each one fits is what separates a useful recommendation from a guess.
1 Β· Naive RAG
simple, uniform dataWhen: well-structured data, a single source, uniform document types, direct queries.
How it works: chunk β embed β vector DB β top-k β generate response.
Fits into: Company FAQ, product documentation, a single manual.
2 Β· Advanced RAG
semantic + exact codesWhen: data that mixes semantic content with exact identifiers (SKU, IDs, ticket numbers); queries blend natural language with structured filters.
How it works: Naive + hybrid search (vector + BM25), query analysis (detect exact patterns), reranking, and direct routing to look up exact fields.
Performance tip: cache deterministic lookups β an SKU query always returns the same result.
3 Β· Modular / Agentic RAG
multi-source with a routerWhen: multiple sources, each with its own structure, update frequency, and characteristics; sometimes the query needs to be broken down.
How it works: orchestration layer with query router, indexes separated by source (not a single index with a filter), parallel retrieval, fusion with source diversity, and reranking.
Fits into: knowledge base spanning Notion, Confluence, Slack, and a SQL database.
4 Β· Graph RAG
relationships between entitiesWhen: highly relational data, where the relationships between entities matter as much as the content; questions involve traversals ("who worked with whom?", "all cases that mention X").
How it works: knowledge graph (entities + relationships) on top of vector searchβentity extraction, graph construction (Neo4j or Apache AGE in Postgres), and a query classifier that routes to graph, vector, or both.
Fits into: legal databases with citations, organizational data, biomedical literature.
π§© Practical tip: don't over-engineer
Match complexity to the problem. Donβt recommend Graph RAG for a simple FAQ. A good recommendation also explains why the alternatives were ruled outβnot just why the winner was chosen.
βοΈ Implementation decisions
Once you choose the architecture, the skill delivers a concrete plan. What separates a good plan from a generic one is justify the magic numbers β chunk size, top-k β with reasoning the user can use to adjust them later.
| Decision | Default | How to justify |
|---|---|---|
| Chunk size | 800β1500 tokens | Based on the text's length, not arbitrarily. |
| Embedding | text-embedding-3-small | Dimension trade-off if relevant. |
| Vector DB | Supabase pgvector | No new infrastructure if itβs already in Supabase. |
| top-k candidates | ~20 before reranking | Sufficient diversity, latency <200ms. |
π Architecture-based recovery method
- Naive: top-k similarity search.
- Advanced: hybrid search + reranking + query preprocessing + exact routing.
- Modular: source-based retrieval with routing logic, parallel execution, and diversity-aware fusion.
- Graph: query classification β graph traversal / vector search / both β merge.
π§© Stack awareness
If the user mentions their stack (Supabase, Vercel, Python, n8n), the recommendation has to use it. If youβre on Supabase, you get pgvector, not βset up a Qdrant.β If youβre using Python, you get a Python scaffold. Generic recommendations that ignore your existing infrastructure are actively useless.
π¦ Runnable Scaffold, Not a Stub
The difference between a useful skill and a boilerplate generator: code that runs on the first try. The skill generates a mini-project with complete SQL migrations and a README that guides you through setup from start to finish.
/rag-setup βββ ingest.ts # chunk, embed, upsert no vector DB βββ query.ts # pipeline de recuperacao + geracao βββ config.ts # modelo, conexao, chunk size + SQL migrations βββ README.md # o que rodar e em que ordem
β Real scaffold
- βComplete SQL in
config.tsβ copy and paste into the Supabase editor and it works. - βEach source has working ingestion and queries, not placeholder comments.
- βOfficial SDKs (
openai,@anthropic-ai/sdk), notfetch()raw. - βREADME with
npm installexact and test command.
β What doesn't count as a deliverable
- β
// implemente a ingestΓ£o aquias the function body. - β"See the recommendation for SQL" instead of the actual SQL.
- βExternal API (Notion) with no real endpoint, header, or parsing.
- βREADME with no dependencies or environment variables.
π Golden rule
A user following the README from top to bottom should end up with a working system. The only values to fill in are marked with // TODO: β everything else is real, runnable code.
π οΈ Packaging your architect
This is the most advanced example in the track: a skill that engineering judgment, not just text generation. The SKILL.md guides five phases β understand the data, recommend, plan, scaffold, and (optionally) build a companion.
Understanding the data
Capture the signals that matter (size, relationships, cardinality, scale, query pattern) β concise, without turning it into an essay.
Recommend
Which of the 4 architectures to use, why it fits, what would go wrong with something simpler, and the honest trade-offs.
Plan & generate scaffold
Chunking/embedding/DB/retrieval plan + the runnable mini-project with migrations and README.
Build companion (optional)
"Want me to help build and test it?" β connect the pipeline to the real data, run test queries, tune chunk/top-k.
--- name: rag-architect description: | Projeta o RAG ideal para os dados do usuario. Analisa estrutura, formato, relacoes e escala para recomendar a arquitetura certa e entregar plano + codigo inicial. Dispara em qualquer pedido sobre RAG, busca vetorial, semantica, "conversar com meus dados", chunking ou qualidade de retrieval β mesmo sem dizer "RAG". --- # RAG Architect Voce e um arquiteto de sistemas RAG. ## Fase 1 β Entender o dado Pergunte: que dado, quanto, que perguntas, qual stack (default: Next.js + Supabase + Claude API). ## Fase 2 β Recomendar 1 de 4 Naive | Advanced | Modular/Agentic | Graph => qual, por que, o que daria errado simples, trade-offs. ## Fase 3 β Plano chunk + overlap (justificado), embedding, vector DB, retrieval. ## Fase 4 β Scaffold rodavel /rag-setup: ingest.ts, query.ts, config.ts (com SQL!), README. SEM stubs. SDKs oficiais. Stack do usuario sempre. ## Fase 5 β Build companion (opcional) "Quer construir e testar?" => wire + tune.
β‘ Prompt to trigger the skill
π Module Summary
π― Hands-on exercises
- For 3 datasets (FAQ; catalog with SKU; legal database with citations), choose the architecture and justify why you ruled out the others.
- Justify a chunk size and a top-k for long docs versus short records β explain the reasoning, not just the number.
- Sketch the
config.tswith SQL migrations for a Naive RAG in pgvector (table + index + match function). - Create a runnable SKILL.md called
rag-architectthat guides you through the 5 phases, recommends 1 of 4 architectures based on a data description, and generates the scaffold/rag-setup. Test it with the prompt from topic 6.
End of Track 5:
Next: Track 6 β Advanced Skills Architecture