PTENES
Skip to content
MODULE 5 POST-RAG

Modern RAG and Hybrid Context

Rethinking RAG in the Long Context Era. When to use it, how to integrate it with fixed context, and strategies for reducing hallucinations with hybrid context.

6
Topics
90
Minutes
6
Exercises
1

When RAG Is Still Necessary

Use cases that justify RAG in 2024+

The Paradigm Shift

With 100K-1M+ token context windows, the question has shifted from "how do you use RAG?" to "when does RAG still make sense?". RAG is no longer the default solution; it has become a specialized tool.

✅ RAG Still Makes Sense

  • • Real-time data: Prices, inventory, status
  • • Base > 1M tokens: When it exceeds the context window
  • • Multi-tenant: Data isolated by user
  • • Cost: Long context is expensive for rarely used data
  • • Frequent updates: Data that changes constantly

❌ Long Context Is Better

  • • Documentation establishes: Manuals, policies, specs
  • • Codebases: Project Source Code
  • • In-depth analysis: When you need to see everything together
  • • Consistency: Answers depend on multiple docs
  • • Rich context: Nuances and interrelationships matter

RAG vs. Long Context Decision Matrix

Criterion RAG Long Context
Data changes constantly ✓ —
Volume > 1M tokens ✓ —
Maximum accuracy required — ✓
Analysis of relationships between docs — ✓
Cost per Query Matters ✓ —
2

RAG as a Complement to Long Context

Intelligent integration of approaches

Modern architecture doesn't choose between RAG and Long Context — it strategically combines. Each has its role in the context system.

Complementary Architecture

// Layered context with complementary RAG
┌─────────────────────────────────────────────┐
│  SYSTEM CONTEXT (Long Context - Fixo)       │
│  • Identidade, regras, políticas            │
│  • Skills disponíveis                       │
│  • Formato de resposta esperado             │
├─────────────────────────────────────────────┤
│  GLOBAL CONTEXT (Long Context - Sessão)     │
│  • Documentação do projeto                  │
│  • Codebase relevante                       │
│  • Histórico resumido                       │
├─────────────────────────────────────────────┤
│  DYNAMIC CONTEXT (RAG - Por query)          │
│  • Dados em tempo real                      │
│  • Resultados de busca                      │
│  • Informações específicas do momento       │
├─────────────────────────────────────────────┤
│  USER QUERY                                 │
└─────────────────────────────────────────────┘

🏗️ Foundation

Use Long Context for anything stable: identity, rules, baseline documentation.

🔄 Dynamic

RAG for changing data: prices, status, updates, external data.

⚡ On Demand

RAG triggered by need: search activated by keywords or a skill.

Example: Support Assistant

FIXED Product manual (50K tokens) + Policies (10K tokens)
RAG Customer request status + Recent ticket history
TRIGGER "Search for similar cases" enabled when escalation skill is called
3

Selective Knowledge Injection

Accuracy over quantity

The Traditional RAG Problem

Classic RAG retrieves chunks based on semantic similarity, but the most “similar” result isn’t always the most “relevant.” Selective injection solves this.

❌ Retrieve the top K by embedding
✓ Retrieve by intent + context

Selective Injection Strategies

1. Intent Classification Injection

# Primeiro: classificar a intenção
intent = classify_intent(user_query)
# → "troubleshooting" | "how_to" | "pricing" | "status"

# Depois: recuperar da fonte correta
if intent == "troubleshooting":
    context = retrieve_from("knowledge_base/errors")
elif intent == "pricing":
    context = retrieve_from("pricing_api", realtime=True)
elif intent == "status":
    context = retrieve_from("orders_db", user_id=user.id)

2. Injection via Active Skill

# Skill determina o que precisa
skill = "code_review"
required_context = skill.required_context()
# → ["file_content", "project_conventions", "recent_changes"]

# Injetar apenas o necessário
for ctx_type in required_context:
    inject_context(ctx_type, scope=skill.scope)

3. Conditional Injection

# Injetar apenas se necessário
if mentions_product(query):
    inject("product_specs", lazy=True)

if mentions_policy(query):
    inject("company_policies", section=detect_policy_type(query))

if requires_calculation(query):
    inject("pricing_formulas")
    inject("current_rates", source="api", cache=60s)

Benefits of Selectivity

🎯
Accuracy
Relevant context, not just similar
💰
Cost
Fewer tokens = lower cost per query
⚡
Latency
Smaller context = faster response
🧠
Focus
LLMs aren’t distracted by irrelevant context
4

Hybrid Context (Fixed + Retrieved)

The best of both worlds

Hybrid context combines the Long Context stability with the RAG dynamism, creating systems that are both consistent and up to date.

Hybrid Context Anatomy

Static
System + Global Context
~40%
Semi-static
Session + Preferences
~20%
Dynamic
RAG + Tools + APIs
~30%
Query
User Input
~10%

Standard: Context Window Budget

System Context 10K max
Global Knowledge 50K max
Session History 20K max
RAG Results 15K dynamic
Response Buffer 5K reserved
Total Budget 100K tokens

Management Strategies

Priority Eviction

Remove the lowest-priority context when it exceeds the budget

Lazy Loading

Load context only when needed

Summarization

Compress old history into summaries

Practical Example: E-commerce Assistant

context_budget = ContextBudget(total=100_000)

# Camada fixa (sempre presente)
context_budget.allocate("system",
    content=system_prompt + skills_definition,
    priority=CRITICAL, evictable=False)

# Camada por sessão
context_budget.allocate("catalog_summary",
    content=get_catalog_summary(),
    priority=HIGH, evictable=True)

# Camada dinâmica (por query)
if user_query.mentions_product():
    context_budget.allocate("product_details",
        content=rag.retrieve(user_query, top_k=5),
        priority=MEDIUM, evictable=True)

if user_query.mentions_order():
    context_budget.allocate("order_status",
        content=api.get_order(user.current_order),
        priority=HIGH, evictable=True)
5

Reducing Hallucinations Through Context

Grounding in concrete facts

Why LLMs Hallucinate

Hallucinations happen when the model needs to fill gaps in its knowledge. The solution isn’t to ask it to "not hallucinate"—it’s to provide the context that fills in the gaps.

Fact Gap

Model makes up data it doesn't know

Pressure to Respond

Prompt implies that there must be an answer

Grounding Strategies

1. Authoritative Context

Explicitly provide the source of truth, with instructions for using it.


{product_data}


Responda APENAS com informações presentes em
PRODUCT_DATABASE. Se a informação não estiver
disponível, diga "Não encontrei essa informação
no catálogo."

2. Mandatory Citation

Force the model to cite its sources, preventing fabrications.

Each factual claim must include [source: document#section].
If you can't cite a source from the provided context,
mark it as [source: general knowledge] or omit it.

3. Permission to Say "I Don't Know"

Reduce the pressure to make up answers.

It’s perfectly acceptable to respond:
- "I don’t have that information"
- "I’d need to check the system"
- "I can help with X, but I don’t have data about Y"

Partial answers are better than made-up ones.

Anti-Hallucination Framework

📚
Rich Context
Complete Data
🏷️
Citation
Traceability
🚫
Limits
Defined scope
✅
Validation
Post-check
6

Contextual Quality Evaluation

Metrics and diagnostics

How can you tell if your context is working? Specific metrics help diagnose problems and optimize the context system.

Quality Metrics

🎯 Relevance

% of context actually used in the answer

Target: > 70%
📊 Coverage

% of questions answered using available context

Target: > 90%
🔍 Accuracy

% of factually correct answers

Target: > 95%
⚡ Efficiency

Tokens used vs. response quality

Continuously optimize

Diagnostic Framework

❌
Frequent “I don’t know” answers

Diagnosis: Insufficient or poorly retrieved context

Action: Expand the knowledge base or improve retrieval

❌
Frequent hallucinations

Diagnosis: Context present but ignored, or gaps

Action: Improve grounding instructions, add mandatory citations

❌
Slow/expensive answers

Diagnosis: Context that is too large or poorly organized

Action: Implement selective injection, compress static context

❌
Inconsistencies between responses

Diagnosis: Context conflicts or lack of prioritization

Action: Establish a clear hierarchy and resolve conflicts at the source

Contextual Quality Checklist

✓ Does the context contain up-to-date information?
✓ Is the priority hierarchy defined?
✓ Conflicts resolved at the source?
✓ Token budget respected?
✓ Quality metrics being monitored?
✓ Fallbacks for RAG failures?

Download this module

Save for offline study

Previous
Module 4: Orchestration
Next
Module 6: Masterclass Preparation