PTENES
Skip to content
MODULE 2 DECISION

Context Design at Scale

This is where strategic decision, not for operational use. Learn when NOT to use context.

1

Large-Scale Context Strategies

Architectural decisions

The Challenge of Scale

In production systems with thousands of users, every context decision affects cost, latency, and quality. The architect defines policies, not prompts.

💰
Cost per Token

100K tokens/req = $$$$

⏱️
Latency

More context = slower

🎯
Relevance

More != better

Allocation Strategies

1
Static Allocation

Fixed context by request type. Predictable, but inflexible.

2
Dynamic Allocation

Context based on query complexity. Efficient, but complex.

3
Tiered Allocation

Context levels (basic, standard, premium). Balanced.

2

Deciding: Long Context, RAG, or Hybrid

Decision framework

Decision Tree

Dados mudam frequentemente?
├─ SIM → Dados > 1M tokens?
│        ├─ SIM → RAG obrigatório
│        └─ NÃO → RAG ou cache curto
│
└─ NÃO → Dados cabem no contexto?
         ├─ SIM → Análise relacional necessária?
         │        ├─ SIM → LONG CONTEXT
         │        └─ NÃO → Qualquer abordagem
         │
         └─ NÃO → Precisão máxima necessária?
                  ├─ SIM → HÍBRIDO (long + RAG seletivo)
                  └─ NÃO → RAG puro

Long Context

  • • Static data
  • • In-depth analysis
  • • Complex relationships
  • • High cost per query is OK

RAG

  • • Dynamic data
  • • Very large knowledge base
  • • Multi-tenant
  • • Cost per query is critical

Hybrid

  • • Fixed base + dynamic data
  • • Maximum accuracy
  • • Balanced cost
  • • Complexity OK
3

Cognitive and Technical Costs of Context

The hidden cost

Technical Costs

  • 💵 Financial Cost ($/1K tokens)
  • ⏱️ Increased latency
  • 📊 Reduced throughput
  • 🔧 Operational complexity

Cognitive Costs

  • 🧠 "Lost in the middle" effect
  • 🎯 Attention dilution
  • ⚡ Information conflicts
  • 🔀 Emerging inconsistencies

Trade-off Calculator

Context Cost/req Latency Accuracy
10K tokens ~$0.03 ~1s Base
50K tokens ~$0.15 ~3s +15%
100K tokens ~$0.30 ~6s +20%
200K tokens ~$0.60 ~12s +22%*

* Diminishing returns after ~100K

4

Context Governance

Organizational Policies

Governance defines who decides what about context. Without governance, each developer makes inconsistent decisions.

Access Policies

  • • Who can add context?
  • • Sensitivity levels
  • • Mandatory audit trail

Quality Policies

  • • Validation before injection
  • • Freshness requirements
  • • Format and structure

Budget Policies

  • • Limits by layer
  • • Quotas by team/project
  • • Excess alerts

Conflict Policies

  • • Priority among sources
  • • Automatic vs. manual resolution
  • • Escalation
5

Persistence and Eviction Policies

Context lifecycle

Lifecycle Strategies

Ephemeral
Request-scoped
Session
Conversation-scoped
Persistent
User/tenant-scoped
Global
System-wide

Eviction Policies

LRU (Least Recently Used)

Remove context that has not been accessed for the longest time

Priority-based

Remove according to set priority

TTL (Time-to-Live)

Remove after a set time

Summarization

Compress instead of removing

6

Systemic Failures Caused by Context

Failure modes

💥

Context Overflow

The system tries to inject more context than the limit allows.

Mitigation: Budget enforcement + graceful degradation

🔀

Context Contamination

One user's data leaks into another user's context.

Mitigation: Strict isolation + tenant boundaries

⚡

Context Conflict

Conflicting information from multiple sources.

Mitigation: Conflict resolution rules + source priority

📉

Context Staleness

Outdated context leads to incorrect answers.

Mitigation: TTL policies + freshness validation

Previous
Module 1: Architecture
Next
Module 3: Skills Governance

Download this module

Save for offline study