How Models Interpret Long Context
Internal mechanisms and limitations
Models with 100k-2M token windows don't process all text equally. Understanding attention mechanisms and their biases is essential to optimizing the use of long contexts.
Attention Biases
Primacy Bias
The beginning of the context carries more weight
Lost in the Middle
Center tends to be ignored
Recency Bias
The end of the context has high priority
Practical Implications
- → Put critical instructions in the beginning of the context
- → Repeat important information in the final (task reminder)
- → Use explicit markers for content in the middle
- → Test information retrieval in different positions
Large Windows vs. Traditional RAG
When to use each approach
With 1-2M tokens available, you can load entire documents into the context window. This fundamentally changes the need for RAG—but doesn't eliminate it.
📚 Full Context Loading
When to use:
- • Corpus fits in the context window (<1M tokens)
- • Requires a holistic view
- • Cross-reference between documents
- • Retrieval latency is a problem
🔍 Selective RAG
When to use:
- • Corpus is very large (TB of data)
- • Data changes frequently
- • Cost per token is critical
- • Requires surgical precision
Decision Matrix
| Scenario | Full Context | RAG |
|---|---|---|
| 50-page manual | ✅ Ideal | Unnecessary |
| 10M-document base | Impossible | ✅ Required |
| 500k-line codebase | ✅ Feasible | ✅ Alternative |
| Real-time news feed | Impractical | ✅ Ideal |
Long-Context Organization Strategies
Structuring 100k-1M tokens
Organization Techniques
📑 Table of Contents
Include an index at the beginning listing sections and their positions. The model uses it as a navigation map.
🏷️ Section Headers
Use clear, consistent delimiters: === SECTION: Name ===
🎯 Relevance Ordering
Sort documents by expected relevance, most relevant first and last.
📋 Metadata Enrichment
Add metadata: [SOURCE: doc1.pdf] [DATE: 2024-01] [RELEVANCE: HIGH]
Example: Codebase Analysis Structure
=== TABLE OF CONTENTS ===
1. Architecture Overview (line 50)
2. API Endpoints (line 500)
3. Data Models (line 2000)
4. Tests (line 5000)
=== SECTION: Architecture Overview ===
[RELEVANCE: HIGH] [TYPE: documentation]
...
Progressive Summarization and Semantic Compression
Preserving information with fewer tokens
Even with large windows, context sometimes exceeds the limit. Compression techniques make it possible to retain essential information while reducing tokens.
Compression Techniques
🔄 Hierarchical Summarization
Summarize in levels: document → section → paragraph. Load the appropriate level.
🎯 Entity Extraction
Extract key entities and relationships; discard narrative text.
📊 Structured Conversion
Convert prose into JSON/tables — semantically denser.
🗜️ Progressive Compression
Older content = more compressed. Recent content = detailed.
Prioritization and Context Discarding
Deciding What to Keep
Prioritization Criteria
Semantic similarity score with the query
More recent information is generally more relevant
Explicit tags: [CRITICAL], [REQUIRED], [OPTIONAL]
Context referenced by other maintained
Long-Context Anti-Patterns
Common errors and how to avoid them
Common Anti-Patterns
❌ "More is Better" Fallacy
Add all available context without curation.
❌ Flat Context
No structure or hierarchy - everything at the same level.
❌ Instructions Burial
Put important instructions in the middle of the context.
❌ Stale Context
Keep outdated or contradictory context.