Context Window Management
Optimizing the use of available context
The context window is the token limit a model can process in a single interaction. Managing this resource efficiently is crucial for getting accurate and relevant responses, especially in complex applications.
Context Limits by Model
Optimization Strategies
Summarize previous conversations while preserving essential information
Include only context relevant to the current task
Keep the N most recent messages and discard the older ones
Use vector databases to store and retrieve context
Introduction to RAG
Retrieval-Augmented Generation
RAG (Retrieval-Augmented Generation) is an architecture that combines the generative capabilities of LLMs with searches across external knowledge bases. This lets the model answer using up-to-date information specific to your domain.
RAG Pipeline
Advantages
- • Up-to-date knowledge
- • Reduces hallucinations
- • Verifiable citations
- • Specific domain
Challenges
- • Retrieval quality
- • Additional latency
- • Knowledge base maintenance
- • Embedding cost
Embeddings and Vector Search
Semantic representation of text
Embeddings are numerical representations (vectors) of text that capture its semantic meaning. Texts with similar meanings will have nearby vectors in vector space, enabling similarity searches.
How It Works
# Original Text
"The cat slept on the couch"
↓ Embedding Model
# Vector (simplified)
[0.23, -0.45, 0.67, 0.12, -0.89, ...]
* Real vectors have 768-3072 dimensions
Popular Embedding Models
Pinecone
Managed Vector DB
Weaviate
Open source
Chroma
Lightweight and simple
Chunking Strategies
Intelligently splitting documents
Chunking is the process of dividing large documents into smaller segments for indexing. The chunking strategy directly impacts retrieval quality in RAG.
Chunking Strategies
📏 By Fixed Size
Split into chunks of N tokens with overlap
📝 By Structure
Preserves the document structure (paragraphs, sections, markdown)
🧠 Semantic
Uses embeddings to identify topic changes
🔄 Recursive
Try larger separators first, then smaller ones
Pro Tip
There is no perfect chunk size. Try values between 256-1024 tokens and empirically evaluate what works best for your specific use case.
Prompt Evaluation
Metrics and testing methodologies
Evaluating the quality of prompts and AI systems is essential for continuous improvement. A good evaluation methodology combines automated metrics with human review.
Evaluation Framework
📊 Automated Metrics
- Relevance: Cosine similarity with gold standard
- Accuracy: % of correct answers
- Recall: % of information retrieved
- Latency: Response time
👥 Human Evaluation
- Usefulness: Does the answer help the user?
- Clarity: Is it easy to understand?
- Factuality: Correct information?
- Completeness: Does it cover all the points?
Testing Methodology
Create a Test Dataset
Collect representative examples with expected answers
Define Metrics
Choose metrics aligned with your business goals
Test Variations
Compare different prompt versions (A/B testing)
Analyze and Iterate
Use insights to improve continuously
Prompts in Production
Deployment, monitoring, and scalability
Taking AI systems to production requires attention to aspects such as versioning, monitoring, security, costs, and scalability. A well-executed deployment ensures reliability and enables continuous improvement.
Production Checklist
Versioning
- • Prompt version control
- • Change history
- • Easy rollback
Monitoring
- • Request logs
- • Quality metrics
- • Anomaly alerts
Security
- • Rate limiting
- • Input sanitization
- • Data protection
Costs
- • Token tracking
- • Prompt optimization
- • Response caching
Reference Architecture
┌─────────────────────────────────────────┐
│ USERS │
└──────────────────┬──────────────────────┘
│
▼
┌─────────────────────────────────────────┐
│ API Gateway (Rate Limit, Auth) │
└──────────────────┬──────────────────────┘
│
▼
┌─────────────────────────────────────────┐
│ Prompt Service (Templates, Cache) │
└───────┬─────────────────────┬───────────┘
│ │
▼ ▼
┌───────────────┐ ┌───────────────────┐
│ Vector DB │ │ LLM API │
│ (RAG) │ │ (OpenAI, etc) │
└───────────────┘ └───────────────────┘
Capstone Project
Applying all the concepts from the technical level
The final project is an opportunity to demonstrate mastery of the techniques you’ve learned. You’ll build a complete Q&A system with RAG over a knowledge base of your choice.
Project Specification
🎯 Objective
Create a specialized Q&A assistant that answers questions about a specific domain (technical documentation, product manual, legal database, etc.)
📋 Requirements
- • Complete RAG pipeline (ingestion, chunking, embedding, retrieval)
- • Professional system prompt with persona and guardrails
- • Structured response formatting
- • Source citations in responses
- • Handling out-of-scope questions
📦 Deliverables
- 1. Architecture Documentation
- 2. Complete, Annotated System Prompt
- 3. RAG pipeline code/configuration
- 4. Test dataset with 20+ questions
- 5. Evaluation Report with Metrics
Evaluation Rubric
🎓 Congratulations!
By completing this project, you will have demonstrated technical-level prompt engineering skills. You are ready for the Masterclass Level!