🧠 Context Window
A context window (context window) is the maximum number of tokens an LLM can process at once. Everything in the prompt—instructions, data, history—competes for that space. Exceeding the window means lose information or truncate critical data.
🎯 Concept: Tokens and Limits
Modern models range from 8K to 2M context tokens, but more context doesn’t mean a better result. The model loses accuracy in the "middle" of very large windows (the "lost in the middle" problem). The strategy is to place only what's relevant.
- • Token: Minimum unit of text. ~1 token = ~4 characters in English, ~3 in Portuguese
- • Compression: Summarize long documents before sending them to the model
- • Hierarchical summarization: Summarize chunks, then summarize the summaries
- • Relevance-based prioritization: Use semantic search to select the most important passages
💡 Practical Tip
For data applications, the rule is: don’t put the entire database in the prompt. Use RAG to retrieve only relevant passages, compress long documents, and reserve at least 30% of the window for the model's response. Cost scales with tokens (input + output). 5 well-selected chunks outperform 50 full pages.
🔎 RAG - Retrieval-Augmented Generation
RAG is the technique that connects LLMs to your data. Instead of relying only on the model’s internal knowledge, RAG retrieves relevant information in your database and injects it into the prompt before generating a response. That’s how chatbots “know” a company’s internal data.
🔄 Complete RAG Pipeline
1. Chunking
Split documents into chunks of ~500-1000 tokens with overlap
2. Embeddings
Convert each chunk into a numeric vector that captures semantic meaning
3. Indexing
Store vectors in a vector database with an index for fast search
4. Query
Convert the user's question into a vector and search for similar chunks
5. Reranking
Reorder results with a cross-encoder model for greater accuracy
6. Generation
Build a prompt with the retrieved context and generate a response with the LLM
# Pipeline RAG completo em Python
docs = split_into_chunks(load_corpus(), tokens=600)
vecs = embed(docs)
index = build_vector_index(vecs)
query = embed("Como normalizar tabelas?")
tops = index.search(query, k=5)
prompt = compose(system, user, context=tops)
answer = llm.generate(prompt)
📊 Reference Numbers
- Typical chunk size: 500-1000 tokens with 10-20% overlap
- Popular embeddings: text-embedding-3-small (1536 dims), BGE, E5
- Typical top-k: 3-10 chunks depending on the size of the window
- Reranking improves accuracy in 10-25% vs. pure vector search
🐘 pgvector in PostgreSQL
pgvector adds the type VECTOR to PostgreSQL, enabling similarity search directly in the relational database you already know. No extra infrastructure, no new database. Ideal for up to ~1 million vectors with good performance.
-- Habilitar a extensao
CREATE EXTENSION IF NOT EXISTS vector;
-- Criar tabela com coluna vetorial
CREATE TABLE doc_chunk (
id BIGSERIAL PRIMARY KEY,
content TEXT NOT NULL,
embedding VECTOR(1536)
);
-- Buscar os 5 chunks mais similares (distancia cosseno)
SELECT id, content
FROM doc_chunk
ORDER BY embedding <=> '[0.010, -0.020, ...]'::vector
LIMIT 5;
💡 Performance: ivfflat and HNSW
Without an index, pgvector performs an exact search (seq scan). For large tables, create an approximate index:
- • ivfflat: Divides vectors into
lists(clusters). In a search, query onlyprobesnearby clusters. Rule:lists = sqrt(n),probes = sqrt(lists). - • HNSW: A navigable hierarchical graph. Faster than ivfflat for searching, slower to build. Best for production with high QPS.
- • Start with ivfflat and migrate to HNSW when you need faster search.
🎯 Distance Operators
<=> Cosine
Most widely used. Compares vector direction, ignoring magnitude.
<-> L2 (Euclidean)
Geometric distance. Sensitive to vector magnitude.
<#> Inner Product
Maximizes similarity. Used with normalized vectors.
🌊 Weaviate
Weaviate and a vector database with native hybrid search: it combines semantic search (vectors) with keyword search (BM25). Purely semantic search can miss exact terms. Hybrid search solves this by combining both rankings.
import weaviate
client = weaviate.Client("http://localhost:8080")
# Busca semantica (nearVector)
result = client.query.get("Document", ["content", "source"]) \
.with_near_vector({"vector": query_vec}) \
.with_limit(5) \
.do()
# Busca hibrida (vetor + BM25)
result = client.query.get("Document", ["content", "source"]) \
.with_hybrid(query="normalizacao tabelas", alpha=0.5) \
.with_limit(5) \
.do()
# alpha=1.0 = puramente vetorial
# alpha=0.0 = puramente BM25
# alpha=0.5 = metade de cada
🔍 Weaviate’s strengths
- • Hybrid search: The alpha parameter controls the balance between semantics and BM25
- • Modules: Integrated modules for automatic embedding (text2vec-openai, text2vec-cohere)
- • Multi-tenancy: Isolates data by tenant without duplicating infrastructure. Good for SaaS
- • GraphQL API: Expressive, well-documented query interface
⚡ Milvus
Milvus and an open-source vector database designed to billions of vectors. When pgvector no longer scales, Milvus is the most mature option. It supports GPU acceleration, distributed sharding, and multiple index types.
from pymilvus import (
connections, FieldSchema, CollectionSchema,
DataType, Collection
)
# Conectar
connections.connect("default", host="localhost", port="19530")
# Definir schema
fields = [
FieldSchema(name="id", dtype=DataType.INT64,
is_primary=True, auto_id=True),
FieldSchema(name="content", dtype=DataType.VARCHAR,
max_length=2000),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR,
dim=1536)
]
schema = CollectionSchema(fields, description="RAG docs")
collection = Collection("documents", schema)
# Criar indice IVF_FLAT
index_params = {
"metric_type": "COSINE",
"index_type": "IVF_FLAT",
"params": {"nlist": 256}
}
collection.create_index("embedding", index_params)
collection.load()
# Buscar
results = collection.search(
data=[query_vector],
anns_field="embedding",
param={"metric_type": "COSINE", "params": {"nprobe": 16}},
limit=5,
output_fields=["content"]
)
🎯 When Milvus Shines
Scale
Billions of vectors distributed across a cluster. Automatic sharding, read replicas.
GPU Acceleration
Native GPU support for indexing and search. Orders of magnitude faster.
Zilliz Cloud
Managed (SaaS) version of Milvus. No ops, scales on demand.
🎯 Qdrant and Redis Vector
Qdrant stores vectors with filterable payloads — ideal for semantic searches with complex filters. Redis with the RediSearch module adds vector search to the cache many already use in production. Two approaches, two different scenarios.
# Qdrant com qdrant-client
from qdrant_client import QdrantClient
from qdrant_client.models import (
VectorParams, Distance, PointStruct
)
client = QdrantClient("localhost", port=6333)
# Criar collection com distancia cosseno
client.create_collection(
collection_name="documents",
vectors_config=VectorParams(
size=1536,
distance=Distance.COSINE
)
)
# Inserir com payload (metadados filtraveis)
client.upsert(
collection_name="documents",
points=[
PointStruct(
id=1,
vector=[0.012, -0.034, ...],
payload={
"content": "Normalizacao ate 3FN",
"source": "aula-sql",
"level": "intermediario"
}
)
]
)
# Buscar com filtro de payload
results = client.search(
collection_name="documents",
query_vector=query_vec,
query_filter={"must": [
{"key": "level", "match": {"value": "intermediario"}}
]},
limit=5
)
# Redis Vetorial com RediSearch
# Criar indice com campo vetorial HNSW
FT.CREATE idx_docs ON HASH PREFIX 1 doc:
SCHEMA
content TEXT
embedding VECTOR HNSW 6
TYPE FLOAT32
DIM 1536
DISTANCE_METRIC COSINE
# Inserir documento
HSET doc:1
content "Normalizacao ate 3FN reduz redundancia"
embedding "\x3f\x80\x00\x00..."
# Buscar KNN (5 vizinhos mais proximos)
FT.SEARCH idx_docs
"*=>[KNN 5 @embedding $vec AS score]"
PARAMS 2 vec "\x3f\x80\x00\x00..."
SORTBY score
RETURN 2 content score
DIALECT 2
Qdrant - Ideal for
- • Semantic search + complex payload filters
- • Native multi-tenancy with isolation
- • Written in Rust, low latency
- • Native REST API and gRPC
Redis Vector - Ideal for
- • Already using Redis? Add vectors without new infrastructure
- • Cache + vector search in the same instance
- • Ultra-low latency (in-memory)
- • Native TTL for temporary vectors
🧪 Chroma and Azure AI Search
Chroma is the lightest vector database for local development and rapid prototyping. Azure AI Search is Microsoft’s managed enterprise solution with an SLA, Azure OpenAI integration, and hybrid search. Dev vs Enterprise.
# Chroma - banco vetorial mais simples
import chromadb
client = chromadb.Client() # In-memory, zero config
# Criar collection
collection = client.get_or_create_collection("my_docs")
# Adicionar documentos (embedding automatico)
collection.add(
documents=[
"Normalizacao ate 3FN reduz redundancia",
"Indices B-tree aceleram buscas por range",
"Transacoes ACID garantem consistencia"
],
ids=["doc1", "doc2", "doc3"],
metadatas=[
{"source": "sql-101"},
{"source": "indices"},
{"source": "transacoes"}
]
)
# Buscar por similaridade
results = collection.query(
query_texts=["como reduzir redundancia?"],
n_results=3
)
print(results["documents"])
// Azure AI Search - Indice com busca vetorial (JSON)
{
"name": "rag-index",
"fields": [
{"name": "id", "type": "Edm.String", "key": true},
{"name": "content", "type": "Edm.String",
"searchable": true},
{"name": "contentVector",
"type": "Collection(Edm.Single)",
"searchable": true,
"dimensions": 1536,
"vectorSearchProfile": "my-vector-profile"}
],
"vectorSearch": {
"algorithms": [
{"name": "my-hnsw", "kind": "hnsw",
"hnswParameters": {
"metric": "cosine", "m": 4,
"efConstruction": 400, "efSearch": 500
}}
],
"profiles": [
{"name": "my-vector-profile",
"algorithm": "my-hnsw"}
]
}
}
🎯 When to Use Each One
Chroma
Prototyping, hackathons, local development, small projects. Zero configuration; runs in memory or persists to SQLite. Not recommended for high-load production.
Azure AI Search
Enterprise production, 99.95% SLA, Azure OpenAI integration, skillsets for automatic enrichment, native semantic reranking.