Module Overview
Click a card to access the full module.
Detailed Content
Explore the topics in each module. Click to expand.
🏗️ Architecture and Performance
OLTP vs OLAP, partitioning, data warehouse, ETL, observability, and data governance.
OLTP processes transactions in real time. OLAP analyzes large volumes of historical data.
Mixing OLTP and OLAP workloads in the same database degrades both. Separating them is engineering.
Transactional vs. analytical, row-store vs. column-store, ETL, ELT
Split large tables into smaller partitions based on a criterion (date, range, hash).
Tables with billions of rows become unmanageable. Partitioning improves queries and maintenance.
Range partition, Hash partition, List partition, partition pruning
Policies that define how long to keep data and how to archive old data.
Compliance (LGPD, GDPR), storage cost, performance. Data that lasts forever is a liability.
Retention policy, cold storage, tiered storage, purge jobs
A centralized repository optimized for analytical queries and BI.
Data spread across 10 systems does not produce insights. A DW consolidates it and enables analysis.
Star schema, Snowflake schema, fact tables, dimension tables, BigQuery, Redshift
Processes that Extract, Transform, and Load data between systems.
Raw data needs cleaning and transformation before it is useful.
ETL vs. ELT, batch vs. streaming, Airflow, dbt, data quality checks
Ability to understand the database's internal state through metrics, logs, and traces.
An unmonitored database is a time bomb. Incidents are inevitable; flying blind isn't.
Metrics (latency, throughput, errors), logs, APM, alerts, SLOs, pg_stat_statements
A framework of policies, processes, and responsibilities for data.
Without governance, data becomes garbage. Quality, security, and compliance depend on it.
Data catalog, data lineage, RBAC, auditing, data classification
🤖 AI Applied to Data
Context window, RAG, pgvector, Weaviate, Milvus, Qdrant, Redis Vector, Chroma, and Azure AI Search.
The maximum number of tokens an LLM processes at once.
Exceeding the window loses information. Knowing how to size it is a requirement for using AI with data.
Tokens, context window, compression, summary, relevance prioritization
Technique that retrieves relevant data and injects it into the prompt before generating a response.
LLMs don't know your internal data. RAG enables chat over your knowledge base.
Chunking, embeddings, vector indexing, reranking, RAG pipeline
An extension that adds the VECTOR type to PostgreSQL for similarity search.
Uses the PostgreSQL you already have. No extra infrastructure for up to ~1M vectors.
CREATE EXTENSION vector, VECTOR(1536), operators <=>, ivfflat, HNSW
Vector database with hybrid search (semantic + keyword).
Purely semantic search misses exact terms. Hybrid search combines the best of both.
nearVector, BM25, hybrid search, modules, multi-tenancy
Open-source vector database designed for billions of vectors.
When pgvector no longer scales, Milvus is the most mature open-source option.
Collection, FLOAT_VECTOR, IVF_FLAT, HNSW, GPU acceleration
Qdrant stores vectors with filterable payloads. Redis FTS adds vectors to Redis.
Qdrant for complex filters + semantics. Redis for those who already use Redis and want vectors.
Qdrant payload, cosine distance, Redis VECTOR HNSW, FT.SEARCH KNN
Chroma is lightweight for local development. Azure AI Search is a managed enterprise solution.
Chroma for rapid prototyping. Azure for production with an SLA and Microsoft integration.
chromadb.Client(), Azure vectorSearch, HNSW, skillset, reranking