PTENES
TRACK 3

🚀 Architecture and AI

Master data architecture, performance at scale, and AI applied to databases. From OLTP/OLAP to vector databases and RAG.

2
Modules
14
Topics
~2h
Duration
🚀
Advanced

Module Overview

Click a card to access the full module.

Detailed Content

Explore the topics in each module. Click to expand.

3.1 ~50 min

🏗️ Architecture and Performance

OLTP vs OLAP, partitioning, data warehouse, ETL, observability, and data governance.

What it is:

OLTP processes transactions in real time. OLAP analyzes large volumes of historical data.

Why learn:

Mixing OLTP and OLAP workloads in the same database degrades both. Separating them is engineering.

Key concepts:

Transactional vs. analytical, row-store vs. column-store, ETL, ELT

What it is:

Split large tables into smaller partitions based on a criterion (date, range, hash).

Why learn:

Tables with billions of rows become unmanageable. Partitioning improves queries and maintenance.

Key concepts:

Range partition, Hash partition, List partition, partition pruning

What it is:

Policies that define how long to keep data and how to archive old data.

Why learn:

Compliance (LGPD, GDPR), storage cost, performance. Data that lasts forever is a liability.

Key concepts:

Retention policy, cold storage, tiered storage, purge jobs

What it is:

A centralized repository optimized for analytical queries and BI.

Why learn:

Data spread across 10 systems does not produce insights. A DW consolidates it and enables analysis.

Key concepts:

Star schema, Snowflake schema, fact tables, dimension tables, BigQuery, Redshift

What it is:

Processes that Extract, Transform, and Load data between systems.

Why learn:

Raw data needs cleaning and transformation before it is useful.

Key concepts:

ETL vs. ELT, batch vs. streaming, Airflow, dbt, data quality checks

What it is:

Ability to understand the database's internal state through metrics, logs, and traces.

Why learn:

An unmonitored database is a time bomb. Incidents are inevitable; flying blind isn't.

Key concepts:

Metrics (latency, throughput, errors), logs, APM, alerts, SLOs, pg_stat_statements

What it is:

A framework of policies, processes, and responsibilities for data.

Why learn:

Without governance, data becomes garbage. Quality, security, and compliance depend on it.

Key concepts:

Data catalog, data lineage, RBAC, auditing, data classification

📄 View Full
3.2 ~70 min

🤖 AI Applied to Data

Context window, RAG, pgvector, Weaviate, Milvus, Qdrant, Redis Vector, Chroma, and Azure AI Search.

What it is:

The maximum number of tokens an LLM processes at once.

Why learn:

Exceeding the window loses information. Knowing how to size it is a requirement for using AI with data.

Key concepts:

Tokens, context window, compression, summary, relevance prioritization

What it is:

Technique that retrieves relevant data and injects it into the prompt before generating a response.

Why learn:

LLMs don't know your internal data. RAG enables chat over your knowledge base.

Key concepts:

Chunking, embeddings, vector indexing, reranking, RAG pipeline

What it is:

An extension that adds the VECTOR type to PostgreSQL for similarity search.

Why learn:

Uses the PostgreSQL you already have. No extra infrastructure for up to ~1M vectors.

Key concepts:

CREATE EXTENSION vector, VECTOR(1536), operators <=>, ivfflat, HNSW

What it is:

Vector database with hybrid search (semantic + keyword).

Why learn:

Purely semantic search misses exact terms. Hybrid search combines the best of both.

Key concepts:

nearVector, BM25, hybrid search, modules, multi-tenancy

What it is:

Open-source vector database designed for billions of vectors.

Why learn:

When pgvector no longer scales, Milvus is the most mature open-source option.

Key concepts:

Collection, FLOAT_VECTOR, IVF_FLAT, HNSW, GPU acceleration

What it is:

Qdrant stores vectors with filterable payloads. Redis FTS adds vectors to Redis.

Why learn:

Qdrant for complex filters + semantics. Redis for those who already use Redis and want vectors.

Key concepts:

Qdrant payload, cosine distance, Redis VECTOR HNSW, FT.SEARCH KNN

What it is:

Chroma is lightweight for local development. Azure AI Search is a managed enterprise solution.

Why learn:

Chroma for rapid prototyping. Azure for production with an SLA and Microsoft integration.

Key concepts:

chromadb.Client(), Azure vectorSearch, HNSW, skillset, reranking

📄 View Full