PTENES
MODULE 2.2

🗜️ Lazy RAG: huge PDF → cheat sheet

McKinsey, BCG, and KPMG playbooks are 28 to 46 pages long—they blow past any context limit and are packed with buzzwords. The solution is simple and low-effort: split them into chunks, summarize each one, and put together a one-page cheat sheet.

6
Topics
~40
Minutes
Interm.
Level
Technical
Type
1

📚 The Context Problem

You want the elite knowledge in these PDFs. But dropping a document from 46 pages whole in Claude Code exceeds the context limit—and worse: it fills the window with summaries, thank-yous, and buzzwords. You need the valuable, not from the volume.

✗ Upload the entire PDF

  • ✗Exceeds the token limit
  • ✗Wastes expensive context on junk
  • ✗Dilutes the signal with noise
  • ✗Generic, superficial response

✓ Distill before using

  • ✓Fits comfortably in the context
  • ✓Only frameworks, numbers, and insights
  • ✓Pure signal, high density
  • ✓Reusable in any project
28–46 pages

Huge PDFs

Limit

of context

Signal vs. noise

value vs. volume

Density

what matters

2

✂️ Chunking into blocks of ~10k tokens

The core idea behind lazy RAG: instead of processing everything at once, you slice. Extracts the text from the PDF, counts the tokens, and splits it into chunks of ~10,000 tokens. Each chunk fits comfortably in an API call.

PDF extract raw text counts tokens cuts block 1 · ~10k tk block 2 · ~10k tk block 3 · rest each block becomes 1 API call processed in parallel →

// the concept of chunking — splitting by tokens

CHUNK_SIZE = 10_000   # tokens por bloco

texto  = extrair_texto_do_pdf("bcg-guide.pdf")   # via PyMuPDF
tokens = tokenizer.encode(texto)                 # cl100k_base

blocos = []
for i in range(0, len(tokens), CHUNK_SIZE):
    pedaco = tokens[i : i + CHUNK_SIZE]
    blocos.append(tokenizer.decode(pedaco))

# 1 PDF de 46 págs ≈ 3-4 blocos → cabem na API um a um
10k tokens

by block

PyMuPDF

extracts text

Tokenizer

cl100k_base

In parts

not all at once

3

🧪 The synthesis prompt

This is where the quality of the cheat sheet comes from. The prompt instructs the model to extract only what's worthwhile — frameworks, statistics, insights, citations, terms — preserving numbers and cutting buzzwords. A block with no value gets the label [MINIMAL_CONTENT] and is discarded afterward.

// actual structure of the synthesis prompt (pdf_to_cheatsheet.py)

Você é um consultor criando um cheat sheet a partir
deste documento. Extraia o conhecimento mais valioso:

## Conceitos & Frameworks   → liste cada um em 1-2 frases
## Estatísticas críticas     → preserve números e fontes
## Insights acionáveis       → o que dá para implementar
## Citações notáveis         → 2-3 frases de impacto
## Terminologia              → defina termos e siglas

Regras: seja conciso, NÃO generalize números, use bullets,
negrito no termo na 1ª vez. Se o bloco for só sumário/
referências, responda: [MINIMAL_CONTENT: ...]

💡 The detail that makes a difference

The rule "DO NOT generalize numbers" is what separates a useful cheat sheet from an empty summary. “AI boosts productivity” is worthless; "40% productivity gain (Google Cloud, 2025)" is ammunition for your deliverable.

5 sections

fixed extraction

Preserve #

doesn't generalize

Cuts the buzzwords

only what's valuable

[MINIMAL]

mark the gap

4

✨ Gemini 2.5 Flash: affordable and massive context window

The model that synthesizes. A deliberate choice: huge context, rock-bottom price. It makes it possible to distill an entire library of PDFs for pennies — with exponential retry and a delay between calls to respect the rate limit.

1

Huge context

Handles 10k-token blocks with room to spare — and responds without cutting off halfway.

2

Very low cost

The entire guide library costs pennies. It’s the same model that writes the deliverables in the Factory.

3

Resilient

Exponential retry (3 attempts) and ~5s delay between calls. A network error won’t bring down the process.

💡 Practical tip

This was a real error during development (in the video, Gemini's chunking failed on a file). Building in public means showing the mess: if an API error occurs, just ask Claude Code to run it again—the process resumes where it left off.

2.5 Flash

the model

Context

huge

Cents

low cost

Retry

exponential

5

🧩 Build the cheat sheet

With the blocks synthesized, you rebuilds: combines them in the right order, discards those marked as empty, and generates a single Markdown file—with a header (source, date), numbered sections, and a "verify against the original" footer.

// reassembling the blocks into a final cheat sheet

cheat = cabecalho(fonte="bcg-guide.pdf", data=hoje)

for i, bloco in ordenar_por_indice(blocos_sintetizados):
    if "[MINIMAL_CONTENT" in bloco and tamanho_util(bloco) < 500:
        continue                      # descarta o vazio
    cheat += f"## Seção {i+1}\n\n{bloco}\n"

cheat += rodape("verifique no documento original")
salvar("bcg-guide_TLDR.md", cheat)    # 1 página, pronto

✓ Well-structured cheat sheet

  • ✓Blocks in the original order
  • ✓Empty entries discarded
  • ✓Header with source and date
  • ✓Footer requesting verification

✗ Poorly structured cheat sheet

  • ✗Blocks out of order
  • ✗Summary and acknowledgments in the middle
  • ✗No source trail
  • ✗Treated as absolute truth
Sort

by index

Discard

the empty ones

Header

source + date

1 page

Final Markdown

6

🐍 pdf_to_cheatsheet.py in practice

You don’t write any of this by hand — the script already exists. It reads a folder of PDFs, saves progress to progress.json, skips blocks already completed, and accepts useful flags. Just ask Claude Code to run it.

// what you type to Claude Code

# dá uma olhada sem gastar API (só conta os blocos)
python pdf_to_cheatsheet.py --dry-run

# processa só um arquivo
python pdf_to_cheatsheet.py --single-pdf "bcg-guide.pdf"

# se o contexto encher ou a API falhar, rode de novo:
# ele retoma pelos blocos que faltam (progress.json)

💡 Why resuming matters

O ProgressTracker saves each completed block. If the process fails at block 7 of 10, when you run it again, it skips the first 6 and continues. You doesn't pay twice through the same block — savings that add up in a large library.

--dry-run

without spending on API calls

--single-pdf

a single file

Resumable

resumes where you left off

progress.json

saves what was done

✅ Module summary

✓
The entire PDF overloads the context — you want what’s valuable, not volume.
✓
Chunking + synthesis is lazy RAG — blocks of ~10k tokens, a prompt that extracts only the gold.
✓
Gemini 2.5 Flash does it for pennies — huge context, retry, resume.
✓
The result is one reusable page — ready to feed into Claude Code.

🎯 Mission 2.2 — Your first cheat sheet

Get 1 consulting PDF (a free guide from McKinsey, BCG, KPMG, Google Cloud…) and turn it into a 1-page cheat sheet:

  1. Ask Claude Code to run the pdf_to_cheatsheet.py in it (or apply the block-by-block synthesis prompt).
  2. Check: did frameworks, numbers, and insights remain—without buzzwords?
  3. Save as fonte_TLDR.md with a header (source + date).

Success: 1 generated 1-page cheat sheet. What you gained: the first piece of your knowledge base—and the technique for distilling any PDF from now on.

Next module:

2.3 — Your consultant brain (organize multiple cheat sheets in a curated, injectable knowledge base)