PTENES
MODULE 2.2

🗑️🔥 Substrate & Context — the flaming garbage truck

The second foundation layer: raw domain knowledge. You start with a dumpster truck on fire — 500 to 5,000 chaotic files. The work is organize before connecting AI, gather what's missing and distill everything into nuggets. Inject only the nuggets.

8
Topics
~55
Minutes
Intermediate.
Level
Practical
Type
0%
0 of 0 topics read · Section 1 of 8

Detailed content

1

🚛 The Flaming Garbage Truck Analogy

O Substrate is your domain’s raw knowledge—and you almost always start with what the author calls "a dumpster fire". Especially in a business, they're 500, 5,000 files of various types and structures: PDFs, spreadsheets, zips, even unusual files like Parquet. Chaos. The Substrate is the layer that turns this chaos into something the OS can use.

🌱 New here?

Substrate = the reference material stored outside the CLAUDE.md, which the OS consults only when needed. Context = what actually goes into the conversation at that moment. The difference is the key to the module: you store a lot (substrate) but inject a little (context), so you don't exceed the context window — the amount of text the model reads at once.

🔥 What it is

The image of the burning garbage truck helps ease anxiety: it’s normal to start out messy. The job isn’t to “have perfect data”; it’s to give chaos shape — and that is design, not heavy engineering.

  • •Lots of files, lots of formats, no order: the usual starting point.
  • •The Substrate is where this material becomes a reference you can consult.

Why learn

Because most people try to connect AI on top of the chaos and gets frustrated with bad answers. Accepting the flaming dumpster truck as a starting point—and treating it methodically—is what separates an OS that knows your domain from one that guesses. In context-heavy domains (Tax OS ~34%), this layer is the most important of all.

Key concepts

Substrate
saved reference
Context
what enters the conversation
500–5,000 files
the usual chaos
Design > engineering
give chaos shape
2

🗂️ Organize BEFORE connecting AI

Before asking Claude Code for anything, do a mini data engineering: organize. The author describes two simple cuts — first by file type (PDFs with PDFs, spreadsheets with spreadsheets), then by age (last 6 months vs. last 2 years). This isolates parts of the material and gives the OS clean ground to work with.

💡 AI can help you organize

"Organize first" doesn’t mean doing everything by hand: you can ask Claude Code to move each file to the folder for its type and then branch by age. The point is that this organizing happens before of any analysis—you prepare the ground, then use it.

The cleanup flow in 3 cuts

1

By type

Each file goes in the folder for its format: pdf/, spreadsheets/, transcricoes/.

2

By age

Within each type, separate recent (the last 6 months) from old (up to 2 years). Recent items usually matter more.

3

Isolate & name

With everything separated, it’s obvious what’s a source and what’s noise — and the OS reads only the right part.

Why learn

Because the quality of the OS is limited by the quality of the Substrate. Dumping 5.000 disorganized files into AI produces vague and expensive answers (lots of tokens). Organizing first is the low-cost step that multiplies everything that follows.

Key concepts

Mini data engineering
organizing is design
By type
1st cut
By age
recent vs. old
Before AI
sets the stage
3

🌾 Harvesting: what the specialist still needs

After organizing things comes the harvesting (harvest). The guiding question is: beyond what I already have, what's missing to create the “expert of my dreams” in this domain? You’ll gather knowledge from where it lives — YouTube, Instagram, experts, laws, lawyers — and bring it into Substrato.

🔥raw chaos organizeby type + age harvestingcollect what's missing ⚙️ workhorsedistills into nuggets 📁 raw/stored 📁 synthesized/only nuggets ↗ context

How to read: the chaos (orange) is organized, then gathered from external sources, then a workhorse model distills everything. Two folders remain: raw/ (kept as a backup, gray) and synthesized/ (the nuggets, blue)—only this one enters the context.

Why learn

Because what you already have is rarely enough. A true specialist knows things that aren’t in your files—they’re in experts’ heads and public sources. Harvesting deliberately closes that gap instead of hoping the model “knows.”

Key concepts

Harvesting
collect what's missing
Dream expert
the domain target
External sources
YouTube, experts, laws
Intentional gap
what's deliberately missing
4

🧠 Tacit knowledge: what generic solutions lack

The target of harvesting is the tacit knowledge: what a generic model no knows why it isn't well documented—it's in the know-how of the people doing the work. In Freedom OS, for example, these are details like which translation agencies are reliable, what needs to be notarized (officially signed) and which documents to translate into Arabic or French depending on the country.

🌱 New here?

Tacit knowledge is the opposite of "explicit": experiential knowledge, rarely written down—shortcuts, exceptions, "what always goes wrong." A generic model was trained on public text, so it has explicit knowledge but not the tacit knowledge of the your case. The Substrate exists to capture precisely that tacit knowledge.

🔎 Where the tacit hides

  • People — experts, lawyers, accountants who answer “it depends” and explain why.
  • Videos & posts — YouTube and Instagram creators who live and breathe the domain, full of unofficial tips and tricks.
  • Your own cases — the emails, forms, and repeated questions you’ve already answered.

Why learn

Because it’s exactly the tacit knowledge that makes the OS feel like a expert and not a chatbot. Without it, the OS repeats what any model would say. Capturing tacit knowledge is the difference between “generic” and “no one does it like my OS”—the test that comes back in the Trilha 4 audit.

Key concepts

Tacit
know-how
Explicit
what the generic version already has
Exceptions
"what goes wrong"
Non-generic
the OS’s differentiator
5

📚 Playbooks & compendiums

Scattered collected material still isn't useful. The next step is to aggregate sources into a compendium (a compendium.md) — a distilled reference document — and in playbooks: playbooks for “how to do X in this domain.” It’s what the OS consults to respond with an expert’s depth.

📘 Compendium

Distilled domain reference: laws, facts, definitions, “how it works.”

  • ✓Answers “what’s true here?”
  • ✓Brings many sources together in one place.

📗 Playbook

Execution plan: steps for carrying out recurring tasks in the domain.

  • ✓Answers “how do you do X?”
  • ✓Becomes raw material for Skills (Track 3).

💡 The Substrate done check

The layer is ready when there is a substrate/compendium.md organized that answers the goal question that you defined for the domain. If the question still has no good answer, there's more to gather or distill.

Why learn

Because compendiums and playbooks are the format that makes knowledge usable. They’re also the bridge to the next track: a good playbook is almost a Skill waiting to be packaged.

Key concepts

Compendium
distilled reference
Playbook
how to do X
Goal question
the done-check
Bridge to Skills
playbook → verb
6

🧊 Raw vs. synthesized — inject only nuggets

The Substrate’s core pattern is separation raw vs synthesized. You keep a folder raw/ with the raw transcripts and documents (kept just in case) and a folder adjacent and complementary with the distilled versions — the nuggets. The golden rule: inject only the nuggets in context.

🌱 New here?

One nugget is a highly distilled summary — think 10 bullets that capture the essentials of a one-hour transcript. Inject into context = put that text into the conversation for the model to use. You want to inject nuggets (small, dense ones), not raw material (huge, sparse), so the OS stays fast and cheap.

📁 raw/ — the pantry

  • •Raw transcripts and documents, just as they came in.
  • •Stored just in case; almost never enters the context.
  • •Your safety net if you need to re-distill.

📁 synthesized/ — the finished dish

  • ✓Nuggets: the essentials in a few bullets.
  • ✓It’s what actually gets injected into the context.
  • ✓Small, dense, cheap in tokens.

Why learn

Because it’s the pattern that repeats in every domain — pantry vs. finished dish. Confusing the two (injecting the raw material) blows up the context window, raises costs, and worsens the answers. Separating raw from synthesized content is what keeps the OS lean and the knowledge retrievable.

Key concepts

raw/
raw, saved
synthesized/
distilled nuggets
Inject only nuggets
the golden rule
Pantry vs. plate
the analogy
7

⚙️ Workhorse model in a synthesis loop

The person doing the distillation doesn’t need to use the most expensive model. The author uses a workhorse model ("workhorse" model)—cheap and efficient, like Gemini Flash — running in a loop for all files: "use this model, go through each file, and create a synthesis of the central nuggets from these transcripts".

🌱 New here?

One workhorse model is a cheap model chosen for repetitive, high-volume tasks — not for difficult reasoning, but to “sweep through” lots of files. In a loop = repeat the same operation file by file. You spend little per item and process hundreds of sources without burning expensive tokens.

🔁 The "knowledge harvester" skill

This loop can become a reusable skill: it asks which handles scrape and how many videos, saves everything to raw/ and runs the workhorse to generate the synthesized version. It becomes a system self-improving with a human in the loop: you keep adding more, distilling, and still back it up (a private copy on GitHub, with .gitignore of what’s sensitive).

Why learn

Because distilling by hand doesn’t scale, and using a top-of-the-line model for it is wasteful. The “cheap workhorse + loop” pair makes synthesis viable at scale—and when you generalize the script, you get a Substrate engine that runs every time new material arrives.

Key concepts

Workhorse
cheap model
Gemini Flash
author's example
In a loop
file by file
Human in the loop
guided self-improvement
8

⚡ Practical: scan raw/ and generate nuggets

Let’s build the synthesis engine. The copy-run prompt below tells Claude Code to run a workhorse model in a loop about everything in raw/ and write one nugget per file in synthesized/. Replace only the sections between <brackets>.

🎯 Objective

Transform the folder raw/ (raw) in a folder synthesized/ of ready-to-inject nuggets—using a cheap model without burning expensive tokens.

Paste into Claude Code — copy-run prompt

prompt
Você é meu pipeline de Substrato. Use o modelo workhorse
<modelo barato: ex. gemini-flash> para destilar meu material bruto.

Passos:
1. Liste todos os arquivos em ./raw/ (transcrições, PDFs já em texto,
   posts). NÃO modifique nada em raw/ — é a minha despensa.
2. Para CADA arquivo, chame o modelo workhorse e gere um nugget:
   - no máximo <N: ex. 10> bullets com o essencial;
   - capture o conhecimento tácito (exceções, "o que dá errado",
     atalhos), não só o resumo óbvio;
   - foque no que responde à minha pergunta-objetivo:
     "<sua pergunta-objetivo do domínio>".
3. Salve cada nugget em ./synthesized/ com o MESMO nome do
   arquivo de origem + sufixo ".nugget.md".
4. Ao final, gere ./synthesized/compendium.md agregando os
   nuggets por tema, e me diga quantos arquivos processou.

Regra: só os arquivos de synthesized/ entram no contexto depois.
Nunca injete o conteúdo de raw/ diretamente.

✔️ How to verify

  • 1.The folder raw/ stayed intact (same file count as before).
  • 2.There is a .nugget.md in synthesized/ for each file of raw/.
  • 3.Each nugget fits in a few bullets and includes at least 1 tacit detail (not just the obvious).
  • 4.O compendium.md answers your goal question (Substrato's done-check).

Why learn

Because this is the operational heart of the Substrate layer. Running it once gives you the synthesized/; generalizing the prompt gives you the “knowledge harvester” skill that you reuse whenever new material arrives — the self-improving system with you in the loop.

Key concepts

raw/ intact
never modify
1 nugget / file
workhorse loop
compendium.md
groups by topic
<brackets>
what you exchange

✅ Module summary

✓
You start with a dumpster fire — 500–5,000 chaotic files; that’s normal.
✓
Organize before you turn on AI — by type and age; mini data engineering.
✓
Harvesting captures tacit knowledge — what the generic model lacks, drawn from experts and sources.
✓
Compendiums & playbooks — the knowledge becomes a searchable reference.
✓
raw vs synthesized; workhorse in a loop — distill cheaply, inject only the nuggets.

Next module:

2.3 — Rules & Hooks: the fences 🚧