PTENES
MODULE 2.2

📥 Choose and prepare the source

Before running any command, you decide what Graphify will read code or documents. Then it downloads the corpus, cleans the folder, and gets everything ready for extraction. Clean source, clean graph.

6
Topics
~35
Minutes
Practical
Level
0%
0 of 6
1

🧭 Code or documents: what to point to

Graphify reads a source and extracts the graph from it. This source can be one of two types: a code base (a code repository) or a document corpus (PDFs, markdown, texts). The module's first decision is which of the two you’ll point to it — because that changes how Graphify works under the hood.

🔰 New here? "Code base" and "corpus"

A code base is simply the folder containing a project’s code (files .py, .ts, etc.). A corpus is a set of documents about a subject—the textual "raw material." Graphify accepts both, but handles each differently.

The difference is technical, and it matters: code is extracted by the AST (the syntax tree that the tree-sitter builds from the code)—it’s deterministic, fast, and doesn't need an API key. Already documents go through a LLM, which reads the meaning of the text and infers entities and relationships. Deciding early prevents rework later on.

your source what should you point to? code or docs code → AST tree-sitter · no key required documents → LLM semantic reading source folder clean and focused two paths, one destination: the folder Graphify scans

↑ The two paths—code (extracted by AST, without a key) and documents (read by a LLM)—they end up in the same source folder clean. In this course, we follow the documents’ path, but the principle applies to both.

💻 Code path (AST)

  • ✓12 languages via tree-sitter.
  • ✓Deterministic and no API key.
  • ✓Nodes are functions, modules, classes.

📄 Document path (LLM)

  • ✓PDF, markdown, text—any document.
  • ✓Reads the meaning, not just the syntax.
  • ✓Inside Claude Code, without your own key.

🔑 Key concepts

code base
Code repository
corpus
Set of docs
AST
Code, no key
LLM
Semantic docs
2

⬇️ Download the Claude Code documentation

To get a real, familiar corpus, the easiest way is to ask Claude Code to download its official documentation to a local folder — for example ./claude-code-docs. You don't need a script: describe the task in plain language, and the agent does the tedious work of saving each file one by one.

copy-run · prompt to paste into Claude Code

Objective: download the official Claude Code documentation into a local source folder, ready to become a graph.

Baixe a documentação oficial do Claude Code numa pasta
chamada ./claude-code-docs no diretório atual.

- Salve uma página por arquivo .md, com nome legível.
- Mantenha só o conteúdo (sem menus/rodapé do site).
- Ao terminar, me diga quantos arquivos foram salvos.

How to verify: count the downloaded files.

ls ~/projetos/segundo-cerebro/claude-code-docs | wc -l

Replace the path with the location where you created the folder. The number of files is up to you, not a constant — it depends on how much documentation there was that day.

Having a known corpus helps you compare your results with the reference. In one run shown in the video, the Claude Code docs yielded 145 documents that became 591 nodes, 685 connections e 67 communities. Use this only as a rough guide ("in the video run, it was…") — your numbers will vary with the documentation version and the scope you choose.

🔰 New here? Why ask the agent to download?

Claude Code already knows how to navigate and save pages. Instead of hunting down each URL yourself, it browses the documentation and saves the .md for you. It’s the same agent that will run Graphify later, so it makes sense for it to build the source now.

🔑 Key concepts

prompt
Text task
real corpus
The official documentation
145 docs
Video example
wc -l
Check the count
3

🗂️ Organize the source folder

A source folder is exactly what Graphify will scan — so it deserves a dedicated root with no clutter. Gather the files that describe the knowledge in one place, separate from the rest of the system. The golden rule: what goes into the folder becomes a node in the graph.

copy-run · create the dedicated folder

Objective: create a dedicated source folder inside the course project.

mkdir -p ~/projetos/segundo-cerebro/claude-code-docs

How to verify: the folder exists and you can list its contents (empty at first, then with your .md).

ls ~/projetos/segundo-cerebro/claude-code-docs | wc -l

Suggested path: ~/projetos/segundo-cerebro/<sua-pasta-fonte>. Use any name you like — just keep everything from the same source together.

🔰 New here? What does the mkdir -p

mkdir creates a folder; the flag -p also creates any missing parent folders and doesn't complain if the folder already exists. It's safe to run again.

🔑 Key concepts

dedicated folder
A single root
structure
Organized
noise
What to avoid
goes in → node
File becomes a graph
4

📏 Size and scope (start small)

The temptation is to point it at the entire corpus at once. Resist. Start with one subset — a dozen representative files—and only expand the scope once the entire workflow works end to end. Smaller extractions are faster and cheaper to test and adjust.

🔰 New here? What does "iterate" mean

Iterate is repeating a short cycle—run, check the result, adjust—instead of trying to get everything right in one go. With a small source, each round takes seconds; across the entire corpus, it takes minutes (and, on the LLM path, tokens).

✓ Start small

  • ✓Runs in seconds: fast feedback.
  • ✓Cheap to repeat until you get the setting right.
  • ✓Errors show up early and are easy to fix.

✗ Entire corpus right away

  • ✗Long wait before seeing anything.
  • ✗If the graph isn't useful, you've already spent everything.
  • ✗Errors only show up at the end and are costly to redo.
1

Point to a subset

About 10–15 files that represent the subject well. Run the entire workflow through to the graph.

2

Check the result

Does the graph make sense? Do the nodes and connections line up? Adjust the source if needed.

3

Expand the scope

Once the workflow is reliable, point it at the entire corpus at once.

🔑 Key concepts

scope
How much to include
subset
Initial sample
iterate
Short cycle
cost
Time and tokens
5

🧹 Cleanup: what to delete

Each additional file in the source folder is potential noise in the graph. Remove binaries, builds, dependencies, duplicates, and anything that doesn’t carry meaning. The goal is to improve the signal-to-noise ratio: keep only the files that describe knowledge, nothing that's a machine byproduct.

🔰 New here? "Signal/noise" and .gitignore

Signal is the useful content; noise is everything that distracts from it. A .gitignore is a list of file patterns to ignore—here, we use the same idea to decide what no goes into the source.

✓ Keep (signal)

  • ✓Markdown, text, reference docs.
  • ✓Source code (if the source is code).
  • ✓What describes concepts and decisions.

✗ Exclude (noise)

  • ✗Binaries, images, node_modules, builds.
  • ✗Logs, caches, temporary files, drafts.
  • ✗Duplicate copies of the same content.
exclusions (illustrative example, .gitignore style)
# what does NOT become a source
node_modules/
dist/
build/
*.png
*.jpg
*.zip
*.log
.cache/
**/tmp/

🔑 Key concepts

exclusions
What to remove
.gitignore
Patterns to ignore
signal/noise
Useful vs. distraction
less is more
Cleaner graph
6

📁 Where to run it (graphify-out goes here)

Graphify saves the output in the folder graphify-out/ inside the current directory (o cwd, from current working directory). That’s why, where you run defines where the files appear. Run it from the project root and you'll always know where to find the graph.json later.

🔰 New here? What is "cwd"

O cwd is the folder your terminal is currently in — what pwd shows. Commands with relative paths (like graphify-out/) are resolved from it.

project root — you run Graphify from here (cwd) claude-code-docs/ source folder — what goes in graphify graphify-out/ graph.json · graph.html — comes out here

↑ A project root is your cwd: both the source folder (input) and graphify-out/ (generated output). A predictable path to find the graph.json in the next module.

🔑 Key concepts

cwd
Current directory
graphify-out/
Output folder
project root
Where to run it
graph.json
The key output

✋ Self-recovery (optional, non-blocking): you are going to point Graphify at a code repository. How does it extract the graph?

📌 Module summary

✓
Code or documents: code goes through the AST (no key); documents go through an LLM.
✓
Download a real corpus: ask Claude Code for the official docs in a folder like ./claude-code-docs.
✓
A clean, dedicated folder: what goes into the folder becomes a node — so only include what has meaning.
✓
Start small: start with a subset, iterate, then expand the scope.
✓
Run at the root: the output goes to graphify-out/ in the cwd — a predictable path.

Next module

2.3 · Run Graphify and read the graph — generate the knowledge graph and learn to read what it produced, from the visualization to the report.