📥 Choose and prepare the source
Before running any command, you decide what Graphify will read code or documents. Then it downloads the corpus, cleans the folder, and gets everything ready for extraction. Clean source, clean graph.
🧭 Code or documents: what to point to
Graphify reads a source and extracts the graph from it. This source can be one of two types: a code base (a code repository) or a document corpus (PDFs, markdown, texts). The module's first decision is which of the two you’ll point to it — because that changes how Graphify works under the hood.
🔰 New here? "Code base" and "corpus"
A code base is simply the folder containing a project’s code (files .py, .ts, etc.). A corpus is a set of documents about a subject—the textual "raw material." Graphify accepts both, but handles each differently.
The difference is technical, and it matters: code is extracted by the AST (the syntax tree that the tree-sitter builds from the code)—it’s deterministic, fast, and doesn't need an API key. Already documents go through a LLM, which reads the meaning of the text and infers entities and relationships. Deciding early prevents rework later on.
↑ The two paths—code (extracted by AST, without a key) and documents (read by a LLM)—they end up in the same source folder clean. In this course, we follow the documents’ path, but the principle applies to both.
💻 Code path (AST)
- ✓12 languages via tree-sitter.
- ✓Deterministic and no API key.
- ✓Nodes are functions, modules, classes.
📄 Document path (LLM)
- ✓PDF, markdown, text—any document.
- ✓Reads the meaning, not just the syntax.
- ✓Inside Claude Code, without your own key.
🔑 Key concepts
⬇️ Download the Claude Code documentation
To get a real, familiar corpus, the easiest way is to ask Claude Code to download its official documentation to a local folder — for example ./claude-code-docs. You don't need a script: describe the task in plain language, and the agent does the tedious work of saving each file one by one.
Objective: download the official Claude Code documentation into a local source folder, ready to become a graph.
Baixe a documentação oficial do Claude Code numa pasta chamada ./claude-code-docs no diretório atual. - Salve uma página por arquivo .md, com nome legível. - Mantenha só o conteúdo (sem menus/rodapé do site). - Ao terminar, me diga quantos arquivos foram salvos.
How to verify: count the downloaded files.
ls ~/projetos/segundo-cerebro/claude-code-docs | wc -l
Replace the path with the location where you created the folder. The number of files is up to you, not a constant — it depends on how much documentation there was that day.
Having a known corpus helps you compare your results with the reference. In one run shown in the video, the Claude Code docs yielded 145 documents that became 591 nodes, 685 connections e 67 communities. Use this only as a rough guide ("in the video run, it was…") — your numbers will vary with the documentation version and the scope you choose.
🔰 New here? Why ask the agent to download?
Claude Code already knows how to navigate and save pages. Instead of hunting down each URL yourself, it browses the documentation and saves the .md for you. It’s the same agent that will run Graphify later, so it makes sense for it to build the source now.
🔑 Key concepts
🗂️ Organize the source folder
A source folder is exactly what Graphify will scan — so it deserves a dedicated root with no clutter. Gather the files that describe the knowledge in one place, separate from the rest of the system. The golden rule: what goes into the folder becomes a node in the graph.
Objective: create a dedicated source folder inside the course project.
mkdir -p ~/projetos/segundo-cerebro/claude-code-docs
How to verify: the folder exists and you can list its contents (empty at first, then with your .md).
ls ~/projetos/segundo-cerebro/claude-code-docs | wc -l
Suggested path: ~/projetos/segundo-cerebro/<sua-pasta-fonte>. Use any name you like — just keep everything from the same source together.
🔰 New here? What does the mkdir -p
mkdir creates a folder; the flag -p also creates any missing parent folders and doesn't complain if the folder already exists. It's safe to run again.
🔑 Key concepts
📏 Size and scope (start small)
The temptation is to point it at the entire corpus at once. Resist. Start with one subset — a dozen representative files—and only expand the scope once the entire workflow works end to end. Smaller extractions are faster and cheaper to test and adjust.
🔰 New here? What does "iterate" mean
Iterate is repeating a short cycle—run, check the result, adjust—instead of trying to get everything right in one go. With a small source, each round takes seconds; across the entire corpus, it takes minutes (and, on the LLM path, tokens).
✓ Start small
- ✓Runs in seconds: fast feedback.
- ✓Cheap to repeat until you get the setting right.
- ✓Errors show up early and are easy to fix.
✗ Entire corpus right away
- ✗Long wait before seeing anything.
- ✗If the graph isn't useful, you've already spent everything.
- ✗Errors only show up at the end and are costly to redo.
Point to a subset
About 10–15 files that represent the subject well. Run the entire workflow through to the graph.
Check the result
Does the graph make sense? Do the nodes and connections line up? Adjust the source if needed.
Expand the scope
Once the workflow is reliable, point it at the entire corpus at once.
🔑 Key concepts
🧹 Cleanup: what to delete
Each additional file in the source folder is potential noise in the graph. Remove binaries, builds, dependencies, duplicates, and anything that doesn’t carry meaning. The goal is to improve the signal-to-noise ratio: keep only the files that describe knowledge, nothing that's a machine byproduct.
🔰 New here? "Signal/noise" and .gitignore
Signal is the useful content; noise is everything that distracts from it. A .gitignore is a list of file patterns to ignore—here, we use the same idea to decide what no goes into the source.
✓ Keep (signal)
- ✓Markdown, text, reference docs.
- ✓Source code (if the source is code).
- ✓What describes concepts and decisions.
✗ Exclude (noise)
- ✗Binaries, images,
node_modules, builds. - ✗Logs, caches, temporary files, drafts.
- ✗Duplicate copies of the same content.
# what does NOT become a source node_modules/ dist/ build/ *.png *.jpg *.zip *.log .cache/ **/tmp/
🔑 Key concepts
📁 Where to run it (graphify-out goes here)
Graphify saves the output in the folder graphify-out/ inside the current directory (o cwd, from current working directory). That’s why, where you run defines where the files appear. Run it from the project root and you'll always know where to find the graph.json later.
🔰 New here? What is "cwd"
O cwd is the folder your terminal is currently in — what pwd shows. Commands with relative paths (like graphify-out/) are resolved from it.
↑ A project root is your cwd: both the source folder (input) and graphify-out/ (generated output). A predictable path to find the graph.json in the next module.
🔑 Key concepts
✋ Self-recovery (optional, non-blocking): you are going to point Graphify at a code repository. How does it extract the graph?
📌 Module summary
Next module
2.3 · Run Graphify and read the graph — generate the knowledge graph and learn to read what it produced, from the visualization to the report.