🧭 Code vs. documents: when to stop at Graphify
Not every graph needs to become a vault. Here you’ll learn the difference between extracting code and extract documents, and deciding—with judgment, not hype—when the graph is enough and when it’s worth taking it to Obsidian.
🧱 Codebase graph: AST, tree-sitter, and no API key
When you point Graphify at a code repository, it doesn’t “read” the code as continuous text. It structural parsing: reads the language syntax and builds a tree of functions, classes, modules, and their calls. The result becomes the graph: nodes are functions and modules, and edges are "calls," "imports," and "defines."
🔰 New here? What are "AST" and "tree-sitter"
AST (Abstract Syntax Tree) is a tree representation of a program's structure: function X contains call Y, which receives argument Z. It's how the computer "understands" the code underneath, without ambiguity.
tree-sitter is the library Graphify uses to generate this AST quickly and reliably, in 12 languages. It's pure parsing — no AI model is called at this stage.
The practical consequence is important: extracting a code base is deterministic, fast, and no API key. There’s no possibility of “hallucination”—the tree reflects the file’s syntax exactly. You can run it on a huge project without spending a model token.
# extrai a AST do código em ./src — tree-sitter, sem API key graphify extract ./src # pergunta direto ao grafo de código já extraído graphify query "quais funções chamam autenticar()?"
↑ graphify extract is the path headless (plain terminal). For code, it runs without any key—the tree-sitter does all the work.
✓ Strong at coding
- ✓Exact structure: functions, classes, imports.
- ✓Fast and cheap — no model call.
- ✓Deterministic: zero hallucinations.
✗ Blind to meaning
- ✗Doesn’t capture the intent behind the code.
- ✗Ignores comments and explanatory prose.
- ✗"Why does this exist?" isn't in the AST.
🔑 Key concepts
📄 Document graph: LLM and semantics
Documents don't have program syntax — a docs page or a PDF is prose. So Graphify changes its strategy: instead of tree-sitter, it uses a LLM to read each document and understand what concepts it covers and how they relate to one another. It’s extraction semantics: by meaning, not by form.
🔰 New here? "LLM," "corpus," and "semantics"
LLM (Large Language Model) is the AI model that understands and generates text—the engine behind Claude. Here, it reads the documents and extracts the concepts.
Corpus is the set of documents you give it to process—your docs folder, PDFs, notes.
Semantics = relative to the meaning. Semantic extraction captures “this text is about authentication and links it to sessions,” even if the exact word doesn’t appear again.
↑ Same Graphify, two engines. Code becomes structure (AST, no key); documents become meaning (LLM, semantic). Inside the /graphify in Claude Code, the session already provides the model—you don’t need a key here either.
🔑 Key concepts
🛑 When to stop at Graphify
There’s a temptation to move every graph into Obsidian “because you can.” But the graph is already useful alone. O graph.html (self-contained interactive visualization) opens in the browser without a server, and graphify query answers questions directly. To explore a specific codebase or get a quick answer, this enough.
💡 The golden rule
If the answer disappears when you close the terminal — and that's fine — stop at the graph. The vault only makes sense when the knowledge needs to persist e coexist with the rest of what you know.
- •One-off question: graph + query, done.
- •Audit a repo that isn’t yours: explore the
graph.html, don’t export. - •With no intention of maintaining it: exporting would be overhead.
↑ Arrows green = yes, red = no. The question is never “can it be exported?”—it’s “does this knowledge need to live outside the terminal?”
🔑 Key concepts
🔮 When to bring it into Obsidian
Obsidian comes in when knowledge stops being a one-off lookup and becomes collection. When you export, Graphify generates one markdown file per node with wikilinks [[nome-do-no]] and backlinks — then this graph can coexist with your other notes, projects, and ideas in a single navigable vault.
# exporta o grafo como vault Obsidian (um .md por nó + backlinks) /graphify ./claude-code-docs --obsidian --obsidian-dir ~/vault/graphify/claude-code
↑ The flag --obsidian only exists along the path skill (/graphify inside Claude Code), not in the graphify extract headless.
✓ Export is worthwhile when…
- ✓The knowledge will be consulted again, many times.
- ✓You want to connect it with your own notes (an integrated second brain).
- ✓Want to navigate the graph visually in Obsidian.
✗ Not worthwhile when…
- ✗It's a single question you won't revisit.
- ✗The repo isn’t even yours—you’re just auditing it.
- ✗You won't keep the vault up to date.
🔰 New here? What is a "vault"
Vault is Obsidian's "box": a folder of interconnected markdown files. Graphify's export fills a vault with one note per concept — and that's where the graph becomes persistent, navigable memory, not just a file in the terminal.
🔑 Key concepts
🔁 Keep code and docs in sync
Code and docs change. A graph frozen on the day of extraction grows stale and misleads. The solution isn't to re-extract everything from scratch every time — it's incremental update: Graphify stores a cache/ per file (by content hash) and only reprocesses what changed.
# reprocessa só os arquivos modificados desde a última vez graphify update . # rebuild automático: fica observando e atualiza ao salvar graphify watch .
↑ graphify update . = on-demand update; graphify watch . = continuous. Inside Claude Code, the equivalent is /graphify . --update e /graphify . --watch.
The source changes
You edit the code or add a new doc to the project.
update reprocesses only the delta
The hash-based cache avoids calling the LLM again for untouched files. Cheap.
Re-export the vault (if you use Obsidian)
The export is regenerated from scratch — so keep your notes away from the imports folder.
⚠️ Watch out when re-exporting
The Obsidian export is regenerated every time. If you manually edited a generated note, that edit is lost on the next export. Keep imports and your notes in separate folders — module 2.5 covers this.
🔑 Key concepts
🧮 Cost, noise, and maintenance
All this structure comes at a price — and it’s real, not just hype. Choosing wisely between code and documents, stopping at the graph or going to the vault, is a account of three variables: extraction time/cost, graph noise, and maintenance effort. Weighing them first keeps you from becoming a collector of abandoned graphs.
Extraction cost
Code (AST) is almost free. Documents call the LLM — slower and, outside Claude Code, uses tokens from your key. The cache/ softens on reruns.
Graph noise
A huge corpus generates hundreds of nodes — not all of them useful. Use the GRAPH_REPORT.md (god nodes, communities) to see whether the signal is good before exporting.
Vault maintenance
A vault only helps if it stays current. If you’re not going to run update once in a while, it may be better to stop at the graph.
💡 The honest question
"Will this graph/vault give me more value than it costs to create and maintain?" If the answer is no, stop at Graphify — or don’t extract anything at all. Criteria beat hype every time.
🔑 Key concepts
✋ Self-recovery (optional, non-blocking): you ran Graphify on a repo that isn't yours, just to ask a quick question about how a function works. Where should you stop?
📌 Module summary
Next module
3.4 · Maintenance and troubleshooting — keep your graph and vault healthy and resolve the most common errors.