PTENES
MODULE 3.3

🧭 Code vs. documents: when to stop at Graphify

Not every graph needs to become a vault. Here you’ll learn the difference between extracting code and extract documents, and deciding—with judgment, not hype—when the graph is enough and when it’s worth taking it to Obsidian.

6
Topics
~35
Minutes
Advanced
Level
0%
0 of 6
1

🧱 Codebase graph: AST, tree-sitter, and no API key

When you point Graphify at a code repository, it doesn’t “read” the code as continuous text. It structural parsing: reads the language syntax and builds a tree of functions, classes, modules, and their calls. The result becomes the graph: nodes are functions and modules, and edges are "calls," "imports," and "defines."

🔰 New here? What are "AST" and "tree-sitter"

AST (Abstract Syntax Tree) is a tree representation of a program's structure: function X contains call Y, which receives argument Z. It's how the computer "understands" the code underneath, without ambiguity.

tree-sitter is the library Graphify uses to generate this AST quickly and reliably, in 12 languages. It's pure parsing — no AI model is called at this stage.

The practical consequence is important: extracting a code base is deterministic, fast, and no API key. There’s no possibility of “hallucination”—the tree reflects the file’s syntax exactly. You can run it on a huge project without spending a model token.

terminal — extract code (illustrative)
# extrai a AST do código em ./src — tree-sitter, sem API key
graphify extract ./src

# pergunta direto ao grafo de código já extraído
graphify query "quais funções chamam autenticar()?"

↑ graphify extract is the path headless (plain terminal). For code, it runs without any key—the tree-sitter does all the work.

✓ Strong at coding

  • ✓Exact structure: functions, classes, imports.
  • ✓Fast and cheap — no model call.
  • ✓Deterministic: zero hallucinations.

✗ Blind to meaning

  • ✗Doesn’t capture the intent behind the code.
  • ✗Ignores comments and explanatory prose.
  • ✗"Why does this exist?" isn't in the AST.

🔑 Key concepts

Code base
Code repository
AST
Syntax tree
tree-sitter
Parser, 12 languages
No key
Doesn’t call an LLM
2

📄 Document graph: LLM and semantics

Documents don't have program syntax — a docs page or a PDF is prose. So Graphify changes its strategy: instead of tree-sitter, it uses a LLM to read each document and understand what concepts it covers and how they relate to one another. It’s extraction semantics: by meaning, not by form.

🔰 New here? "LLM," "corpus," and "semantics"

LLM (Large Language Model) is the AI model that understands and generates text—the engine behind Claude. Here, it reads the documents and extracts the concepts.

Corpus is the set of documents you give it to process—your docs folder, PDFs, notes.

Semantics = relative to the meaning. Semantic extraction captures “this text is about authentication and links it to sessions,” even if the exact word doesn’t appear again.

Code → AST (tree-sitter) .py / .ts file parser module def login() def parse() class API deterministic · no API key Documents → concepts (LLM) docs / PDF (corpus) LLM reads concept authentication session token semantic · uses the model

↑ Same Graphify, two engines. Code becomes structure (AST, no key); documents become meaning (LLM, semantic). Inside the /graphify in Claude Code, the session already provides the model—you don’t need a key here either.

🔑 Key concepts

Corpus
Set of docs
LLM
Reads and understands
Semantics
By meaning
Concept
Idea node
3

🛑 When to stop at Graphify

There’s a temptation to move every graph into Obsidian “because you can.” But the graph is already useful alone. O graph.html (self-contained interactive visualization) opens in the browser without a server, and graphify query answers questions directly. To explore a specific codebase or get a quick answer, this enough.

💡 The golden rule

If the answer disappears when you close the terminal — and that's fine — stop at the graph. The vault only makes sense when the knowledge needs to persist e coexist with the rest of what you know.

  • •One-off question: graph + query, done.
  • •Audit a repo that isn’t yours: explore the graph.html, don’t export.
  • •With no intention of maintaining it: exporting would be overhead.
STOP AT THE GRAPH OR TAKE IT TO OBSIDIAN? Graph ready (graph.json) Just wants 1 answer / explore something specific? yes 🛑 Stop at Graphify graph.html + query no Will coexist with your notes/projects? yes 🔮 Bring it into Obsidian --obsidian → vault no ⏳ Stays in the graph for now — no export

↑ Arrows green = yes, red = no. The question is never “can it be exported?”—it’s “does this knowledge need to live outside the terminal?”

🔑 Key concepts

graph.html
Self-contained visual
query
Ask the graph
Silo
Isolated graph, okay
Overhead
Export cost
4

🔮 When to bring it into Obsidian

Obsidian comes in when knowledge stops being a one-off lookup and becomes collection. When you export, Graphify generates one markdown file per node with wikilinks [[nome-do-no]] and backlinks — then this graph can coexist with your other notes, projects, and ideas in a single navigable vault.

inside Claude Code — /graphify (illustrative)
# exporta o grafo como vault Obsidian (um .md por nó + backlinks)
/graphify ./claude-code-docs --obsidian --obsidian-dir ~/vault/graphify/claude-code

↑ The flag --obsidian only exists along the path skill (/graphify inside Claude Code), not in the graphify extract headless.

✓ Export is worthwhile when…

  • ✓The knowledge will be consulted again, many times.
  • ✓You want to connect it with your own notes (an integrated second brain).
  • ✓Want to navigate the graph visually in Obsidian.

✗ Not worthwhile when…

  • ✗It's a single question you won't revisit.
  • ✗The repo isn’t even yours—you’re just auditing it.
  • ✗You won't keep the vault up to date.

🔰 New here? What is a "vault"

Vault is Obsidian's "box": a folder of interconnected markdown files. Graphify's export fills a vault with one note per concept — and that's where the graph becomes persistent, navigable memory, not just a file in the terminal.

🔑 Key concepts

--obsidian
Export flag
Vault
Notes folder
Wikilink
[[backlink]]
Integrate
Work with notes
5

🔁 Keep code and docs in sync

Code and docs change. A graph frozen on the day of extraction grows stale and misleads. The solution isn't to re-extract everything from scratch every time — it's incremental update: Graphify stores a cache/ per file (by content hash) and only reprocesses what changed.

terminal — keep in sync (illustrative)
# reprocessa só os arquivos modificados desde a última vez
graphify update .

# rebuild automático: fica observando e atualiza ao salvar
graphify watch .

↑ graphify update . = on-demand update; graphify watch . = continuous. Inside Claude Code, the equivalent is /graphify . --update e /graphify . --watch.

1

The source changes

You edit the code or add a new doc to the project.

2

update reprocesses only the delta

The hash-based cache avoids calling the LLM again for untouched files. Cheap.

3

Re-export the vault (if you use Obsidian)

The export is regenerated from scratch — so keep your notes away from the imports folder.

⚠️ Watch out when re-exporting

The Obsidian export is regenerated every time. If you manually edited a generated note, that edit is lost on the next export. Keep imports and your notes in separate folders — module 2.5 covers this.

🔑 Key concepts

update
Only what changed
watch
Rebuild on save
cache
Hash per file
Incremental
Doesn’t redo everything
6

🧮 Cost, noise, and maintenance

All this structure comes at a price — and it’s real, not just hype. Choosing wisely between code and documents, stopping at the graph or going to the vault, is a account of three variables: extraction time/cost, graph noise, and maintenance effort. Weighing them first keeps you from becoming a collector of abandoned graphs.

💸

Extraction cost

Code (AST) is almost free. Documents call the LLM — slower and, outside Claude Code, uses tokens from your key. The cache/ softens on reruns.

📢

Graph noise

A huge corpus generates hundreds of nodes — not all of them useful. Use the GRAPH_REPORT.md (god nodes, communities) to see whether the signal is good before exporting.

🧹

Vault maintenance

A vault only helps if it stays current. If you’re not going to run update once in a while, it may be better to stop at the graph.

💡 The honest question

"Will this graph/vault give me more value than it costs to create and maintain?" If the answer is no, stop at Graphify — or don’t extract anything at all. Criteria beat hype every time.

🔑 Key concepts

Cost
Time + tokens
Noise
Useless nodes
GRAPH_REPORT
Audit the signal
Maintenance
Ongoing effort

✋ Self-recovery (optional, non-blocking): you ran Graphify on a repo that isn't yours, just to ask a quick question about how a function works. Where should you stop?

📌 Module summary

✓
Code becomes structure: AST via tree-sitter, deterministic and without a key.
✓
Documents become meaning: semantic extraction via LLM, over the corpus.
✓
Stop at the graph when the answer is standalone: graph.html + query are enough.
✓
Go to Obsidian when knowledge needs to persist and coexist with your notes.
✓
Keep it in sync and weigh the cost: update/watch + criteria beat hype.

Next module

3.4 · Maintenance and troubleshooting — keep your graph and vault healthy and resolve the most common errors.