PTENES
MODULE 1.3

⚙️ Graphify under the hood

What the tool does, how it extracts knowledge, what it produces, and where it runs into limits. No jargon left unexplained: each term (CLI, skill, AST, tree-sitter, graph.json) is defined as it comes up.

6
Topics
~40
Minutes
Basic
Level
0%
0 of 6
1

⚙️ What Graphify is (CLI + skill)

O Graphify is the tool that builds the second brain. It has two sides that seem like the same thing but do different jobs: one CLI (a program you run in the terminal — the package is called graphifyy, with two "y"s) and a skill that it installs inside in Claude Code. Understanding this duality avoids 90% of the confusion in the next tracks.

🔰 New here? What is a "CLI"

CLI = Command Line Interface, or “command-line program.” It’s a program without windows: you type its name in the terminal and it does the work. Here the program is called graphify (the package you install is graphifyy).

🔰 New here? What is a "skill"

A skill is an instruction file (SKILL.md) that teaches Claude Code how to perform a task. Once installed, you call it with a slash command — a command that starts with “/”, such as /graphify. The skill is the “inside Claude Code” face; the pure CLT is the “terminal” face.

In practice, you install it once and get both sides. The installation command puts the CLI on your system; a second command registers the skill in Claude Code (it writes a SKILL.md in ~/.claude/skills/graphify/). Inside Claude Code, you doesn't need an API key: the session itself provides the model.

terminal — install (illustrative)
# 1) instala a CLI (recomendado, com uv)
uv tool install graphifyy

# 2) registra a skill /graphify dentro do Claude Code
graphify install

Once you’ve done that, you can trigger the skill in Claude Code by pointing it to a folder. The <isto-voce-troca> is the path to the source you want to map:

inside Claude Code — skill (illustrative)
/graphify <./pasta-da-fonte>
source code or docs graphify extracts the graph graphify-out/ graph.json source of truth graph.html GRAPH_REPORT.md

↑ One source goes in, the graphify processes, and everything lands in graphify-out/. O graph.json is the center: the visual and the report (topic 3) derive from it.

🔑 Key concepts

CLI
Terminal program
skill
SKILL.md in Claude
graphifyy
The package (2 "y")
/graphify
The slash command
2

🌳 How it extracts (AST vs. LLM)

Graphify uses two extraction engines, and which one runs depends on what you pointed to. If the source is code, it reads the AST via tree-sitter — fast, deterministic, and no API key. If the source is document (Markdown, PDF, image), it uses a LLM to understand the meaning from the text and extract the entities and relationships.

🔰 New here? What are "AST" and "tree-sitter"

AST = Abstract Syntax Tree, the code’s “syntax tree”: a structured representation that says “this is a function,” “this is a call,” “this imports that.” The tree-sitter is the library that reads the code and builds this tree — it understands 12 languages. Since it’s just structure (not interpretation), doesn't need an LLM or a key.

🔰 New here? Why a document needs an LLM

Prose doesn’t have a mechanical “tree” like code. To know that “context window” is related to “tokens,” you need to understand what's written — and that's the job of an LLM (the AI model). That's why document extraction is called semantics: it works by meaning, not by form.

Code → AST (no key) code tree-sitter AST · 12 languages graph fast structure · deterministic · no API key Documents → LLM (semantics) docs md · PDF · img LLMusing the meaning graph meaning · headless requires ANTHROPIC_API_KEY

↑ Two paths, the same destination (the graph). Code goes through form (AST, no key); documents go through the meaning (LLM). Inside Claude Code, the model already comes from the session— no key; only headless use in a plain terminal requires ANTHROPIC_API_KEY.

🔑 Key concepts

AST
Code tree
tree-sitter
Reads 12 languages
Semantics
LLM, by meaning
No key
Code in Claude
3

📂 The outputs: graph.json, graph.html, GRAPH_REPORT.md

When Graphify finishes, it doesn’t return “an answer” — it creates a folder call graphify-out/ with several files. Knowing which is which prevents the panic of “now what do I open?” on the first run. There are three main characters and one supporting character.

1

graph.json — the source of truth

The complete graph: all entities, relationships, and communities. Everything else derives from it. It’s the file Obsidian and Claude Code will consult.

2

graph.html — the interactive visual

A self-contained page: opens directly in the browser, no server. It's where you "see" the graph, drag nodes, and explore connections with your mouse.

3

GRAPH_REPORT.md — the text audit

A readable report: the god nodes (most connected nodes), the bridges between communities, and—the gold—a list of suggested questions to start exploring.

🔰 New here? What is "self-contained"

Self-contained = “everything in one file.” The graph.html already includes all the JavaScript and data it needs, so you can double-click it to open — no installation or server required.

graphify-out/ (illustrative)
graphify-out/
├── graph.json          # fonte de verdade — tudo deriva daqui
├── graph.html          # visual interativo (abre no navegador)
├── GRAPH_REPORT.md     # auditoria: god nodes + perguntas sugeridas
└── cache/              # cache por arquivo (não re-chama o LLM à toa)

The supporting player is the folder cache/: it stores the result per file (identified by the hash of the content) so you don’t pay for the LLM call again when nothing has changed. That’s why the second run is usually much faster than the first.

🔑 Key concepts

graph.json
Source of truth
graph.html
Visual in the browser
GRAPH_REPORT.md
Audit + questions
cache/
Doesn’t call the LLM again
4

🧱 Code vs. documents

Graphify points to two source types, and the choice changes everything that comes after. You can map a code base (a code repository) or a document corpus (a collection of PDFs, markdown files, notes). It’s not “better or worse”—they have different goals.

🔰 New here? "Code base" and "corpus"

A code base (or "codebase") is the set of program files in a project. A corpus is a collection of texts treated as a whole—here, the documentation you want to turn into a map. In the course, we use a corpus: the Claude Code official documentation.

🧩 Code base (code)

  • ✓Extraction by AST (tree-sitter), no key.
  • ✓Nodes like Function, Module.
  • ✓Good for understanding an architecture.

📄 Corpus (documents)

  • ✓Extraction semantics via LLM.
  • ✓Nodes like Concept (concepts).
  • ✓Good for understanding a subject or documentation.

To map a code repository in a plain terminal, the command is the CLI's headless interface. Change the <.> using the project path (the dot means “current folder”):

terminal — extract from code (illustrative)
# extrai o grafo do repositório na pasta atual
graphify extract <.>

# ou de uma pasta específica
graphify extract <./src>

Choosing between the two corpora also determines where the result goes: code usually produces a graph you query in Graphify itself (the subject of Track 3); documentation is what’s worth taking to Obsidian, where it fits into the project’s broader context. In this course, we follow the document path.

🔑 Key concepts

Code base
Repository
Corpus
Docs collection
Ontology
The node type
extract
Headless subcommand
5

📦 Obsidian mode and wiki mode

The raw graph (graph.json) is great for machines, but you still want a browse this knowledge. That's where two export modes come in—and they exist only in the skill (/graphify, inside Claude Code), not in the headless CLI. They are the --obsidian e o --wiki.

--obsidian

  • •A .md per node, with wikilinks [[nome]] for the neighbors.
  • •Generates a graph.canvas: the communities become groups in the Obsidian Canvas.
  • •It's the course's secret sauce: the graph becomes a vault browsable.

--wiki

  • •Articles by community, in Wikipedia style.
  • •Less atomic than Obsidian: groups related concepts in one text.
  • •Good for continuous reading; not for the backlink graph.

🔰 New here? "Wikilink" and "backlink"

A wikilink is a link in the format [[nome-da-nota]] that Obsidian understands. When note A links to note B, note B automatically gets a backlink ("who mentions me")—that's how the connections graph appears automatically in Obsidian.

The Obsidian mode command inside Claude Code points to the source and tells it where to save the vault. Change the source path and the --obsidian-dir using whichever destination you want:

inside Claude Code — export (illustrative)
# exporta um vault Obsidian (um .md por nó + canvas)
/graphify <./docs> --obsidian --obsidian-dir <~/vault/graphify/claude-code>

# ou: artigos por comunidade, estilo Wikipédia
/graphify <./docs> --wiki

🔑 Key concepts

--obsidian
1 md per node
--wiki
Article by community
graph.canvas
Communities in Canvas
backlink
"Who Mentions Me"
6

🫙 The graph in a vacuum: the limitation Obsidian solves

Here’s the honest truth of this module: on its own, the Graphify graph lives in a vacuum. It knows that corpus you provided inside out—and only it. It doesn't know about your other notes, your other projects, or the broader context where this knowledge should live. It's a silo: complete on the inside, isolated on the outside.

💡 The limitation, in one sentence

O graph.json saves the provenance (which source file each entity came from), but the source documents no are copied alongside the graph — they stay in the project. The graph points outside, but doesn’t bring the rest inside.

  • •Complete: knows everything about the corpus it received.
  • •Isolated: can't see anything beyond that corpus.
  • •Static: doesn't connect to your workflow on its own.

And that's exactly why the course doesn't stop at Graphify. Taking the graph to the Obsidian (the mode --obsidian from topic 5) is what takes knowledge out of a vacuum: inside the vault, it fits into broader context — alongside your other notes, linkable and integrable. The graph stops being an island and becomes a neighborhood in your second brain.

🔰 New here? What is a "silo"

In technology, silo is information kept in one place, closed off from everything else. The isolated graph is a knowledge silo; Obsidian opens the doors to that silo and connects it to what you already have.

🔑 Key concepts

Silo
Self-contained
Vacuum
No external context
Provenance
Where the node came from
Integration
What Obsidian provides

✋ Self-recovery (optional, non-blocking): of Graphify's three outputs, which is the source of truth from which everything else derives?

📌 Module summary

✓
Two sides: the CLI graphifyy in the terminal and the skill /graphify inside Claude Code (no key).
✓
Two engines: code goes through the AST/tree-sitter (no key); documents go through an LLM (semantic).
✓
Three outputs in graphify-out/: graph.json (truth), graph.html (visual), GRAPH_REPORT.md (audit).
✓
Navigable export: --obsidian (1 md per node + canvas) or --wiki (articles by community).
✓
The limitation: on its own, the graph is a silo; Obsidian takes it out of the void and fits it into the larger context.

Next module

1.4 · Obsidian as memory — why the vault is the right place for the graph to live and how the agent queries it.