PTENES
MODULE 3-6

🧠 Brains — the engine that thinks and the organized memory

"Brain" has two meanings when it comes to Jarvis, and both matter. The first is the LLM engine — the model that thinks, which you can swap without changing the code. The second is the organized memory — the “second brain” divided into zones. In this module, the last one in Anatomy, you’ll understand both—and why there is NO “magic model council” voting.

6
Topics
~50
Minutes
Intermediate
Level
Practical
Type
1

⚙️ The brain = the LLM engine

Throughout Anatomy, you built channels (how it communicates), identity (who it is), tools (its hands), skills (the recipes), and agents (the loop that works on its own). What’s missing is the piece that actually thinks: o brain. In the first sense, "brain" is the LLM engine — the language model that receives the context, reasons, and returns text. Everything you built only comes to life when this engine is plugged in.

The big idea in this module: this engine is a pluggable piece. It isn’t welded to the rest of Jarvis. The agent code talks to the engine through an interface standardized (a “socket”), so you can unplug one model and plug in another without rewriting anything. That’s why the phrase that sums it all up is: "one brain, many mouths" — and, here, multiple possible brains, a single architecture.

🧩 The engine under the hood

Think of Jarvis as a car you've already built: the body (channels), the dashboard (identity), the wheels (tools). The engine and is the LLM. The entire car depends on it—but the engine is just one more part, and parts can be swapped.

  • •The engine receives: context (your message + SOUL.md + memory + tool results).
  • •The engine returns text — a response, or a tool request for the agentic loop.
  • •The rest of Jarvis doesn’t "know" which engine is running—it just talks to the outlet.

New here? LLM ("Large Language Model") is a program trained on a lot of text that predicts the next word — it’s the brain running behind ChatGPT, Claude, or local models. Provider and it’s who provides this brain: Anthropic, OpenAI, OpenRouter (which brings several together), or your own PC via Ollama. Interface here is the default "fit" through which the code talks to any of them in the same way.

Key concepts

LLM engine

The model that actually thinks: it receives context and returns text.

Plug-in component

The brain isn’t soldered in: you plug it in and unplug it.

One brain, many mouths

The same model serves Telegram, voice, and web — different voices, one brain.

Provider

Who provides the engine: Anthropic, OpenRouter, local Ollama...

2

🔌 Switch brains without changing the code

If the engine is pluggable, switching brains should be as simple as changing a light bulb—and it is. In a well-built Jarvis, the model doesn’t appear throughout the code: it’s declared in a single place, the file .env. Change two lines there, restart, and the entire Jarvis starts thinking with a different model. This is called hot-swap (hot swap).

Why does this matter so much? Because every task has an ideal engine. A simple question (“what time is it in Tokyo?”) doesn’t need the most expensive model in the world; a tough reasoning problem deserves a frontier model. Switching brains becomes a economic decision — you choose power or cost based on the task, without being held hostage by a single provider.

📊 Three engines, three profiles

  • •Powerful cloud (Anthropic Claude, GPT): better reasoning, paid for token (each little piece of text that comes in or goes out — you pay by volume). For difficult tasks.
  • •Low-cost cloud (via OpenRouter): access to dozens of inexpensive models with a single key. For medium-sized tasks.
  • •Local (Ollama on your PC): $0 per use, private, offline. Models by RAM (3B@8GB, 8B@16GB); CPU is slow (30–60s).
Copy-run · change the brain in .env font-mono

Objective: make Jarvis think with a different model without touching the code — only by editing the file .env in the project root.

1) Powerful cloud model (Anthropic). Paste it into your .env:

LLM_PROVIDER=anthropic
LLM_MODEL=claude-sonnet-4-5
ANTHROPIC_API_KEY=<sua-chave-anthropic>

2) Want cheap cloud access via OpenRouter? Change only these lines:

LLM_PROVIDER=openrouter
LLM_MODEL=<modelo-barato-ex-meta-llama/llama-3.1-8b-instruct>
OPENROUTER_API_KEY=<sua-chave-openrouter>

3) Want 100% local and free? Start Ollama and point to it:

LLM_PROVIDER=ollama
LLM_MODEL=<modelo-local-ex-llama3.2>
OLLAMA_HOST=http://localhost:11434

How to check: restart Jarvis and send in the channel: "which model are you using right now?". If the code logs the provider at startup, you’ll see something like [brain] provider=ollama model=llama3.2. Behavior changes (speed, quality), but you haven’t touched a single line of code — only in the .env. That’s hot-swapping in action.

Key concepts

Hot-swap

Change the engine by changing the .env and restarting — without touching the code.

.env

Configuration/secrets file; the only place where the model is declared.

Economic decision

Inexpensive model for simple tasks, premium model for difficult ones.

OpenRouter / Ollama

A cloud model aggregator; a model runtime on your PC.

3

🪧 The myth of the "model council"

When someone hears “multiple brains,” their imagination runs wild: five models sitting around a table, debating, voting, and a judge choosing the best answer. It’s a nice image—and in the real projects in this course, it doesn't exist. There’s no magic router that consults N models at the same time and averages the results. Be honest about it: that’s marketing, not architecture.

⚠️ Be careful with the "advice" promise

Running several frontier LLMs in parallel for every message is expensive, slow, and rarely improves the response reliably. Anyone selling you an "AI committee vote" is usually selling complexity you don't need — and can't audit.

This course’s thesis is the opposite: less is more. Code you understand is worth more than advice you’ll never be able to debug.

So, when projects talk about “multiple models,” what does that really mean? Two concrete and honest things: replacement (the hot-swap from topic 2—replacing one with another), and portability — the same skill running in different runtimes (Claude Code, Codex) via the Agent Skills spec and tools like the polyskill. It’s not simultaneous voting; it’s the same source running on more than one brain.

✓ Honest “multiple models”

  • ✓Replacement: switch the engine via .env when it makes sense.
  • ✓Portability: the same skill runs in Claude Code AND Codex.
  • ✓Subagent delegates to its own model and returns only the answer (T3-5).
  • ✓You understands and audits each piece.

✗ The advice myth

  • ✗5 models voting on each message.
  • ✗A magical “judge” that chooses the best answer.
  • ✗Cost and latency multiply without guaranteed benefits.
  • ✗A black box you can't debug.

New here? Runtime and the "environment" where the AI runs (Claude Code, Codex, Cursor...). Portability and the same recipe works in more than one runtime without rewriting it. polyskill and a tool that packages a skill into a single source and makes it run in different runtimes—the real "multiple models."

Key concepts

Model Council (myth)

The idea of N LLMs voting live—the one real projects don’t use.

Replacement

One engine at a time, switched based on the task.

Portability (cross-runtime)

The same skill in Claude Code, Codex, etc. (polyskill).

Less is more

Auditable code beats impressive complexity.

4

📒 The "second brain" becomes memory

Now for the second meaning of “brain.” You’ve already seen in Identity (T3-2) that the conversation forgets when closing — the context window is like RAM; it empties. For Jarvis to truly remember you, it needs persistent memory: notes that persist between sessions. This collection is what we call second brain — borrowing the term from personal productivity (“Building a Second Brain” by Tiago Forte): an external system where you store what doesn’t fit in your head.

But here's the danger. If you simply throw EVERYTHING into a pile of notes—decisions, facts, feelings, reminders—the second brain grows and turns into a messy attic. When Jarvis needs to find something, it finds noise. The solution isn't to have less memory; it's to organize it into zones, each with a purpose. That's where the idea of the "3 brains" (next topic) comes from.

1

Without memory

Every conversation starts from scratch. It treats you like a stranger every day.

2

Memory as one big pile

It remembers—but all in one bucket. It finds noise, mixes facts with feelings, and gets bloated.

3

Memory organized into zones

Everything in its place (Project / Self / Knowledge). It finds things quickly and reasons better.

New here? Persistent memory and the one that survives when you close the conversation—saved in files on disk, not in the context window. Second brain and it’s the name of this external collection of notes. Inbox (inbox) is where every new note lands first, before being filed in the right zone.

Key concepts

Second brain

External collection of notes that persists across sessions.

Persistent memory

What stays on disk, unlike the context window, which empties.

Single-file clutter

Putting everything in the same bucket creates noise; the problem isn't quantity, it's order.

Organize into zones

Separating by purpose gives you better order and makes searching easier.

5

🧠 The 3 brains (memory)

Here’s the heart of the second meaning of "brains." Instead of a pile of notes, Jarvis’s memory is divided into three zones, each with a different nature. Think of three drawers in a filing cabinet: what changes all the time doesn't live with what almost never changes.

BRAIN 1 · the LLM engine (swappable) anthropic openrouter ollama (local) .env agent the outlet BRAIN 2 · the memory (3 zones) 📥 single inbox/triagem routes Projectepisodicwhat happened,decisions Selfidentityvalues, questionschanges slowly Knowledgereferencefacts that onlyaccumulate [[wikilinks]] connect the three zones

On the left, the LLM engine and it’s interchangeable: three providers plug into the agent’s same "socket" via .env. On the right, the memory and is divided: one single inbox receives everything and the skill /triagem routes to Project (what happened), Self (who it is) or Knowledge (facts). The [[wikilinks]] weave the three together.

📦 Project (episodic)

What happened: tasks, decisions, how things are progressing. It changes all the time — it’s Jarvis’s “diary.”

🪪 Self (identity)

Who it is and who you are: values, tone, preferences, open questions. Changes slowly — stable by design.

📚 Knowledge (reference)

Facts that only accumulate: recipes, contacts, study notes. They almost never get deleted — they only grow.

🔗 The glue between the zones

A single inbox receives every new note; one skill /triagem reads and routes each one to the right zone — you don’t need to decide on the spot. And the [[wikilinks]] (that “[[name]]” in double square brackets, inherited from wikis and Obsidian) link one note to another across the three zones: a Project decision can point to a Self value and a Knowledge fact.

Key concepts

(episodic) Project

What happened and the decisions—the memory that changes all the time.

Self (identity)

Values and questions — the memory that changes slowly.

Knowledge (reference)

Facts that only accumulate.

Inbox + /triagem + [[wikilinks]]

One entry point that routes, and links that cross the three zones.

6

🗂️ Memory = filesystem + index

In practice, how does this memory exist on disk? The answer is disconcertingly simple—and that’s why it’s robust. Jarvis’s memory is text files (format .md, Markdown) that you can open and read yourself. These files are the truth. Running on top of them is a search index (SQLite with FTS5/BM25), which is just a shortcut for finding the right note quickly — if the index disappears, you rebuild it from the files.

THE TRUTH · .md files projeto/decisoes.md self/valores.md conhecimento/notas.md readable · versionable · text only indexes rebuilds THE SHORTCUT · index (disposable) SQLite FTS5 / BM25 keyword search disappeared? reconstruct it from the .md files THE ROUTER · CLAUDE.md CLAUDE.md / AGENTS.md read at the START of each session points to where to find everything

The .md files are the truth — readable and versionable. The SQLite index and it’s just a search shortcut: if it disappears, it’s rebuilt from the files. It’s the CLAUDE.md e o router read at the start of each session, telling Jarvis where to find each memory brain.

🧭 CLAUDE.md is the router

At the start of each session, Jarvis reads a master file — the CLAUDE.md (or AGENTS.md). It doesn’t store everything: it points. "Decisions live in projeto/. My values live in self/. To find facts, use the index." It's the map that shows where the three brains are — which is why it's the first thing read.

It's the same logic as the Identity contract (T3-2): the file at the start defines the rules.

💡 Why this survives the hype

AI trends come and go, but text and text. Your file-based memory .md opens in any editor, goes into Git, migrates to any future tool. You’re not held hostage by a proprietary database or a format only one company understands. “It survives because it’s just text.”

New here? Filesystem and it’s simply your computer’s file system—folders and files. SQLite and a database that fits in a single file. FTS5/BM25 are SQLite’s “full-text search” mechanism and the formula that ranks results by relevance (search by word). A vector database and an alternative that searches for meaning instead of a word—useful when you can’t remember the exact term.

Key concepts

.md files = the truth

Readable text on disk; everything else is derived from it.

SQLite FTS5/BM25

A word-search index; a disposable, rebuildable shortcut.

CLAUDE.md = router

Read at the start of the session; points to where each brain lives.

Vector database

Search by meaning, an alternative to or complement for FTS5.

Self-check (optional): in this course’s model, what does “brains” mean?

🎯 Module summary

✓
The brain = the LLM engine — the part that thinks, pluggable through a standard socket. “One brain, many mouths.”
✓
Hot-swap via .env — switch engines (anthropic / openrouter / ollama) without changing the code; make cost-effective decisions by task.
✓
There is no "model council" — "multiple models" in the honest sense = substitution + portability (polyskill), not simultaneous voting.
✓
The second brain becomes memory — one big pile gets bloated; organizing it into zones brings order.
✓
The 3 brains — Project (episodic), Self (identity), Knowledge (reference) + inbox/triage + [[wikilinks]].
✓
Memory = filesystem + index — .md files are the source of truth, SQLite FTS5 is the shortcut, and CLAUDE.md is the router. It survives because it's all text.

You completed Anatomy (Track 3):

Channels · Identity · Tools · Skills · Agents · Brains. The 6 layers that turn a chat into an assistant that sees, remembers, and acts. Next: in Track 4, you’ll bring all this together in an architecture and build it.