🧬 Anatomy
The 6 layers that turn an ordinary chat into an assistant that sees, remembers, and acts. Each module opens up a "part of the body" of your Jarvis—Channels, Identity, Tools, Skills, Agents, and Brains—so you can understand the whole piece by piece.
Read from left to right: you talks to the Jarvis, which triggers the 6 layers in order—and the result is a action in the world. Each module in this track opens up a layer.
Learning path map
Detailed content
📡 Channels — how you talk to it
Your Jarvis’s input and output channels. Why Telegram became the favorite, how WhatsApp and voice fit in, and why the whitelist is the first security measure.
A channel is the "doorway" you use to talk to Jarvis: it could be Telegram, WhatsApp, voice, a web page, or the terminal itself. The brain stays the same; only the doorway changes.
Choosing the right channel determines whether you’ll talk to it from your phone, your PC, or anywhere—and also how much that exposes the system.
Channel, input/output, multichannel, the single brain behind the doors.
Telegram only asks for one token (a password generated by @BotFather). It uses long polling: it's your Jarvis fetching messages on Telegram instead of opening a web server and waiting for a connection.
Long-polling = no open port to the internet = one less port for attackers. That makes it the safest channel to start with—plus it's free and already mobile-friendly.
@BotFather, token, long-polling, zero exposed servers.
Beyond Telegram, you can plug in WhatsApp (text, image, and voice, via Evolution API), email, and web. Multichannel means having the same brain respond in multiple places.
Everyone uses a different app; knowing the brain is independent of the channel lets you choose where your Jarvis appears.
WhatsApp, Evolution API, email, web, one brain / many mouths.
A whitelist and a list of who can talk to the bot. Only YOUR user ID is served; everyone else is silently ignored—they don't even get an error response.
A public bot that replies to everyone spends your money and can be manipulated. A whitelist is the simplest defense and the first one you turn on.
Whitelist, user ID, single-owner, silently ignore.
The same channel carries text and audio (Telegram sends voice messages as a file .ogg) and image. Voice is an input/output optional, decided on the spot.
Understanding that voice is just another type of media on the channel prepares you for Trail 5, where the voice pipeline is actually built.
Text, .ogg audio, image, optional voice decided at runtime.
Since the channel is a messaging app, your Jarvis is accessible from any device where you have that app—phone, tablet, PC—without needing its own app in the store.
It’s the shortcut that saves you from building and publishing an app: “Telegram is already mobile.” This comes back in a big way in Track 5.
No native app, multiple devices, synced by the app itself.
🪪 Identity — memory, persona, and who it is
What gives Jarvis character and memory: its soul (SOUL.md), its contract (AGENTS.md), your profile, the memory that survives after you close the conversation — and why “who it serves” is also part of its identity.
O SOUL.md and a text file containing the owner’s personality, values, tone, and priorities. It is ALWAYS injected into the system prompt (the invisible instructions the model reads before anything else).
“Without the brain rewire, the architecture is just a folder”: SOUL.md is what makes the assistant sound like YOURS, not a generic chatbot.
SOUL.md, system prompt, character profile, tone, and values.
O AGENTS.md says what it DOES and NEVER does, in explicit "ALWAYS" and "NEVER" lists. It's the behavior contract.
A written rule is better than hoping the model guesses. This is where you set limits — which matter a lot in the children's learning path.
AGENTS.md, ALWAYS/NEVER lists, explicit rules, guardrails.
The file USER stores who you are: name, context, preferences, people, and projects that matter. It's the owner's profile.
Without this, Jarvis treats you like a stranger in every conversation. With a profile, it already knows your name, your time zone, and what you usually ask for.
USER, the owner's profile, personal context, preferences.
The conversation forgets everything when you close it. The persistent memory are files .md + an index SQLite FTS5/BM25 (keyword search) and/or a vector database (search by meaning).
It's the difference between a chat that starts from scratch every time and an assistant that remembers what you agreed on last week.
Persistent memory, .md, SQLite FTS5/BM25, vector database.
Who the Jarvis serves is part of who it is: the ID whitelist (from module 3.1) plus the secrets (API keys) stored only in the file .env.
Identity isn’t just personality — it’s also a boundary. Mixing persona and authentication makes it clear that security is part of the agent’s character.
Whitelist, secrets, .env file, identity boundary.
The same Jarvis can switch modes (work, study, “story time mode”) simply by changing which SOUL/profile is active. A persona and a set of interchangeable behaviors.
Instead of maintaining several bots, you maintain one and change its “outfit.” This is the basis of the children's tutor in Trail 6.
Persona, modes, active SOUL, one agent / multiple outfits.
🔧 Tools — Jarvis’s hands (tools + MCP)
How Jarvis goes from "just talking" to taking action: what tools are, how the model calls them, and why MCP became the "USB" that safely connects any tool to any agent.
On its own, the LLM only produces text. The tools (tools) give it hands and eyes: search the web, read/write files, send email, check the calendar, run a command.
It’s what separates “talking” from “doing.” Without tools, it “responds like a stranger”; with them, it becomes a real assistant.
Tool, action in the world, real data, hands and eyes.
Function-calling and the mechanism where the model "asks" for a function (e.g.: enviar_email(...)); the system actually runs it and returns the result so the loop can continue.
It’s the plumbing underneath it all. Understanding that the model only REQUESTS (it doesn’t execute) explains how confirmation and sandboxing are possible.
Function-calling: the model requests / the system executes, then returns to the loop.
O MCP (Model Context Protocol, by Anthropic, Nov/2024) is a standard where each integration lives in a MCP server separate and auditable. It’s the “USB” for AI tools.
With MCP, you can plug in Gmail, Notion, or GitHub without rewriting the agent—and you can read exactly what each server does.
MCP, MCP server, universal standard, pluggable integrations.
Downloading third-party "skill files" is risky: in OpenClaw, they found 341 malicious skills. MCP isolates each integration in an auditable server with clear permissions.
The course’s security lesson: prefer the auditable standard (MCP) over code from a stranger that you’ve never read.
341 malicious skills, third-party code, auditability, MCP-only.
Dangerous actions (running shell commands, deleting a file) ask for confirmation and they run on a sandbox (isolated environment). And every input is treated as a potential prompt injection (text that tries to hijack the agent).
Giving an agent hands without safeguards is dangerous. Confirmation + sandboxing let you grant power without losing control.
Confirmation, sandbox, prompt injection, zero trust.
Without real data (your calendar, your email), Jarvis only gives generic answers. Connecting tools turns it from a “conversationalist” into an ASSISTANT that acts on your life.
It’s the “why” behind the whole layer: connecting to the real world is what gives the agent practical value.
Real data, generic vs. personal, Connections, practical value.
🧩 Skills — packaged abilities
Reusable recipes: what a skill is, why to package a step-by-step guide, how it uses almost no context until it’s used, and the honest truth about portability between models.
A skill and a recipe: a Markdown file (SKILL.md) that packages a repeated step-by-step process. Example: “write a LinkedIn post” = research → chart → text → review → publish.
It’s how you teach a procedure once and reuse it forever—the foundation of your Jarvis’s capability library.
Skill, SKILL.md, recipe, packaged step-by-step guide.
Instead of explaining the 5 steps every time, you say one phrase (“make my post for the week”) and the skill carries out the entire process.
Packaging turns repeated instructions into a short command—less effort from you, more consistent results.
Abstraction, short commands, consistency, reuse.
O progressive loading (progressive disclosure) makes the skill take up only the frontmatter (~100 tokens with a name and description) until it’s invoked. Only then does the full step-by-step process enter the context.
You can have dozens of skills without stuffing the context window—the agent "knows they exist" without loading them all.
Progressive disclosure, frontmatter, ~100 tokens, context savings.
A tool and an atomic action ("send email"). A skill and a procedure that ORCHESTRATES several tools and uses judgment across them.
Confusing the two leads to designing them incorrectly. A tool = a single verb; a skill = the recipe that chains the verbs together.
Atomic action vs. procedure, orchestration, judgment.
The spec for Agent Skills + o polyskill let ONE skill source run in both Claude Code and Codex. “Multiple models” here means portability, not a council of models voting.
Avoids the common myth: the real advantage is not being tied to one runtime, not "multiple brains making decisions together".
Agent Skills, polyskill, cross-runtime, honest portability.
Each run can improve the recipe: the agent notices what failed and updates SKILL.md. The skills library becomes your collection of capabilities that grows with use.
Skills aren't static—they're living assets. That changes how you think about "teaching" your Jarvis over time.
Self-correction, a library of capabilities, continuous improvement.
🤝 Agents — the loop that works on its own
The autonomy layer: the agentic loop as a component, sub-agents that save context, multi-agent systems, the cadence (heartbeat/cron) that acts without you—and the safeguards that keep everything secure.
A agent and it’s an LLM running on loop: receives a request → thinks → calls a tool → reads the result → calls more tools → responds. Here it becomes a layer of the anatomy, with autonomy.
It’s the engine you saw in Track 1; seeing it again as a layer connects the pieces (tools, skills) into something that acts on its own.
Agentic loop, autonomy, request→think→tool→response.
The main agent delegates a heavy task to a sub-agent, which uses up its OWN context and returns only the final answer. The main session stays lightweight.
It’s the trick for big tasks (reading 50 files, doing lots of research) without clogging the main agent’s context window.
Subagent, delegation, isolated context, only the response comes back.
Multi-agent and it means having several specialists (research, writing, review) working in coordination. More power, but also more parts to coordinate and more things that can go wrong.
Knowing when it's worthwhile (and when it only adds complexity) helps you avoid over-engineering — a common mistake when you discover the pattern.
Multi-agent, specialists, coordination, complexity cost.
Cadence are time triggers — heartbeat, cron, routines — which make Jarvis act on its own schedule: the 7 a.m. summary, monitoring the inbox, “with the laptop closed.”
It’s what turns the assistant from “responds when I call” into “works while I sleep.” It returns as the 4th CLAWS layer in Track 4.
Heartbeat, cron, routines, scheduled action.
An autonomous loop needs guardrails: iteration limit (so it doesn’t run forever), approval gates for serious actions, a audit log (a forensic record of every action).
Unlimited autonomy is like letting the car drive without brakes. These controls are what make "working on its own" reliable.
Iteration caps, approval gates, forensic audit log.
An agent reactive responds when you call. A proactive notifies you before you ask ("your flight was delayed, so I’ve already reorganized your schedule"). That’s the leap from assistant to partner.
It’s the ultimate goal of the agent layer—and it only makes sense with cadence, memory, and limits in place.
Reactive vs. proactive, anticipation, partnership.
🧠 Brains — the engine that thinks and the organized memory
Honestly: “brains” doesn’t mean a panel of models voting. It means (a) a swappable LLM engine and (b) memory organized into zones. Here you’ll see how to switch brains without changing the code and how the “second brain” becomes memory.
The “brain” is the LLM engine — the model that thinks. It is pluggable: you can hot-swap (hot swap, without restarting everything) through .env between anthropic, openrouter, or ollama. “One brain, many mouths.”
Treating the model as a replaceable component frees you from any single provider — the course’s central thesis.
LLM engine, hot-swap, .env, pluggable, one brain/many mouths.
The same interface accepts different models: an inexpensive local model for simple tasks, a powerful cloud model for difficult ones. The choice becomes an economic decision for each response.
It’s how you balance cost and quality without rewriting anything—just by pointing to the .env for another model.
Same interface, low-cost vs. powerful, cost per response.
There is no judge or magical router in these projects, with N models voting together. Honestly, "multiple models" means replacement e portability (cross-runtime), not simultaneous voting.
It’s a common misconception in AI marketing. Demystifying it saves you from chasing a feature that doesn’t exist.
The myth of the council, replacement, portability, no voting.
The other meaning of “brains” is the memory of Jarvis, organized like a "second brain." A pile of notes grows and turns into a mess; separating them into zones brings order.
Disorganized memory is just as bad as no memory. Structure is what helps the agent find what matters.
Second brain, organized memory, zones, avoiding clutter.
Memory is divided into 3: Project (episodic: what happened, decisions), Self (identity: values, questions — changes slowly) and Knowledge (reference: facts that only accumulate). One unified inbox + one /triagem skill routes things, and [[wikilinks]] cross all three.
It’s the concrete system you’ll use to keep your Jarvis’s memory from turning into a dumping ground.
Project/Self/Knowledge, inbox + /triagem, [[wikilinks]].
At its core, memory is files .md readable (the truth) + an index SQLite FTS5/BM25 (the quick search). The CLAUDE.md/AGENTS.md and the ROUTER read at the start of every session.
“It survives the hype because it’s just text”: understanding this frees you from depending on exotic databases and lets you own your memory.
Filesystem .md, SQLite FTS5/BM25 index, CLAUDE.md/AGENTS.md router.