🧱 From zero to your first Jarvis (the minimum recipe)
You already understand the pieces (Track 3). Now let’s put them together. This module gives you the minimum recipe — the “hello world” of a Jarvis—and the right order for stacking the building blocks: the 5 Build Levels. You’ll learn to build WITH an AI and test each piece before moving on, instead of putting everything together at once and hoping for the best.
🧱 The minimum recipe
There's a temptation to start big: voice, ten integrations, sub-agents, a beautiful dashboard. That's the fastest path to giving up. Jarvis starts much smaller than you imagine. The minimum recipe has exactly five ingredients—and not one more. If even one of these pieces is in place, you already have an assistant that truly converses and can grow safely.
🍳 The 5 ingredients of Jarvis's "hello world"
- 1.A channel — the port you use to talk to it. Use Telegram: it’s just a token, with no exposed server.
- 2.A brain — a single AI model (LLM) behind it, in the cloud or locally. Just one, not three.
- 3.A
soul.md— a text file that says who it is (name, tone, values). - 4.The agentic loop — the “think → act → read the result → respond” cycle we saw in Track 1.
- 5.A tool — just one (e.g., web search). It gives the agent a "hand" to do something in the world.
Notice what no is on the list: voice, multiple channels, vector database, ten plugins, dashboard. All of this is legitimate—but it comes later. The principle that organizes the entire path is brick-by-brick (brick by brick): build one piece, prove it works, and only then add the next. It’s the opposite of dumping everything into one giant file and hoping it works.
New here? "Brick-by-brick" (literally “brick by brick”) is the philosophy of building software in small, testable blocks, one at a time. Each block you understand and test becomes a solid foundation for the next. The opposite—pasting 100 thousand lines no one has read—was exactly the problem with the bloated systems we saw in Track 2.
Key concepts
Channel + brain + soul + loop + 1 tool. Nothing more.
The smallest Jarvis that already “lives” and responds for real.
Build and test one brick before stacking the next.
The think-act-respond cycle, the heart we saw in Track 1.
🪜 The 5 Build Levels
The minimum recipe says what goes in. The Build Levels say in what order. There are five steps, and the order matters: each one only makes sense if the previous one is already working. You don’t add voice to a bot that still doesn’t respond, or cron to a bot that still has no memory. Climbing one step at a time turns a daunting project into five small, rewarding tasks.
Each step is safe to take only when the one below is already working. In amber, the levels you’re building now (Foundation and Memory, in this module); in cyan, the ones you’ll see in Tracks 3 and 5 (voice, tools), which complete the cadence (heartbeat).
Foundation—it responds
A Telegram bot connected to a model. You send "hi," it replies. That alone proves channel + brain + loop.
Memory — it remembers
Add the soul.md (who it is) + a conversation database. It stops treating you like a stranger with every message.
Voice — it speaks and listens
Audio messages come in (transcription) and go out (voice). This is Track 5 — optional, decided at runtime.
Tools / MCP — it acts in the world
Connects real tools (email, calendar, GitHub) via MCP. This was the “Tools” layer of Track 3.
Heartbeat — it acts on its own
Cron/routines make it act on schedule ("7 a.m. summary"), even with the laptop closed. The leap to being proactive.
💡 Practical tip
In this module, you only put together the Levels 1 and 2. Resist jumping to voice or tools before you have a bot that responds and remembers. Most projects that die tried to climb three steps at once and stumbled.
Key concepts
The 5 steps: Foundation, Memory, Voice, Tools/MCP, Heartbeat.
Each step assumes the previous one is working.
Only Levels 1 and 2—a solid foundation before growing.
Every step is worth celebrating; it keeps the project alive.
🏗️ Level 1: Foundation (it responds)
The first step is the most exciting: the first time you send a message and the thing responds. Foundation has just three connected pieces — a channel (Telegram), a brain (a model), and the loop that connects the two — plus two safeguards that are in place from minute zero: the whitelist e o .env.
📊 Anatomy of Level 1
- •Channel: a bot created in
@BotFatherfrom Telegram gives you a token. No web server, no open port — the bot fetches messages by long polling. - •Brain: one cloud model (Claude/GPT via API key) or local model (Ollama on your PC). Just one.
- •Whitelist: a list with YOUR chat ID. Messages from anyone else are silently ignored.
- •
.env: a file where secrets are stored (token, API key). Never inside the code, never in Git.
New here? A .env ("dot env") is a plain text file with pairs CHAVE=valor that stores secrets — passwords, tokens, API keys. The program reads these values at startup, but they remain outside of the code. That way, you can publish the code without leaking your keys. The whitelist ("allowlist") is the opposite of a denylist: only those on it are served; everything else is blocked by default.
✓ Foundation done right
- ✓Token and API key live in the
.env. - ✓Whitelist with your ID, enabled from the very first line.
- ✓Telegram via long polling: no open ports.
- ✓One model, predictable behavior.
✗ Foundation done wrong
- ✗Token pasted directly into the code (and pushed to Git).
- ✗Without a whitelist: anyone who finds the bot can talk to it (using your account).
- ✗Web server exposed just to “make things easier” — an open door to the internet.
- ✗Three models and ten resources before the first "hi" works.
The security gaps in the red column aren't hypothetical: in Track 2, we saw that 42,665 instances of a popular system were exposed on the internet, 93.4% with no authentication. The whitelist is the .env aren't "extras" — they're part of Level 1. Security isn't added at the end; it starts with the first building block.
Key concepts
Your Telegram bot’s “password,” generated with a command.
The bot asks "Any messages?"—no exposed server, no port.
Only your ID is served; everyone else is silently ignored.
Secrets file, outside the code and outside Git.
🧠 Level 2: Memory (it remembers)
The Level 1 bot responds—but it has amnesia. End the conversation and it forgets everything; ask your name again and it doesn't know. Level 2 fixes this with two pieces: the soul.md, that says who it is, and a bank SQLite, which stores what's already been said. Together, they turn a generic chatbot into your assistant.
🪪 Two memories, two roles
Don't confuse the two. One is identity (changes slowly, written by you); the other is history (grows on its own with every conversation).
- •
soul.md— the soul: an injected Markdown file always at the beginning (the system prompt). Name, tone, values, priorities. “Without the brain rewire, the architecture is just a folder.” - •SQLite — the notebook: a database in a single file that stores every message. With an index FTS5/BM25, it searches by word ("what did we say about project X?") in milliseconds.
New here? O system prompt and the "briefing" the model reads before any conversation—and where the soul.md goes in. SQLite and a database that fits in a single file (no server, no complicated setup)—perfect for a personal Jarvis. FTS5/BM25 is SQLite's "full-text search": it finds the right message by keyword, like a turbocharged Ctrl+F. Since everything is text files plus a simple database, the memory "outlives the hype": it's just data you can read and take with you.
soul.md in the project folder and paste the block below, replacing the parts <assim> with yours.
# SOUL — quem voce e
Voce e o <Nome-do-seu-Jarvis>, o assistente pessoal de <Seu-nome>.
## Tom de voz
- Direto, caloroso e sem enrolacao.
- Responde em <portugues-do-Brasil>.
- Quando nao sabe, diz "nao sei" em vez de inventar.
## Prioridades do dono
- <ex.: me ajudar a organizar o dia e escrever melhor>
## Como agir
- SEMPRE confirma antes de qualquer acao que mande mensagem ou apague algo.
- NUNCA compartilha meus dados com terceiros.
soul.md is being read. Done: Level 2 has an identity.
📊 What changes when it remembers
- •It stops introducing itself with every message — it knows who you are.
- •Picks up old topics: "what happened with that email from yesterday?".
- •Memory is just one file — you can read, copy, back up, and take it with you.
Key concepts
The “character sheet” always injected into the system prompt.
The briefing the model reads before every conversation.
Database in a single file — stores the history.
Keyword search across the history, very fast.
🤝 Build WITH an AI
Here’s the turning point that makes all this accessible to a nontechnical user: you don’t need to write the code yourself. You build Jarvis talking with a coding AI — Claude Code, Antigravity, Cursor. Instead of typing every line, you describes what you want in a initialization prompt, and the AI generates the project brick by brick, with you in control.
The key is to give it a prompt good: to lock in the architecture decisions from the start (Telegram only, with a whitelist and secrets in the .env) so the AI doesn’t “improvise” an exposed web server or add a thousand dependencies. The prompt below is your starting point—ready to paste into Claude Code.
<assim>.
Voce vai me ajudar a construir meu agente de IA pessoal, tijolo por
tijolo (brick-by-brick). Eu sou leigo: explique cada passo em
portugues simples e nao avance sem eu confirmar.
Stack e decisoes (NAO mude sem perguntar):
- Canal: SOMENTE Telegram, via long-polling. SEM servidor web,
SEM portas abertas.
- Cerebro: um unico modelo, provedor <anthropic | openrouter | ollama>.
- Seguranca: whitelist com o meu chat ID = <seu-chat-id>.
Qualquer outro remetente e ignorado em silencio.
- Segredos: token e chaves SO no arquivo .env (nunca no codigo,
nunca no Git). Crie um .gitignore que ignore .env.
Construa nesta ordem e pare para eu testar a cada etapa:
1. Level 1 (Foundation): bot que responde "oi" com a whitelist ativa.
2. Level 2 (Memory): adicione soul.md (injetado no system prompt)
e um SQLite que guarda o historico da conversa.
Comece pelo Level 1. Antes de escrever codigo, liste os arquivos
que vai criar e me explique cada um.
⚠️ The mistake to avoid
Ask for “a complete Jarvis with voice, calendar, email, and ten plugins” all at once. The AI will generate a bunch of code you can't understand or test — exactly the “100,000 lines nobody reads” from Track 2. Ask one Level at a time, read what it explains, and move on only when the current building block passes the test.
Key concepts
Claude Code, Antigravity, Cursor — they write code with you.
The description that locks in the stack and security before the first line.
Telegram only, whitelist, secrets in .env—the AI doesn't improvise.
Confirms each step; understands every building block it adds.
🧪 Test each building block
"Brick-by-brick" only works if you prove that each brick bears weight before you stack the next one. The good news: testing a Jarvis is simple and doesn't require any tools — you send a message and look at the response. The key is to have, for each Level, an acceptance test: a question whose correct answer proves that brick is standing.
The loop that powers the course: builds a brick, runs the acceptance test, and it's just stacks the next one if it passes. Did it fail? Go back and fix it — never build on top of a cracked building block.
✅ Acceptance tests by Level
- L1Foundation: send "hi"—it replies. From another phone (outside the whitelist), send "hi"—it ignores. Both need to be true.
- L2Memory: say "my name is <Ana>". In a new message afterward, ask "what's my name?". It gets it right = memory works.
💡 Practical tip
Write the test before to ask for the building block. “Level 1 is ready when the bot only responds to me” is a clear criterion—you know exactly when you’re done and when to move on. Without a criterion, every project becomes an eternal “almost done.”
Self-check (optional): What is the right order for building your first Jarvis?
Key concepts
The question whose correct answer proves the brick works.
Define “done” before building; it prevents the endless “almost done.”
Confirm that it SERVES you and BLOCKS everyone else.
Did one building block fail? Fix it before adding the next.
🎯 Module summary
Next module:
4-2 — Architecture of a robust solution (the 4 C layers: Context · Connections · Capabilities · Cadence)