🛠️ Building your pocket Jarvis (project)
Time to put it all together. In this module, you’ll go from “I understand the theory” to “it’s working on my phone”: a personal Telegram bot, with a brain, memory, soul, a useful skill, voice, and a heartbeat that sends you the day's summary at 7 a.m. Five short steps, each testable before the next. No app store, no exposed server, no mixing of keys.
🗺️ Project overview
Everything you've seen in the previous tracks comes together here in one sentence: let's build a personal Telegram bot that talks with you, remembers who you are, runs a useful routine, and even speaks — accessible from any phone, without installing an app from an app store. Telegram is already on your phone; it becomes the entry point. The brain (the AI model) and memory run on a computer of your own (at home) or in the cloud. The phone is just the channel.
The project’s golden rule is brick-by-brick (brick by brick): you build ONE piece, test that it works, and only then move on to the next. Don’t build everything at once and then try to figure out what broke. There are five steps, and each one brings your Jarvis a little more to life.
O cell phone never holds the brain — it only talks, through Telegram, to the agent that lives on your PC or in the cloud. Each cyan-blue piece is one of the five steps: you connect them one at a time.
New here? A bot and a program that communicates through a messaging app (here, Telegram) instead of its own screen. Agent and it’s our Jarvis: the program that receives your message, thinks with the AI model, and responds—sometimes using tools. Heartbeat ("heartbeat") is a scheduled task that makes the agent act on its own at a set time, without you asking.
🧰 What you’ll need (once)
- •Your own computer (Windows, Mac, or Linux) with Node.js or Python installed — where the agent will run.
- •The account for Telegram that you already use on your phone.
- •A model key (cloud: OpenRouter/Anthropic) or o Ollama installed to run locally — choose in Step 2.
- •~40 minutes and the willingness to test each building block.
Key concepts
The door you use to talk to Jarvis—here, Telegram.
Build and test one piece at a time, never everything at once.
"Telegram is already mobile" — no need to publish a native app.
The brain lives on the PC/cloud; the phone just talks to it.
🤖 Step 1: the bot
The first building block is creating the bot. On Telegram, there’s an official bot called @BotFather that creates other bots for you. You talk with it and receive a token (a long password that identifies your bot), and that’s it—the channel exists. Next, you add the first security measure: the whitelist (allowlist), so that only you be served.
Talk to @BotFather
On Telegram, search for @BotFather, open the conversation and send /newbot. It asks for the name and username (which needs to end in bot, e.g.: meu_jarvis_bot).
Save the token
It responds with something like 7891234:AAF...xYz. This is the token — treat it like a password. It will live in a file .env, never in the code or in a public printout.
Find your user ID
Talk to @userinfobot: it replies with your user number (e.g., 123456789). That’s the number you add to the whitelist.
.env
Objective: store the token and whitelist outside the code. Create a file called .env in the project folder and paste the block below, replacing the pieces between < > with your values.
# .env (NUNCA suba este arquivo pro GitHub)
TELEGRAM_BOT_TOKEN=<cole-seu-token-do-BotFather>
TELEGRAM_ALLOWED_IDS=<seu-user-id-do-userinfobot>
How to check: run the agent (e.g., npm start or python main.py), send /start to your bot on your phone and see if it responds. Then ask a friend to message the same bot: they should be silently ignored. If that happens, the whitelist is working.
New here? A file .env ("environment") is a block of text where you store secrets (tokens, keys) separately from the code. Long-polling and how Telegram works without opening any ports on your PC: your agent is the one that asks to Telegram, "got a message?", instead of waiting for someone to knock on your machine. That's why the bot is safe even when running at home.
Key concepts
The official Telegram bot that creates your bot and gives you the token.
The bot’s password. It lives in the .env, never in the code.
List of allowed ID(s); any others are ignored.
No exposed port: the agent asks Telegram if anything has arrived.
🧠 Step 2: the brain
The bot already receives messages—but it still doesn't think. Now you plug in the brain: the AI model that will generate the responses. Here you make the most important choice in the project, and it's just one line in the .env. You can use the cloud (powerful, pay-per-use) or run local with Ollama (free after downloading, private, slower). The best part: switching between the two doesn’t change the code — it’s hot-swap (hot swap).
☁️ Cloud (OpenRouter / Anthropic)
- ✓Fast, strong responses, even on a weak PC.
- ✓No hardware setup.
- ✗You pay per token; the data leaves your machine.
- ✗You need internet access.
🏠 Local (Ollama)
- ✓$0 per token after download; data stays at home.
- ✓Works offline.
- ✗Requires RAM (3B@8GB, 8B@16GB); CPU is slow (30-60s).
- ✗Responses that are a little weaker than top cloud models.
.env
Objective: choose the brain by changing just two lines. Add ONE of the blocks below to your .env (leave the other one commented out with #).
# --- opcao A: NUVEM (potente, paga por uso) ---
LLM_PROVIDER=openrouter
LLM_MODEL=<ex: anthropic/claude-3.5-sonnet>
LLM_API_KEY=<cole-sua-chave-do-openrouter>
# --- opcao B: LOCAL com Ollama (gratis, privado) ---
# LLM_PROVIDER=ollama
# LLM_MODEL=<ex: llama3.2>
# (rode antes, no terminal: ollama pull <ex: llama3.2> )
How to check: restart the agent and message it from your phone: “in one sentence, who are you?”. If you get a coherent response, the brain is connected. To test hot-swapping, switch the active block (comment out A, uncomment B), restart, and send it again — same conversation, different brain, without touching the code.
⚠️ Classic caution: don't mix keys
OpenAI key isn't Anthropic key, and neither one is the Telegram token. Each service has its own. Mixing them up is the number one cause of "it doesn't work and I don't know why." If an authentication error appears (401/403), the first place to look is which key went into which variable in .env.
Key concepts
The model that thinks in text and generates responses.
Change the model by changing the .env, without touching the code.
A program that runs AI models on your own computer.
Local for simple tasks, cloud for difficult tasks.
💾 Step 3: memory + soul
Now the bot can think, but it’s a stranger: it forgets everything when you close it and has no personality. Two files solve that. The soul.md ("the soul") of character — who it is, its tone of voice, what it prioritizes — and it’s injected always at the beginning of the conversation. The memory makes it remember: conversations become saved text and an index SQLite lets you search what has already been said. Without the soul, you have "a folder of code"; with it, you have a Jarvis.
soul.md
Objective: give the bot an identity. Create a file soul.md in the project folder and paste the block, replacing the pieces between < >.
# soul.md — a alma do meu Jarvis
## Quem sou
Eu sou o <nome-do-seu-jarvis>, assistente pessoal de <seu-nome>.
## Tom de voz
Direto, caloroso e breve. Sem enrolacao. Trato <seu-nome> pelo nome.
## Prioridades
- Responder em portugues.
- Quando nao souber, dizer que nao sabe (nunca inventar).
- Lembrar do contexto das conversas anteriores.
## Sobre <seu-nome>
Fuso: <ex: America/Sao_Paulo>. Trabalho: <ex: professor>.
Prefere respostas em topicos curtos.
How to check: restart and send "what’s your name and who am I to you?". It should reply with the name of the soul.md and call you by your name. Then send “my favorite dish is lasagna”, close Telegram, open it again later, and ask "what’s my favorite dish?". If it remembers, SQLite memory is recording.
🧩 How memory works, without the mystery
- •Each conversation becomes text saved in a file or database — the “truth” is readable by you.
- •An index SQLite FTS5/BM25 (keyword search) quickly finds the relevant passage when you ask something.
- •Before responding, the agent searches memory and pastes the passage into context. That’s why it “remembers.”
New here? SQLite and a database that's just one file in your folder (no server). FTS5/BM25 and how it searches for text—you ask for "lasagna" and it finds the sentence where you mentioned it. System prompt and it’s the invisible text attached to the start of every conversation: that’s where the soul.md goes in so the model “knows who it is.”
Key concepts
The soul: personality, tone, and priorities, always injected.
Conversations turn into text that survives closing the app.
Simple file database; the memory index.
The opening text where the soul becomes the model’s identity.
🧩 Step 4: a useful skill
So far, your Jarvis can talk and remember. Time for it to do something concrete. A skill and a "recipe": a short file that packages steps you'd repeat every time. Ours: "summary of my day" — it pulls up your calendar (via a tool MCP from the calendar) and returns a bullet-point summary. Instead of explaining five steps, you say one sentence and the skill executes them.
MCP (Model Context Protocol) is the "USB for AI tools": each integration (calendar, email, GitHub) is a separate, standardized server that you plug in without rewriting the agent. It’s safer than downloading "community skills" from questionable sources — you know exactly what each server does.
skills/resumo-do-dia/SKILL.md
Objective: give Jarvis the recipe for summarizing your day. Create the folder and file SKILL.md below, swapping the pieces between < >.
---
name: resumo-do-dia
description: Resume os compromissos de hoje em topicos curtos.
Acionar quando o usuario disser "resumo do dia", "como esta
meu dia" ou "agenda de hoje".
---
# Resumo do meu dia
Passos:
1. Use a ferramenta de calendario (MCP) para listar os
eventos de HOJE no fuso <ex: America/Sao_Paulo>.
2. Para cada evento, pegue horario + titulo.
3. Responda em ate <ex: 5> topicos curtos, do mais cedo
ao mais tarde. Comece com "Bom dia, <seu-nome>!".
4. Se nao houver eventos, diga que a agenda esta livre.
How to check: restart the agent and message it from your phone "summary of my day". It should activate the skill, check the calendar, and reply with a bulleted list of events (or say that you’re free). If it gives a generic answer without checking the calendar, make sure the calendar MCP server is connected.
📊 Skill vs. tool — what's the difference?
- •Tool: an atomic action — "list calendar events." It's a hand.
- •Skill: a procedure that orchestrates tools + judgment—"put together the day's summary." It's the recipe that uses its hands.
- •The skill costs almost nothing until it’s called: the agent reads only the frontmatter (~100 tokens) and only loads the rest when needed.
New here? O frontmatter and the block between --- at the top of the file (name + description). The agent uses it to decide whether to activate the skill. Progressive loading means the entire recipe is read only when needed—saving context and keeping Jarvis lightweight.
Key concepts
A recipe packaged in a markdown file.
The "USB for tools": plugs in standardized integrations.
Name + description at the top; decides when to activate the skill.
The recipe is read only when needed.
🎙️ Step 5: voice and cadence
The last two building blocks make Jarvis human and proactive. Voice: you send an audio message on Telegram, and the agent transcribes it with Whisper (understands your speech), thinks, and — if you want — responds out loud with TTS (speech synthesis). Cadence: a heartbeat scheduled (cron) makes it act on its own—for example, sending you the “summary of my day” every day at 7 a.m., even with your phone in your pocket and your laptop closed.
O audio turns into text (Whisper), the LLM thinks in text, and speaking back (TTS) is optional. Remember: "voice is the interface, the orchestrator is the brain" — you decide for each message whether to respond in text or voice.
.env + terminal
Objective: connect voice and schedule the daily summary. Add to the .env and test audio conversion in the terminal.
# .env (voz e cadencia)
VOICE_ENABLED=true
STT_PROVIDER=whisper
OPENAI_API_KEY=<chave-da-OpenAI-so-para-Whisper>
# heartbeat: rodar a skill "resumo-do-dia" todo dia as 7h
HEARTBEAT_CRON=<ex: 0 7 * * *>
HEARTBEAT_SKILL=resumo-do-dia
# --- teste rapido da conversao de audio (no terminal) ---
# ffmpeg -i <audio-do-telegram.ogg> -ar 16000 saida.wav
How to check: (voice) record an audio message in Telegram saying “tell me an interesting fact” — the agent should transcribe and respond. (cadence) so you don’t have to wait until 7 a.m., temporarily change HEARTBEAT_CRON for 2 minutes from now (e.g., if it’s 2:32 p.m., use 34 14 * * *), restart, and watch the message arrive on its own. Did it work? Set the cron back to 0 7 * * *.
⚠️ The Whisper key is different
Whisper (transcription) usually uses the OPENAI_API_KEY, which is NOT your brain’s key if you chose Anthropic/OpenRouter, nor the Telegram token. Three services, three credentials. If the voice “won’t transcribe,” it’s almost always because this key is missing or has been swapped.
New here? STT ("speech-to-text") means turning speech into text; Whisper and it’s the model that does this. TTS ("text-to-speech") is the reverse process: text becomes voice. ffmpeg and an audio/video converter that prepares the Telegram file for Whisper. cron and the scheduling notation: 0 7 * * * means "at 7:00 a.m. every day."
Key concepts
Turns the audio you send into text.
Optional output: Jarvis responds out loud.
A schedule that makes the agent act on its own at a set time.
From “responds when I call” to “alerts me beforehand.”
Self-check (optional): in your pocket Jarvis, where does the brain (the model) live?
🎯 Module summary
Next:
Return to the track—you’ve completed Track 5 (Jarvis on your phone). Time to review your journey and move on to Track 6.