📲 Where Jarvis lives on your phone (the channels)
You don’t need to download an app from a store to have an AI assistant in your pocket. The secret is simple and even a little surprising: you use a messaging app that already runs on your phone—Telegram or WhatsApp—as the "door" to your Jarvis. In this module, you’ll discover where it actually lives, how your phone becomes just the screen for an agent that works 24/7, and why it’s more YOURS than Siri or Gemini.
📲 The shift: no app store app
When someone imagines “an AI assistant on a phone,” they almost always think of downloading a new app from the App Store or Play Store, creating another account, accepting another terms-of-use agreement. The turning point in this module is realizing that none of this is necessary. You already have a messaging app installed and logged in that opens on any phone in the world: Telegram or WhatsApp. That app becomes the door to your Jarvis.
Think of it this way: your assistant isn't an app; it's a contact. You message it like you message a friend, and it replies in the same place. The phrase that sums up this entire track is: "Telegram is already mobile". Since the messaging app runs on iPhone, Android, computers, and in the browser, your Jarvis automatically goes along too—without you writing a single line of app code.
🚪 What is a “channel”
A channel and it’s simply the door you use to talk to the assistant. Instead of building an entire app, you reuse a door that already exists and that everyone already knows how to use.
- •You don’t install anything new—you use the messaging app that’s already on your phone.
- •The same assistant is accessible from your phone, tablet, and PC at the same time.
- •Zero app maintenance: no build, no app store approval, no version updates.
New here? "Native app" is a program made specifically for iPhone or Android that you download from the app store. A "channel" is the assistant's conversation entry point. The key insight is to use an existing channel (Telegram/WhatsApp) instead of build a native app from scratch.
Key concepts
The conversation interface for Jarvis (Telegram, WhatsApp, voice, web).
The messaging app brings your assistant to any device.
No build, no store, no approval — just the channel that already exists.
You talk to it like a friend in your chat list.
✈️ Telegram (agentejax)
O Telegram and is the preferred channel for a pocket Jarvis, and the project used as a reference here is called agentejax — a TypeScript agent that lives on Telegram, is available 24 hours a day, and is built to be simple. Why Telegram? Because creating a robot (a "bot") there is incredibly easy: you talk to an official robot called @BotFather, it gives you a token (a long password that identifies your bot), and that’s it—you already have a channel.
The second reason is technical, but important for your security: Telegram uses long polling. Instead of your agent opening a "door" on your computer for the entire internet to knock on (which is risky), it’s the your agent that calls Telegram every moment asking "did I get a new message?" Result: no exposed ports, no web server open to the world — one fewer security vulnerability.
🛡️ The first safeguard: the whitelist
A Telegram bot can, by default, be found by anyone. Without protection, a stranger could message it and use YOUR assistant (and YOUR AI account). The safeguard is called whitelist: a list with your user ID (your Telegram account’s unique number). The agent only responds to people on the list; everyone else is silently ignored.
- •Only your ID: only you are answered.
- •The others disappear: messages from outside don't even get a response—there's no way to guess a bot is there.
- •Security by default: the door stays closed to everyone except the owner.
Reading the diagram: your message leaves the cell phone, goes through the Telegram (which doesn’t open any ports on your agent, thanks to long polling), reaches agent running in the cloud or on your PC, and only there does it activate the tools. The phone is just the display window.
Objective: create your own Telegram bot, get the token (the bot's password), and find your user ID for the whitelist. These are 3 conversations right inside Telegram — no need to install any software.
/start
/newbot
<Your-Jarvis-Name> # e.g.: My Jarvis
<bot_username>_bot # must end in "_bot", e.g.: myjarvis_bot
# @BotFather replies with something like this — save this line:
# "Use this token to access the HTTP API:"
123456789:AAExxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx # <-- THIS is your TELEGRAM_BOT_TOKEN
# 2) Find YOUR user ID: open @userinfobot and send anything.
# It replies "Id: 987654321" -> that number is yours <YOUR_USER_ID>
# 3) In the agent’s .env file (agentejax), fill in:
TELEGRAM_BOT_TOKEN=<paste-the-token-from-step-1>
TELEGRAM_WHITELIST=<YOUR_USER_ID> # only this ID will be served
How to check: with the agent running, open the conversation with your bot and send hi. If it responds, the token is right. To test the guardrail, ask a friend (who is NOT on the whitelist) to message the bot: they shouldn’t get any response. If your “hi” doesn’t come back, check that the token was copied in full (with no spaces) and that the agent is up.
Key concepts
The official Telegram bot that creates bots and provides the token.
The long password that identifies and authorizes your bot. Never expose it.
The agent connects to Telegram; no port is left open to the world.
List of authorized IDs; only the owner is served.
💬 WhatsApp (agentevoz)
In Brazil, the dominant messaging channel is WhatsApp. The reference project for this is called agentevoz — a Python agent that serves you through WhatsApp and was designed for real customer support: it receives text, images, and even audio, then decides on the spot how best to respond. The technical bridge between your agent and WhatsApp is called Evolution API.
🔌 Evolution API: the WhatsApp adapter
WhatsApp wasn’t originally designed for bots like Telegram. The Evolution API and an intermediate program (a "bridge") that connects your WhatsApp account to your agent: it receives incoming messages and forwards them, then sends the agent's replies back. You connect it to your number by scanning a QR code, much like WhatsApp Web.
In other words: the agent doesn't speak “WhatsApp” directly; it talks to the Evolution API, which translates for WhatsApp. Switching channels is almost just switching the adapter.
The agentvoz’s biggest differentiator is the Policy Engine ("rules engine"): a layer that decides, message by message, how respond. It chooses between text, voice, or hand off to a human, and also between a low-cost model or a premium model. That way, the assistant doesn’t waste an expensive model on answering “good morning,” and knows when it’s time to call a real person.
✓ What the Policy Engine decides
- ✓Responding text (fast and cheap) or in voice.
- ✓Use a model inexpensive for the simple part, premium for the hard part.
- ✓Recognize when it’s better hand off to a human.
- ✓Savings: spend more only when it’s worth it.
✗ Without a Policy Engine
- ✗Every "hi" hits the most expensive model — inflated bill.
- ✗Always the same format, even when voice would make more sense.
- ✗The bot tries to solve everything on its own—and gets stuck on what only a human can solve.
- ✗No clear rule for when to escalate or conserve resources.
New here? "Evolution API" is the bridge program that connects WhatsApp to your agent. "Policy Engine" is the decision layer that chooses text/voice/human and a cheaper or more expensive model for each message. The "model" is the AI brain (we covered it in Track 1). Voice itself, with Whisper and TTS, is the subject of the next module (5.2).
Key concepts
The bridge that connects your WhatsApp account to the agent.
Decides text/voice/human and low-cost/premium model for each message.
The same brain serves multiple channels at once.
When the bot recognizes that only one person can resolve it.
🖥️ Where does the agent run?
This is the question that confuses newcomers the most: if the phone is just the door, where does the assistant actually run? The answer: the agent program—the code that thinks, remembers, and acts—runs on a computer that stays on, and your phone simply talks to it. You have two main choices for a “home” for that agent.
In the cloud (e.g., Railway)
You move the agent to a service that keeps the program running 24/7 for a small, fixed monthly fee. It doesn’t depend on your PC being on; the assistant responds even while you’re asleep. The Railway and a popular example: you deploy it and forget about it.
On your home PC
The agent runs on your own computer (or a dedicated mini PC) that stays on at home. Hosting costs nothing, and your data never leaves your home. The tradeoff: the PC needs to stay on for the assistant to be available.
📊 The mental rule
Always separate two things that seem like one:
- •The channel (Telegram/WhatsApp) lives on your phone — it’s the door.
- •The agent (the program that thinks) lives in the cloud OR on your PC — it’s the engine.
- •Because the brain runs on a machine that’s always on, it can work 24/7 — even while you’re not looking at your phone.
💡 Practical tip
Start on your PC to learn and test for free; when you want full availability (answering in the middle of the night, sending the 7 a.m. summary on its own), move to the cloud. The code is almost the same—the only thing that changes is where you run the program.
Key concepts
The door (phone) and the engine (server) are separate things.
Host the agent 24/7 without relying on your PC being on.
Zero hosting costs and data that never leaves your home.
How the agent lives on a machine that's always on and acts all the time.
🔒 Privacy in your pocket
Here’s the best part of "Jarvis on your phone": having the assistant in your pocket doesn't mean hand your data over to some random cloud. Because the channel (on your phone) and the brain (in the agent) are separate, you can keep the 100% local brain — running on your home PC — and leave only the channel traveling with you on your phone.
The piece that makes this possible is called Ollama: a program that runs AI models on your own machine, for free, without internet access and without sending anything outside. You chat through Telegram on your phone; the message reaches the agent on your PC; the agent asks a model that lives right there in your home. The data goes from your phone to your computer and back—never to OpenAI or Anthropic.
🏠 The brain stays at home; the doorway goes in your pocket
It’s the best of both worlds: the convenience of talking from your phone anywhere, with the privacy of a brain that never leaves your home.
- •Channel on your phone: the Telegram you already use.
- •Local brain: Ollama running on your home PC, $0 per token.
- •Result: an assistant in your pocket, data in your living room.
Technical honesty: the message still passes through Telegram/WhatsApp servers to reach you—that’s true of any conversation in these apps. The privacy benefit here is that the reasoning and the memory of the assistant happen on your machine, not on an AI company's server. For complete end-to-end privacy, the next module (5.2) covers true on-device processing.
Key concepts
The model runs on your PC; the data stays at home.
A program that runs AI models locally, for free and offline.
Lets you use a phone portal and home intelligence at the same time.
A local model doesn’t charge per use — explore as much as you like.
🚧 Limitations of native assistants
"But my phone already has Siri (or Gemini, or Alexa). Why go through all this work?" Good question. The difference comes down to one word: ownership. Siri and Gemini are native assistants — closed, controlled by Apple and Google. You can’t see the code, choose the model, define the personality, or decide where your data goes. Your Telegram/WhatsApp bot, on the other hand, is your.
✗ Native assistant (Siri / Gemini / Alexa)
- ✗Black box: you don’t read the code or know what it does under the hood.
- ✗Without your personality: you can't rewrite who it is.
- ✗Company data: what you say goes to their servers.
- ✗Limited tools: only does what the manufacturer allows.
✓ Your Telegram / WhatsApp bot
- ✓Open: you see, understand, and change every part of the code.
- ✓Your persona: sets its tone, values, and personality (its soul, in SOUL.md).
- ✓Brain of your choice: cloud or local, you decide where the data lives.
- ✓Available tools: connect your calendar, email, whatever you want, via MCP.
⚠️ The point no one talks about
With a built-in assistant, if the company changes its mind, disables a feature, or raises the price, the problem is yours and there's nothing you can do. With your bot, you own the rules. Convenient doesn't mean free—and this learning path is about a convenient assistant e free.
Key concepts
Siri/Gemini/Alexa: closed and controlled by the manufacturer.
You can’t see or control what happens inside.
Your bot is yours: code, persona, data, and tools.
You can have both—an assistant that's easy to use and truly yours.
Self-check (optional): in the “pocket Jarvis,” where does the agent actually run?
🎯 Module summary
Next module:
5.2 — Voice and on-device: speaking, listening, and running locally