📱 Jarvis on your phone
Your assistant in your pocket, without downloading an app from the store: it lives inside a messaging app you already use (Telegram or WhatsApp). Here you’ll see the mobile channels, how it listens and speaks (voice), what truly running locally means — and build your pocket Jarvis step by step.
Read from left to right: the cell phone and is just the channel (a messaging app). The brain runs in the cloud or on your PC and gives you access to voice, memory, and tools — resulting in a 24/7 assistant in your pocket, without installing any app from the store.
Learning path map
Detailed content
📲 Where Jarvis lives on your phone (the channels)
The shift: you don’t install an app from the store. You use Telegram or WhatsApp, which already runs on your phone, as the gateway to an agent that lives in the cloud or on your PC—24/7, private, and yours.
You don’t need to publish an app on the App Store or Play Store. Your Jarvis lives inside a messaging app you already have installed—and you chat with it like you would with a contact.
It’s the shortcut that makes an “assistant in your pocket” possible today, without becoming an app developer, getting store approval, or paying publishing fees.
“Telegram is already mobile,” a messaging app as a channel, zero native apps.
On Telegram, you talk to the @BotFather (an official bot), it gives you a token (your bot’s password), and you put your user ID in a whitelist — the list of people the bot serves. The reference project is called agentejax.
It’s the preferred approach: just one token, no exposed web server (it uses long polling), free, and accessible from any device where you have Telegram.
@BotFather, token, whitelist (user ID), long-polling.
On WhatsApp, the connection usually goes through Evolution API (a bridge that connects your number to the agent). One Policy Engine (rules engine) decides whether to respond by text, by voice, or hand it off to a human. The reference project is called agentevoz.
WhatsApp is where most people already are — ideal for multichannel support, with text, images, and audio in the same place.
Evolution API, Policy Engine, multichannel, text/voice/human decisions.
The phone is just the channel; the brain runs somewhere else. It could be on a cloud platform (e.g., Railway) or on your own computer at home, running 24 hours a day.
Separating the “channel” from the “brain” is the central idea: you leave home with your phone, but the agent keeps working 24/7 wherever you host it.
Channel ≠ brain, hosting (Railway/PC), 24/7 execution.
You can keep the brain 100% local—with a model running via Ollama on your home PC — and use only the channel on your phone. Your conversations are processed on your machine, not in a third-party cloud.
It’s the best of both worlds: the convenience of your phone with local privacy—you choose where your data stays.
Local Ollama, local-first, a brain at home + a channel in your pocket.
The assistants already on your phone (Siri, Gemini, Alexa) are black boxes: you don't control the memory, persona, tools, or where the data ends up.
It shows you what you gain by using your own bot: complete control. The Telegram/WhatsApp bot is YOURS — you define everything.
Closed assistant, black box, control/ownership of your bot.
🎙️ Voice and on-device: speak, listen, and run locally
How Jarvis listens (STT), how it speaks (TTS), and why voice is just an interface—not the brain. Plus: what it means to run "truly on-device" with Ollama and how far real time goes.
Voice is just the way in and out: audio becomes text on input, and text becomes audio on output. The orchestrator (the agent) does the thinking. "Voice is the interface, the orchestrator is the brain; speaking is an optional output decided in the moment."
Avoids the common mistake of thinking that "voice" is a different AI. It’s the same brain as always, just with a microphone and speaker plugged in.
Interface vs. brain, orchestrator, optional runtime output.
STT = Speech-to-Text (speech to text). Telegram delivers the audio as a file .ogg; o ffmpeg converts the format; and the Whisper (transcription model) writes down what you said, in Portuguese.
It’s the first step for any voice feature. One practical detail: Whisper usually uses an OpenAI key, which is different from the Anthropic key.
STT, ffmpeg, Whisper, .ogg, OpenAI key ≠ Anthropic.
TTS = Text-to-Speech (text to speech). A voice model (e.g., tts-1 with the “nova” voice, or a TTS running locally) takes the text response and generates the audio that Jarvis “speaks” back.
It’s what completes the voice conversation loop—you speak, it understands, thinks, and responds out loud.
TTS, voice model, cloud or local voice.
O Policy Engine decides, with each response, whether to output text, speak, or call a human—and also whether to use a low-cost or premium model, depending on how difficult the task is.
It’s what saves money: voice and premium models cost more; using them only when it’s worth it keeps the bill low without sacrificing quality where it matters.
Policy Engine, per-response decisions, cheap vs. premium.
On-device = run the model on the device/PC itself, without the cloud. Via Ollama, the right model depends on your RAM (e.g., ~3B with 8 GB, ~8B with 16 GB). On CPU alone, responses are slow (30-60s); a GPU or cloud speeds things up.
Sets expectations: “local” is real and private, but it comes at the cost of speed. Knowing that prevents frustration and helps you choose the right model.
On-device, Ollama, model based on RAM (3B/8B), CPU vs. GPU, latency.
The voice roadmap goes from text (v1) to asynchronous voice (v2, Whisper + TTS) to real-time phone calls (v3), with services like Twilio or LiveKit and latency below 1.5 seconds.
Shows the upper limit of what's possible — and that talking to Jarvis "live" by phone is advanced, not the starting point. Set expectations.
Real time, Twilio/LiveKit, latency <1,5s, roadmap v1→v3.
🛠️ Building your pocket Jarvis (project)
The guided end-to-end project: from a Telegram bot to voice and cadence. Five steps to go from “nothing” to a personal assistant with memory and a useful skill, accessible from your phone.
The goal: a personal Telegram bot with memory and a useful skill that you can access from your phone. Built in 5 steps, one brick at a time ("brick-by-brick").
Having the full picture before you start keeps you from getting lost along the way and shows how the pieces from the previous tracks fit together in a real product.
Project scope, brick by brick, recap of the pieces (channel/brain/memory/skill).
Create the bot in @BotFather, copy the token and add your user ID to the whitelist — exactly what you saw in module 5.1, now in practice.
It’s “Level 1 / Foundation”: without the channel up and running, there’s nothing to test. Everything else builds on this first brick.
@BotFather, token, whitelist, Foundation (Level 1).
Choose the model that thinks: cloud (powerful, pay per use) or Ollama local (free, private). The choice lives in a file .env, without touching the code.
This is where the bot stops being an “echo” and starts responding for real. And changing the brain later is as simple as editing one line.
Brain (LLM model), cloud vs. Ollama, hot-swap via .env.
Add a soul.md (the personality: name, tone, values) and a database SQLite to store the conversation. From here on, it remembers who you are and what’s already been said.
It’s the leap from “a stranger who responds” to “an assistant who knows you”—memory creates continuity between conversations.
soul.md, persistent memory, SQLite, Memory (Level 2).
Plug in a skill (a packaged recipe) that does something concrete—for example, “daily summary” pulling your calendar via MCP (the “USB” for AI tools).
This is when the bot stops just chatting and starts doing something useful for you—the first real capability of your pocket Jarvis.
Skill, MCP, tool (agenda), Tools/MCP (Level 4).
Connect voice (audio message via Whisper/TTS, from module 5.2) and the cadence: a heartbeat that runs on its own — for example, a summary every day at 7 a.m. Then, how to test everything end to end.
Brings the project to a close: Jarvis stops just reacting and starts acting on its own at the right time — and you confirm that each building block works.
Voice (Whisper/TTS), cadence, heartbeat/cron, end-to-end test.