PTENES
TRACK 5

📱 Jarvis on your phone

Your assistant in your pocket, without downloading an app from the store: it lives inside a messaging app you already use (Telegram or WhatsApp). Here you’ll see the mobile channels, how it listens and speaks (voice), what truly running locally means — and build your pocket Jarvis step by step.

Telegram your phone only the channel Agent (brain) cloud (Railway) or your PC voice (listen/speak) memory + soul tools / skill 24/7 in your pocket without a store app

Read from left to right: the cell phone and is just the channel (a messaging app). The brain runs in the cloud or on your PC and gives you access to voice, memory, and tools — resulting in a 24/7 assistant in your pocket, without installing any app from the store.

3
Modules
18
Topics
~2h30
Duration
Intermediate
Level
Path progress0%
0 of 18 topics

Learning path map

Detailed content

5.1~50 min

📲 Where Jarvis lives on your phone (the channels)

The shift: you don’t install an app from the store. You use Telegram or WhatsApp, which already runs on your phone, as the gateway to an agent that lives in the cloud or on your PC—24/7, private, and yours.

0 of 6 · 0%
What it is:

You don’t need to publish an app on the App Store or Play Store. Your Jarvis lives inside a messaging app you already have installed—and you chat with it like you would with a contact.

Why learn:

It’s the shortcut that makes an “assistant in your pocket” possible today, without becoming an app developer, getting store approval, or paying publishing fees.

Key concepts:

“Telegram is already mobile,” a messaging app as a channel, zero native apps.

What it is:

On Telegram, you talk to the @BotFather (an official bot), it gives you a token (your bot’s password), and you put your user ID in a whitelist — the list of people the bot serves. The reference project is called agentejax.

Why learn:

It’s the preferred approach: just one token, no exposed web server (it uses long polling), free, and accessible from any device where you have Telegram.

Key concepts:

@BotFather, token, whitelist (user ID), long-polling.

What it is:

On WhatsApp, the connection usually goes through Evolution API (a bridge that connects your number to the agent). One Policy Engine (rules engine) decides whether to respond by text, by voice, or hand it off to a human. The reference project is called agentevoz.

Why learn:

WhatsApp is where most people already are — ideal for multichannel support, with text, images, and audio in the same place.

Key concepts:

Evolution API, Policy Engine, multichannel, text/voice/human decisions.

What it is:

The phone is just the channel; the brain runs somewhere else. It could be on a cloud platform (e.g., Railway) or on your own computer at home, running 24 hours a day.

Why learn:

Separating the “channel” from the “brain” is the central idea: you leave home with your phone, but the agent keeps working 24/7 wherever you host it.

Key concepts:

Channel ≠ brain, hosting (Railway/PC), 24/7 execution.

What it is:

You can keep the brain 100% local—with a model running via Ollama on your home PC — and use only the channel on your phone. Your conversations are processed on your machine, not in a third-party cloud.

Why learn:

It’s the best of both worlds: the convenience of your phone with local privacy—you choose where your data stays.

Key concepts:

Local Ollama, local-first, a brain at home + a channel in your pocket.

What it is:

The assistants already on your phone (Siri, Gemini, Alexa) are black boxes: you don't control the memory, persona, tools, or where the data ends up.

Why learn:

It shows you what you gain by using your own bot: complete control. The Telegram/WhatsApp bot is YOURS — you define everything.

Key concepts:

Closed assistant, black box, control/ownership of your bot.

View Full
5.2~50 min

🎙️ Voice and on-device: speak, listen, and run locally

How Jarvis listens (STT), how it speaks (TTS), and why voice is just an interface—not the brain. Plus: what it means to run "truly on-device" with Ollama and how far real time goes.

0 of 6 · 0%
What it is:

Voice is just the way in and out: audio becomes text on input, and text becomes audio on output. The orchestrator (the agent) does the thinking. "Voice is the interface, the orchestrator is the brain; speaking is an optional output decided in the moment."

Why learn:

Avoids the common mistake of thinking that "voice" is a different AI. It’s the same brain as always, just with a microphone and speaker plugged in.

Key concepts:

Interface vs. brain, orchestrator, optional runtime output.

What it is:

STT = Speech-to-Text (speech to text). Telegram delivers the audio as a file .ogg; o ffmpeg converts the format; and the Whisper (transcription model) writes down what you said, in Portuguese.

Why learn:

It’s the first step for any voice feature. One practical detail: Whisper usually uses an OpenAI key, which is different from the Anthropic key.

Key concepts:

STT, ffmpeg, Whisper, .ogg, OpenAI key ≠ Anthropic.

What it is:

TTS = Text-to-Speech (text to speech). A voice model (e.g., tts-1 with the “nova” voice, or a TTS running locally) takes the text response and generates the audio that Jarvis “speaks” back.

Why learn:

It’s what completes the voice conversation loop—you speak, it understands, thinks, and responds out loud.

Key concepts:

TTS, voice model, cloud or local voice.

What it is:

O Policy Engine decides, with each response, whether to output text, speak, or call a human—and also whether to use a low-cost or premium model, depending on how difficult the task is.

Why learn:

It’s what saves money: voice and premium models cost more; using them only when it’s worth it keeps the bill low without sacrificing quality where it matters.

Key concepts:

Policy Engine, per-response decisions, cheap vs. premium.

What it is:

On-device = run the model on the device/PC itself, without the cloud. Via Ollama, the right model depends on your RAM (e.g., ~3B with 8 GB, ~8B with 16 GB). On CPU alone, responses are slow (30-60s); a GPU or cloud speeds things up.

Why learn:

Sets expectations: “local” is real and private, but it comes at the cost of speed. Knowing that prevents frustration and helps you choose the right model.

Key concepts:

On-device, Ollama, model based on RAM (3B/8B), CPU vs. GPU, latency.

What it is:

The voice roadmap goes from text (v1) to asynchronous voice (v2, Whisper + TTS) to real-time phone calls (v3), with services like Twilio or LiveKit and latency below 1.5 seconds.

Why learn:

Shows the upper limit of what's possible — and that talking to Jarvis "live" by phone is advanced, not the starting point. Set expectations.

Key concepts:

Real time, Twilio/LiveKit, latency <1,5s, roadmap v1→v3.

View Full
5.3~50 min

🛠️ Building your pocket Jarvis (project)

The guided end-to-end project: from a Telegram bot to voice and cadence. Five steps to go from “nothing” to a personal assistant with memory and a useful skill, accessible from your phone.

0 of 6 · 0%
What it is:

The goal: a personal Telegram bot with memory and a useful skill that you can access from your phone. Built in 5 steps, one brick at a time ("brick-by-brick").

Why learn:

Having the full picture before you start keeps you from getting lost along the way and shows how the pieces from the previous tracks fit together in a real product.

Key concepts:

Project scope, brick by brick, recap of the pieces (channel/brain/memory/skill).

What it is:

Create the bot in @BotFather, copy the token and add your user ID to the whitelist — exactly what you saw in module 5.1, now in practice.

Why learn:

It’s “Level 1 / Foundation”: without the channel up and running, there’s nothing to test. Everything else builds on this first brick.

Key concepts:

@BotFather, token, whitelist, Foundation (Level 1).

What it is:

Choose the model that thinks: cloud (powerful, pay per use) or Ollama local (free, private). The choice lives in a file .env, without touching the code.

Why learn:

This is where the bot stops being an “echo” and starts responding for real. And changing the brain later is as simple as editing one line.

Key concepts:

Brain (LLM model), cloud vs. Ollama, hot-swap via .env.

What it is:

Add a soul.md (the personality: name, tone, values) and a database SQLite to store the conversation. From here on, it remembers who you are and what’s already been said.

Why learn:

It’s the leap from “a stranger who responds” to “an assistant who knows you”—memory creates continuity between conversations.

Key concepts:

soul.md, persistent memory, SQLite, Memory (Level 2).

What it is:

Plug in a skill (a packaged recipe) that does something concrete—for example, “daily summary” pulling your calendar via MCP (the “USB” for AI tools).

Why learn:

This is when the bot stops just chatting and starts doing something useful for you—the first real capability of your pocket Jarvis.

Key concepts:

Skill, MCP, tool (agenda), Tools/MCP (Level 4).

What it is:

Connect voice (audio message via Whisper/TTS, from module 5.2) and the cadence: a heartbeat that runs on its own — for example, a summary every day at 7 a.m. Then, how to test everything end to end.

Why learn:

Brings the project to a close: Jarvis stops just reacting and starts acting on its own at the right time — and you confirm that each building block works.

Key concepts:

Voice (Whisper/TTS), cadence, heartbeat/cron, end-to-end test.

View Full