PTENES
MODULE 3-1

📡 Channels — how you talk to it

The first of the Anatomy’s six layers is the easiest to understand and the easiest to get wrong: the door. Before Jarvis has a brain, memory, or hands, he needs a place where you can reach him and he can answer you. In this module, you’ll learn what a channel is, why Telegram became the favorite, what a whitelist is (the first security barrier), and how the same brain reaches any device.

6
Topics
~35
Minutes
Intermediate
Level
Practical
Type
1

📨 What is a channel

Imagine your Jarvis as a very capable person locked in a room. They think, remember, know how to use tools — but you can only talk to them if there’s a door. A channel and that's exactly what it is: the entry and exit point where you talk to the assistant and it responds to you. Without a channel, even the smartest brain in the world stays silent.

In practice, a channel is almost always an app you already use: a messaging app (Telegram, WhatsApp), email, a web page, or even the developer's own terminal/Claude Code. The important point is that the channel isn't the brain—it just carries messages. The same "person in the room" can have several doors at once.

New here? Channel = the communication channel between you and Jarvis (Telegram, WhatsApp, voice, web, terminal). Brain = the AI model that thinks (we saw it in Track 1). The channel carries the message; the brain decides how to respond. They’re separate things — and that’s great, because you can change one without touching the other.

🔌 The channel is the outlet, not the device

Think of an electrical outlet in the wall. It doesn't "do" anything on its own—but it's how power reaches the blender, charger, or TV. The channel is the outlet: it carries your message to the brain and brings the response back. You can plug several devices (channels) into the same Jarvis.

  • •Input: you send text, audio, or a photo through the channel.
  • •Output: Jarvis responds through that same channel.
  • •Independence: switching channels doesn't change who it is — it just changes the doorway.
ONE BRAIN · MANY CHANNELS 🤖 Jarvis the brain ✈️ Telegram 💬 WhatsApp 🎙️ Voice 🌐 Web / terminal

At the center, the Jarvis's brain; at the ends, the channels. The diagram’s lesson: you don’t build several Jarvises — you connect several doors to the same brain. Each ray is a door; the brain is always the same.

Key concepts

Channel

The entry and exit point where you talk to Jarvis.

Input and output (I/O)

You send through the door; it responds through the same door.

Channel ≠ brain

The door transports; the model thinks. Changing one doesn’t affect the other.

Multi-channel

The same Jarvis can have several doors open at once.

2

✈️ Why Telegram is the favorite

Among all the possible channels, almost all the projects in our family (GravityClaw, Intelecto, agentejax) choose Telegram as the main gateway. It's not a trend or personal preference: it's an engineering and security decision. Telegram solves, for free, several problems that other channels hand you.

New here? Token and a long, unique password that Telegram gives you for your bot—whoever has the token "is" the bot. Long-polling ("polling") is when YOUR program periodically asks Telegram, "is there a new message?" instead of Telegram knocking on your door. The difference seems small, but it’s the key to security — I’ll explain in a moment.

✓ Why Telegram shines

  • ✓Only one token: get it from @BotFather in 2 minutes, with no company registration.
  • ✓Long-polling: your program POLLS Telegram; you don’t need to open any ports on your PC.
  • ✓No exposed web server: nothing from the internet can "reach" your machine.
  • ✓Free and cross-platform: PC, phone, web — it already runs everywhere.
  • ✓"Already mobile": you don’t need to publish any app.

✗ What a “web server” charges

  • ✗You need to open a public port so the internet can reach you.
  • ✗That door becomes a target: anyone can knock on it.
  • ✗Requires a domain, certificate, and configuration — real friction.
  • ✗That was exactly the vulnerability in OpenClaw (we'll look at the numbers).

📊 The data that settles the discussion

In the Overview (Track 2), we saw that OpenClaw, the original system, exposed a web server. The result was a security disaster: researchers found 42,665 public instances on the internet, e 93.4% of them with no password — anyone could talk to someone else’s “Jarvis.”

The lean family’s reaction (GravityClaw and company) was direct: Telegram only, with long polling and zero open ports. If no port is exposed, there’s nothing to scan. The channel choice has become a security decision.

💡 Practical tip

The mental rule for choosing a channel: “does my PC need to be accessible from outside?” If the answer is no, you’re on the safe path. Telegram with long-polling gives you that for free—so it’s the recommended starting point, and why almost every project in the family starts there.

Key concepts

Token (@BotFather)

Your bot’s unique password, created through Telegram’s official bot.

Long-polling

Your program asks Telegram; nobody knocks on your door.

No open ports

Nothing from the internet can reach your machine—minimal attack surface.

"Already mobile"

Telegram runs everywhere; you don’t publish an app to an app store.

3

💬 WhatsApp and the rest

Telegram is the starting point, but it isn’t the only destination. In Brazil in particular, many people live on WhatsApp — and it makes sense to want Jarvis there. The project agentevoz connects the assistant to WhatsApp through a tool called Evolution API, which connects WhatsApp to your code. Text, images, and even audio pass through the same gateway.

Besides these two, there are other options: email (Jarvis reads and replies to your inbox), a web page (a chat in the browser) and terminal / Claude Code (the developer's channel). The big idea behind the channel layer is that these paths are interchangeable and composable: the same brain can serve several at once.

New here? WhatsApp doesn’t offer a simple “bot token” like Telegram. The Evolution API and an intermediate program (open source) that connects to WhatsApp and exposes an interface your Jarvis understands. It's a bridge—you get WhatsApp's reach, but add one more component to maintain.

Channel Ease Default security When to use
✈️ Telegram Very high (token only) Excellent (zero ports) Always the starting point.
💬 WhatsApp Media (requires Evolution API) Depends on the bridge When your audience lives on WhatsApp.
📧 Email Media Good (asynchronous) Tasks that aren’t urgent.
🌐 Web Low (requires a server) Requires caution (open port) When it needs a visual interface.
⌨️ Terminal High (for developers only) Excellent (local) Development and testing.

📊 The advantage of multiple channels

Since the channel is separate from the brain, adding entry points is easy: the same Jarvis, with the same memory and personality, can serve you on Telegram at home and WhatsApp at work. You don't duplicate the assistant—you just open another door to the same room. This "one brain, many mouths" design is at the heart of Anatomy.

Key concepts

Evolution API

The open-source bridge that connects WhatsApp to your Jarvis.

Interchangeable channels

Changing the port doesn't change who the assistant is.

Multi-channel

Several doors to the same brain, at the same time.

Cost of each bridge

More ports = more reach, but more pieces to maintain.

4

🚪 Whitelist = the first lock

Here’s the most important security concept in this module. A Telegram bot is, by default, public: anyone who finds out its name can start chatting with it. If this bot is your Jarvis — with access to your calendar, emails, and files — that's dangerous. The solution is simple and powerful: the whitelist (allowlist).

New here? Every Telegram user has a Numeric ID single (for example, 123456789). A whitelist and a short list of IDs that Jarvis is allowed to serve. If a message comes from an ID that is NOT on the list, Jarvis simply ignores it—in silence, without even replying "access denied." The opposite of a whitelist would be a blacklist (block some); this approach is safer: blocks everyone, allows only the people you chose.

THE WHITELIST DECIDES WHO GETS THROUGH messageany ID ID is in thewhitelist? YES ✓ Jarvis thinks and responds NO ✗ silently ignores

Every message first passes through the whitelist check: known ID continues to the brain; strange ID and discarded without a reply. That’s why it’s the “1st guardrail” — it happens before before the model is even invoked, and silently ignoring it gives an attacker no clues.

⚠️ Without a whitelist, the damage is real

Remember the 93.4% of instances with no authentication of OpenClaw: in practice, they were open "Jarvis" assistants anyone could use. Without the first safeguard, a stranger could ask your assistant to read your emails, spend your AI credits, or run commands on your machine. The whitelist is the line between "my personal assistant" and "a robot that obeys anyone."

⌨️ Copy-run · create the bot and lock down the whitelist font-mono

# GOAL: create your Telegram bot, find your ID, and

# leave only YOU on the allowlist (the first safeguard).

# ------------------------------------------------------------

# STEP 1 — in the Telegram app, talk to @BotFather and type:

/newbot

# it asks for a name and a @username; at the end, it gives you the token:

8123456789:AAH<rest-of-your-secret-token>

# STEP 2 — find YOUR ID: talk to @userinfobot; it replies:

ID: <your-numeric-id> (e.g., 123456789)

# STEP 3 — put both in your Jarvis .env file:

TELEGRAM_BOT_TOKEN=<your-botfather-token>

TELEGRAM_ALLOWED_IDS=<your-numeric-id>

# (to allow someone else, separate with a comma: 123456789,987654321)

How to check: send "hi" to your bot from YOUR account—it replies. Ask a friend to send "hi" through the same bot—they doesn't answer anything. If your “hi” works and the stranger’s is silently ignored, the whitelist is in place. Treat the token like a password: whoever has it controls the bot.

Key concepts

Whitelist

Short list of IDs Jarvis accepts; everything else is ignored.

User ID

The unique number that identifies you on Telegram.

Ignore silently

Don't answer the stranger—or give attackers clues.

Secrets in .env

Token and IDs are stored in a secrets file, outside the code.

5

🎙️ Text, voice, and media

A channel carries more than just text. Telegram, for example, transports text, audio, and image through the same door. When you record a voice message, the app sends an audio file (format .ogg); when you send a photo, it delivers the image. The channel is the same — only the type of "payload" passing through it changes.

Here’s a phrase worth keeping: "voice is an interface, not a brain". The model keeps thinking in text. Voice is just a way to get input and output — an optional layer that can be turned on or off at runtime (in real time, depending on the situation). Talking to Jarvis or typing to him takes you to the same “room”: the brain doesn't change.

New here? For the model to “hear” your voice, the audio goes through a STT ("speech-to-text," speech becomes text). To "speak," its response text goes through a TTS ("text-to-speech," text becomes speech). In between, the brain only sees text. "At runtime" means the decision to use voice is made on the spot — you don’t need to decide that when building the system.

1

You record an audio message

Telegram delivers a file .ogg through the same door as always.

2

STT transcribes (speech → text)

A transcription tool (e.g., Whisper) converts the audio into text in Portuguese.

3

The brain thinks in text

The model receives text and returns text — exactly as if you had typed it.

4

TTS responds with voice (optional)

If voice is enabled, the reply text becomes audio and comes back through the door. Otherwise, it responds with text.

💡 Practical tip

Start with text only. Voice is more "wow," but it adds components (transcription, audio generation, cost). The healthy path is: text first, then turn on voice once it's working. And remember that voice often uses different services from the brain—be careful not to mix up their keys (the transcription service's key isn't always the same as the model's).

Key concepts

Multimodal

The same port carries text, audio, and images.

STT / TTS

Speech becomes text (STT); text becomes speech (TTS).

Voice and interface

The brain thinks in text; voice just goes in and out.

Decision at runtime

Connecting voice is an optional choice made in the moment.

6

📱 One channel, many devices

There’s one phrase that appears in almost every project: "Telegram is already mobile". It hides one of the biggest advantages of the channel layer. Since Telegram (and WhatsApp) already runs on your phone, tablet, laptop, and the web, choosing this channel gives you access to Jarvis of any device — without you having to build, publish, and maintain your own app in the store.

The detail that confuses beginners: Jarvis’s brain doesn’t “live” on your phone. It runs somewhere fixed—on your home PC when it’s on, or on a cheap cloud server—and the phone is just another screen that connects to it through the channel. You start a conversation on your PC, continue on your phone on the bus, and finish on your tablet at night: it's always the same room, the same memory, the same assistant.

✓ What “Telegram is already mobile” means

  • ✓Access from phone, PC, tablet, and web — all already available.
  • ✓No app store app to publish or approve.
  • ✓Same conversation and memory on any device.
  • ✓The brain stays in one fixed place; the device is just the screen.

✗ If it were a native app

  • ✗You’d have to develop for Android AND iOS.
  • ✗Go through app store approval (and fees).
  • ✗Ask the user to download and update.
  • ✗Maintain all of this yourself, forever.

🧭 Where this takes you

This idea is the bridge to Track 5 ("Jarvis on your phone"): there is no "Jarvis app"—there’s a brain running somewhere and a channel (Telegram/WhatsApp) you already have in your pocket. Your phone becomes the remote control for an assistant that can run 24 hours a day, even with your laptop closed.

Save: channel = where you talk; brain = where it thinks. Separating the two is what lets Jarvis follow you across all your devices without turning into six different apps.

Key concepts

"Already mobile"

The channel already works on every device—no dedicated app needed.

Fixed brain, mobile screen

The model runs somewhere else; the device just connects.

Continuity

Same conversation and memory on any device.

No app store app

Nothing to publish, approve, or ask to download.

Self-check (optional): why do so many projects choose Telegram as their main channel?

🎯 Module summary

✓
Channel = the door — where you talk to Jarvis; separate from the brain, which only thinks.
✓
Telegram is the preferred choice — just one token, long polling, zero open ports; free and already mobile.
✓
WhatsApp and the rest — WhatsApp via Evolution API, email, web, terminal: add-on channels for the same brain.
✓
Whitelist = the first safeguard — only your ID is served; everything else is silently ignored before the model acts.
✓
Text, voice, media, and multiple devices — the same port carries everything; “voice and interface,” and the channel follows you to any device.

Next module:

Module 3-2: Identity — memory, persona, and who it is