PTENES
MODULE 5-2

🎙️ Voice and on-device: speak, listen, and run locally

Sending audio to Jarvis and hearing it respond in a voice feels like magic—but it's just a chain of simple pieces: audio becomes text, text becomes a response, and the response becomes audio. In this module, you build this pipeline in your mind, understand why voice is just the "door" (and not the brain), and discover how to run the AI engine at home, on your own computer.

6
Topics
~35
Minutes
Intermediate
Level
Practical
Type
1

🎚️ Voice and interface, not the brain

Before any code, here’s one sentence that organizes the entire module: "voice is an interface; the orchestrator is the brain; speaking is an optional output decided on the spot". The voice doesn’t think. It only carries what you say in and brings the answer out. The one who thinks—who decides what to answer, which tool to use, what to remember—is the same orchestrator of text you already know from the other tracks.

This completely changes how you think about the project. Voice isn’t “a different AI”: it’s a shell around the Jarvis you already have. You build the brain first (text), and only then plug in ears and a mouth. If the voice breaks, the text assistant keeps working.

New here? A orchestrator and the "conductor" of Jarvis: the piece of code that receives your message, talks to the AI model, calls tools (calendar, search, email), and puts together the response. The interface and it’s just the channel this goes in and out through—it can be text on Telegram, voice, or both. The brain is always the same; only the door changes.

🧩 Three layers, separate roles

  • •Voice input (ears): turns the audio you sent into text. Optional.
  • •Orchestrator (brain): reads the text, thinks, acts, decides on the response. Always present.
  • •Voice output (mouth): turns the response into audio. Optional, decided for each message.

📊 Why “optional” matters

You don’t always want to hear Jarvis speak. A shopping list is better as text (you can scan it at a glance); a "reminder while I’m driving" is better by voice. Since speaking is optional and decided in the moment, the same assistant works for both situations without you switching apps.

Key concepts

Interface

The door through which the conversation comes and goes (text or voice).

Orchestrator

The brain that thinks, acts, and decides on the response.

Optional output

Speaking is chosen for each message; it isn’t always on.

Voice shell

Voice wraps around the text assistant; it doesn’t replace it.

2

👂 Listen: from audio to text (STT)

When you hold down the microphone in Telegram and send an audio message, the app delivers a file in a format called .ogg (with compression Opus)—great for sending over the network, but transcribers don’t always accept it directly. The “listen” path has three steps: receive the audio → convert the format → transcribe it to text. Only after that text is ready does the brain come in.

The voice pipeline, end to end 🎤 audio .ogg / Opus ffmpeg → .wav 16kHz Whisper STT (pt) orchestrator LLM + tools "the brain" TTS text → voice 🔊 response audio — LISTEN (STT) — — SPEAK (TTS) —

The complete path: audio comes in on the left, the brain thinks in the middle, and the voice output on the right. Notice that the LLM at the center is the same one that would respond in text — voice just adds the endpoints.

The most widely used transcription engine is called Whisper (from OpenAI). You send it the converted audio, and it returns the text, choosing the language (in our case, pt in Portuguese). A gotcha that trips up a lot of people: the Whisper API key is the OpenAI key (OPENAI_API_KEY), which is DIFFERENT from the Anthropic key (ANTHROPIC_API_KEY) used for the brain. Two companies, two keys, two accounts.

⚠️ The classic case of swapped keys

If you put the Anthropic key where Whisper expects the OpenAI key (or vice versa), transcription fails with an “invalid authentication” error, and it looks like everything broke—when it’s just the wrong key in the wrong box. Keep this in mind: listening/speaking through OpenAI = OPENAI_API_KEY; thinking through Anthropic = ANTHROPIC_API_KEY. (Whisper also runs locally and for free, with no key required — see topic 5.)

Key concepts

STT

“Speech-to-Text”: turning speech into text.

.ogg / Opus

The audio format Telegram provides.

ffmpeg

The tool that converts audio from one format to another.

Whisper

The model that transcribes audio into text.

3

🗣️ Speaking: from text to audio (TTS)

The way back mirrors "listening." The brain produces the response as text; a voice model reads that text aloud and generates an audio file that Telegram plays back to you. This is called TTS — "Text-to-Speech," text converted to speech. You still choose which voice to use: deep, high, warm, neutral.

Here you have two broad families of options, and the choice echoes the theme of the entire course (cloud vs. local). A cloud (for example, the model tts-1 from OpenAI, with voices named like "nova") sounds very natural and doesn't require hardware — but it costs per use and sends the text outside. A voice local runs on your computer, is free and private, but tends to sound a little more robotic and puts a strain on your machine.

✓ Cloud voice (e.g.: tts-1 / “nova”)

  • ✓Sounds very natural, almost human.
  • ✓It doesn't require anything from your hardware.
  • ✓Several voices ready to choose from.
  • ✗It costs per use, and the text stays on your machine.

✓ Local voice (runs on your PC)

  • ✓Free after installation, forever.
  • ✓Private: the text never leaves your home.
  • ✓Works offline, without the internet.
  • ✗A slightly more robotic voice that puts more strain on the machine.

💡 Practical tip

Start with cloud voice just to see the result working from end to end—it's easier to plug in. Once the whole pipeline is up and running, swap the voice for the local version if privacy or cost matters. Remember: changing the voice doesn't affect the brain; it's just the part at the end.

Key concepts

TTS

“Text-to-Speech”: turning text into speech.

Voice (preset)

The selected voice—for example, "nova" in tts-1.

Cloud TTS

Natural and requires no hardware, but costs money and isn't private.

Local TTS

Free, private, and offline; sounds more robotic.

4

🚦 Decision by response: the Policy Engine

If the mouth is optional and the brain can be cheap or expensive, someone needs to decide for each message: should I respond in text or voice? Use the cheaper model or the premium one? Forward it to a human? This "decision gatekeeper" has a name in the ecosystem: Policy Engine (policy engine). It looks at each request and chooses the least expensive path that still solves it well.

The engine behind every answer request (text or voice) Policy Engine “is it simple? is it urgent?” channel: text vs. voice list = text · "in the car" = voice model: low-cost vs. premium trivial = cheap · difficult = premium escalates to a human sensitive case = real person right answer at the right cost

O Policy Engine looks at each request and decides three things: how to respond (text or voice), with which model (low-cost or premium) and whether you need a human. Result: the right answer at the right cost.

The logic is just economic common sense. “What time is it?” doesn’t need the world’s most expensive model or a produced voice—short text, a cheap model, done. But “help me rewrite this difficult email to my boss” deserves the premium brain. Making the choice for each response, rather than once and for all, is what keeps the bill low without sacrificing quality where it matters.

Key concepts

Policy Engine

The “gatekeeper” that decides where each message goes.

Decision for each response

Choose case by case, not with a fixed config.

Budget vs. premium

Simple model for the easy stuff, powerful model for the hard stuff.

Escalate to a human

Sensitive cases go to a real person.

5

🏠 Truly on-device: the local brain with Ollama

So far, the brain has lived in the cloud. But you can put the AI engine inside your home, on your own computer. The standard tool for this is called Ollama: you install it, download an open model, and it responds—for free, privately, and offline. The term for this is on-device: the intelligence runs on the device, not on a distant server.

📊 Which model fits your hardware (the RAM rule)

Models are measured in "B" (billions of parameters — the “size of the brain”). Bigger = smarter, but heavier. The practical rule uses your RAM (the computer’s working memory):

  • •~3B (e.g., llama3.2): runs comfortably with 8 GB of RAM.
  • •~8B: asks for 16 GB of RAM so it won’t bog down.
  • •Without a graphics card (only CPU): it works, but it’s slow— 30 to 60 seconds per response. A GPU (graphics card) or the cloud speed things up considerably.

The beauty of the course design is that switching the brain from cloud to local doesn’t mean rewriting Jarvis. You edit a configuration file called .env (where the project’s “keys and settings” live) and points the provider to Ollama. The rest of the system — channels, memory, voice — doesn’t even notice. This is the hot-swap: swap the engine without taking the car apart.

COPY-RUN · listen to audio + run the local brain terminal + .env

Objective: convert a Telegram audio message to text with Whisper and, in parallel, connect the local brain (Ollama) by changing one line of the .env. Nothing leaves your machine.

// 1) install the local engine and download a model that fits in your RAM

curl -fsSL https://ollama.com/install.sh | sh
ollama pull <modelo-que-cabe-na-sua-RAM>     # ex.: llama3.2 (8GB) ou llama3.1:8b (16GB)
ollama serve                                  # deixa o motor local ouvindo (porta 11434)

// 2) the audio arrived as .ogg/Opus -> convert it to .wav 16kHz with ffmpeg

ffmpeg -i <seu-audio>.ogg -ar 16000 -ac 1 voz.wav

# transcreva para texto (Whisper local, gratis, idioma pt) — sem chave nenhuma:
whisper voz.wav --language pt --model small --output_format txt

// 3) point Jarvis’s brain to Ollama in the .env file (hot-swap)

LLM_PROVIDER=ollama
LLM_MODEL=<o-mesmo-modelo-que-voce-baixou>
OLLAMA_HOST=http://localhost:11434

# (se um dia quiser voltar pra nuvem, troque so estas linhas — o resto fica igual)

How to check:

  • • A file appeared voz.txt with the right transcript of what you said? Listening (STT) works.
  • • Run ollama run <seu-modelo> "diga oi" — if it responds in the terminal, the local brain is up and running.
  • • Restart Jarvis: no traffic should go out to the internet in its first response (check offline, in airplane mode)—a sign that the brain has truly gone local.

Key concepts

On-device

AI runs on your device, not on a server.

Ollama

The tool that runs open models locally.

The RAM rule

3B@8GB, 8B@16GB; CPU is slow, GPU speeds things up.

Hot-swap (.env)

Switch from cloud to local by changing only the config.

6

📞 Real time (advanced): from voice message to phone call

Everything we've seen so far is voice asynchronous: you send an audio message, wait, and receive another audio message. It’s like exchanging voice messages on WhatsApp. The next dream is voice in real time — a phone call for real with Jarvis, where you speak and it responds almost instantly, and you can even interrupt it. That’s the top of the mountain, and it’s worth setting expectations: it’s still difficult.

v1

Text only

The starting point: Jarvis reads and writes. Robust, affordable, always available.

v2

Asynchronous voice

Voice messages (the Whisper + TTS pipeline from this module). You send audio, receive audio.

v3

Real-time connection

Fluid conversation over the phone, via services like Twilio or LiveKit, with a latency target below 1.5 seconds.

⚠️ Why real time is hard: latency

Latency and the delay between you finishing speaking and Jarvis starting to respond. On a call, more than ~1.5 seconds of silence already feels broken, strange. But that time has to fit the ENTIRE pipeline: listen + think + speak. That's why real time is on the roadmap, not your first project. Start with asynchronous voice, which can handle delays—and already delivers 90% of the magic.

🧭 The honest path

Climb the ladder one step at a time: solid text first, then asynchronous voice, and only tackle real time when everything else is in good shape. No Module 5-3 you’ll build your pocket Jarvis from end to end—the bot, brain, memory, skill, and voice—using exactly the components in this module.

Self-check (optional): in the voice pipeline, who REALLY thinks and decides on the response?

🎯 Module summary

✓
Voice and interface, not the brain — the text orchestrator thinks; voice only adds ears and a mouth, and the mouth is optional.
✓
Listen (STT) and speak (TTS) — .ogg audio → ffmpeg → Whisper → text; and text → TTS → audio. Careful: OpenAI key ≠ Anthropic.
✓
Decision for each response — the Policy Engine chooses text/voice, a cheaper/premium model, and when to escalate to a human.
✓
On-device and real time — Ollama runs the brain locally (RAM rule, hot-swap in .env); real-time connection is on the roadmap, so calibrate the latency.

Next module:

5-3 — Building your pocket Jarvis (end-to-end guided project)