PTENES
Skip to content
MODULE 2.5 · TEACHING MODE

🎙️ Communication & dictation

The most underrated skill for working with AI isn’t programming—it’s communicate. You no longer type the code: you says the intent, and the agent writes. Someone who masters dictation and clarity is, in Matt Pocock's words, "so much faster". Here you understand why and how to train this way of speaking. Each new term is explained as it comes up.

6
Topics
~40
Minutes
Zero
Prerequisite
Practice
Type
Progress: 0% 0 of 6

📖 Living glossary (read first — come back whenever you need to)

This learning path is about human skills that multiply AI. The new topic here is voice communication. Learn these terms—they appear throughout the module:

Dictation — speak aloud and have the computer transcribe it to text in real time. Instead of typing the prompt, you says the prompt.
Whisper / Whisper Flow — Whisper is OpenAI's speech recognition model; Whisper Flow is the app Matt uses to dictate into any text field on the computer.
Token — the little "piece" of text that AI processes (roughly a syllable/word). Here we use it as a metaphor: every idea that comes out of your brain as a word is a token.
Put the vision into words — put into clear words what you want built (the intent, the “why”), not just a series of disconnected commands.
Overpowered — slang for “extremely powerful,” an outsized advantage. Pocock says communicating well is “ridiculously overpowered” today.
Bandwidth — how much information passes through a channel in a given amount of time. Speaking has more bandwidth than typing: more ideas per minute.
1

🗣️ Articulate the vision

🧠 Imagine it this way: you have a perfect house in your head, with every room in its place. The bricklayer is fast and skilled — but can only build what you can describe. If you just say “build a house,” anything could come out. If you describe each room, you get your house. The AI is the bricklayer; your words are the blueprint.

In the previous track, you saw that AI "ate" the work tactical (writing the code) and that its value shifted to the side strategic (deciding what and why). But deciding isn't enough — you need to transmit the decision. And here's the key insight from this module: how you communicate your intent to the agent has become the skill that most separates those who take off from those who get stuck. Pocock calls this, plainly, communication overpowered — "ridiculously overpowered in the world of development."

Put the vision into words is different from giving an order. “Do the login” is a vague order. “I want an email and password login that reuses our authentication pattern, with friendly error messages and test coverage” is a verbalized vision: includes what, why, and the done criteria. The agent is only as good as how clearly you express yourself. The reason of this is simple: AI can’t read your mind—it only has the words you gave it. The clearer the vision in your words, the less AI has to guess—and guessing is exactly where it goes wrong.

VISION in your head clear explanation (vision)what + why + done vague order"log in" right buildyour home wrong buildanything

The same vision can become the right house or just anything—it depends on how clearly you express yourself.

Conceptual illustration: a person speaks, and the speech transforms into a glowing construction blueprint that the agent follows

⚠️ Common beginner mistake

Think that "short prompt = good prompt." Too short becomes a vague instruction, and AI fills in the gaps by guessing. Articulating the vision is the opposite of being telegraphic: it's being clear about intent and the definition of done — even if it's just a paragraph.

In one sentence: the agent builds exactly what you can put into words — so the clarity of what you say is the limit of the result.

Going deeper (optional): why does "verbalizing" also help you think?

There’s a known effect called "rubber duck debugging": explaining a problem out loud (even to a rubber duck) often makes the solution appear. Talking through your vision with the agent has the same benefit—as you speak, you find gaps in your own plan before the AI even tries. Speech isn’t just output; it’s also a reasoning tool.

2

⚡ Dictated as speed

🧠 Imagine it this way: sending a text message takes 30 seconds; a voice recording with the same idea takes 8. Speaking is a multi-lane highway; typing is a one-way street. People who dictate take the highway.

The average person types ~40 words per minute and speaks ~130. That’s more than 3× faster. When your work stopped being typing code and became describe intent for the agent, this speed difference translates directly into a productivity difference. That’s why Pocock is emphatic: "Anyone who's not doing dictation is just so much slower." It’s not motivational exaggeration — it’s the math of bandwidth.

Dictation is exactly this: you speak and the computer transcribes your prompt in real time, in any text field. Instead of typing three paragraphs of context, you speak three paragraphs in a third of the time—and still share details you’d never have the patience to type. The reason is important: the bottleneck in working with AI is no longer how quickly AI can write (it’s instant); the bottleneck has become you being able to express what you want. Dictating breaks this bottleneck.

⌨️ type ~40 words/min 🎙️ speak ~130 words/min — 3× faster Same idea, one-third the time. You’re no longer the bottleneck.
Illustration: sound waves from speech turning into lines of text flowing quickly to a screen

Quick recall: why does Pocock say that people who don't use dictation are "so much slower"?

In one sentence: the AI is already instant; dictating removes you in the critical path by tripling how much you can say per minute.

3

🔁 Brain → tokens → brain

🧠 Imagine it this way: working with AI is a two-way conversation with a water hose. You send ideas out (a stream flowing outward) and receive answers (a stream flowing inward). If the hose is thin (typing), it drips. If it's thick (speaking + listening/reading quickly), it gushes.

Pocock has a phrase that sums up the whole game: what matters is "how fast you can output tokens from your brain and input them back into your brain." In other words: the speed at which you can remove ideas from your head in the form of tokens (words) and then return to your mind the answers the AI produced. Working with an agent is this cycle in motion: brain → tokens → agent → tokens → brain, repeatedly.

The two sides of the cycle have different bottlenecks. On the output, the limit is how quickly you can speak—and that’s where dictation helps (3× more words per minute). On the input, the limit is how quickly you read and absorb the response—and that’s where speed reading, well-formatted diffs, and even listening to a summary come in. The reason to think about both sides: there’s no point in speaking quickly if you spend 10 minutes reading each response. The goal is to widen the pipe in both directions.

🧠 BRAIN you 🤖 AGENT the AI output: tokens (speak → dictation 3×) input: tokens (read/listen to the response)

Working with an agent is this loop in motion. Widen the pipe in both directions.

🔬 Worked example: the same bug, two channels

You need to explain a pagination bug to the agent. Same bug, two ways to close the loop:

Thin hose (typing)

You type "fix the pagination" in 5s, without context (you get tired of typing). The AI guesses, gets it wrong, and you slowly reread everything. A long loop, with lots of back and forth.

Thick hose (dictating)

In 20s, you says: “pagination repeats the last item when you change pages; I think the offset starts at 1 instead of 0; check the list component and run test X.” The AI gets it right on the first try; you read the diff and close the loop.

In one sentence: your AI productivity = the speed of the brain → tokens → brain loop; dictation speeds up half of the output.

4

💬 The communication skill

🧠 Imagine it this way: two managers with the same team. One explains the task in a way everyone understands the first time; the other rambles and has to redo everything three times. Same team, opposite results. The difference is purely communication.

Here's the central thesis, and it's provocative: communicating well is "ridiculously overpowered" in the world of software development—in Pocock’s words. For decades, “communication” was treated as a soft skill, secondary to technical talent. With agents, it became a out-of-the-box hard skill: it’s the channel through which all your strategy gets carried out. If you’re clear, the agent gets it right; if you’re unclear, no model in the world can save you.

The good news: communication is a skill, not a gift—you can train and improve it. And it has three practical muscles: (1) sharpness — say what, why, and the completion criteria; (2) structure — order your thinking (context → objective → constraints) instead of dumping everything together; (3) brevity without losing meaning — cut what doesn’t change the decision; keep what does. The common mistake is confusing “being technical” with “being clear”: dense jargon impresses people, but confuses the agent just as much as it would confuse a junior.

1 · clarity what · why · done 2 · structure context → objective → limit 3 · concision cut what doesn't help you decide
Illustration: a human manager directing a team of AI agents with clear, glowing instructions

✓ Overpowered communication

  • • State the goal and why, not just the instruction.
  • • Structure: context, then request, then limits.
  • • Cut the extras; keep what drives decisions.

✗ Communication that stalls

  • • A vague instruction without context (“do this”).
  • • Dense jargon that sounds technical but confuses people.
  • • Everything dumped together, with no order.

In one sentence: communicating well has become a hard skill — it’s the channel through which all your strategy reaches the agent, and you can train it.

5

🛠️ Dictation tools

🧠 Imagine it this way: a magic button that, while you hold it and speak, turns your voice into text anywhere on the screen — in the agent chat, the editor, the search field. That's what a dictation app does.

The tool Pocock uses is Whisper Flow. The name comes from Whisper, OpenAI’s speech recognition model—extremely good at transcription, including accents and technical terms. “Flow” is the app that plugs this engine into your whole system: you trigger a shortcut, speak, and the text appears wherever your cursor is. Whether it’s the Claude Code terminal, an email field, or the editor—it works anywhere.

There are alternatives (the operating system’s own built-in dictation, or other apps based on Whisper), but the point isn’t the brand — it’s adopt the habit. O reason a dedicated tool instead of native dictation: the good ones have low latency, automatic punctuation, formatting, and work globally (in any app), which makes the difference between “use it once in a while” and “use it all day.” The common mistake is installing it, trying it once, finding it strange, and giving up—the learning curve takes a few days; after that, there’s no going back.

🎙️ your voice Whisper Flowtranscription engine terminal (Claude Code) code editor email / search
Going deeper (optional): why do punctuation and technical context matter so much?

Weak dictation transcribes “use the list view component” as “use the list. view component” or gets variable names wrong. When you send a technical prompt to the agent, these errors clutter the instruction and the AI misinterprets it. Good engines (like Whisper) handle jargon well and add punctuation automatically — so recognition quality is no small detail: it protects the clarity of what you say all the way to the agent.

In one sentence: Whisper Flow (or similar) turns your voice into text in any app—the important thing isn’t the brand, but adopting the habit.

6

🎯 Practice speaking

🧠 Imagine it this way: no one becomes fluent in a language by reading about grammar—they get fluent by speaking every day, even if they make mistakes at first. Dictation is the same: it feels strange for the first few days, then becomes automatic after a week.

Wrapping up the module: overpowered communication doesn’t come out of nowhere—it’s a built habit. At first, talking to the computer feels strange (everyone goes through this), and you’ll want to go back to the keyboard. Stick with it for a week. The goal isn’t to speak eloquently; it’s to speak structured and clear automatically, the way you'd explain it to a capable junior. Below is a mental checklist—a "human skill" to follow until the habit comes naturally. Copy and paste it onto a sticky note next to your screen:

treinar-a-fala.txt
TREINO DE COMUNICAÇÃO OVERPOWERED — 1 semana

1. INSTALE um app de dictation (ex.: Whisper Flow) e defina um atalho global.
2. REGRA: por 1 semana, todo prompt longo é DITADO, não digitado. Sem exceção.
3. ESTRUTURA ao falar (sempre nesta ordem):
   - CONTEXTO: onde estamos / o que já existe.
   - OBJETIVO: o que quero alcançar (a VISÃO, não só a ordem).
   - PORQUÊ: por que isso importa.
   - LIMITES: restrições, padrão a seguir, o que NÃO fazer.
   - PRONTO: como sei que terminou (teste / critério).
4. REVISE a transcrição 1x antes de enviar (pega erro de reconhecimento).
5. LOOP: feche o ciclo lendo o diff/resposta rápido; não releia tudo devagar.

Meta: falar nitido e estruturado vira automatico. Ai voce esta "so much faster".
Day 1strange Day 3lighter Day 7automatic—"so much faster"

Quick recall: what's the most faithful way to describe Pocock's thesis about communication?

In one sentence: dictate for a week with a fixed structure; the strange becomes automatic, and you become "so much faster".

🧾 Module Summary

✓
Put the vision into words — the agent builds only what you can describe (what + why + done).
✓
Dictation is speed — speaking has ~3× the bandwidth of typing; “anyone who doesn't dictate is so much slower.”
✓
Brain → tokens → brain — the work is this loop; widen the pipe in both directions.
✓
Communication is overpowered — it became a hard skill; use Whisper Flow and practice speaking for a week.

Next module:

2.6 — You’re in control of the product: AI is weak at original ideas; the vision and features are yours.