PTENES
MODULE 2.3

📝 Script & TTS narration

Learn to write a SCRIPT dot M D with 6–9 scenes in a dramatic arc, generate local PT-BR narration with Kokoro, expand acronyms for TTS, and measure durations with ffprobe—all without an API key.

6
Topics
~30
Minutes
Practical
Level
Hands-on
Type
Script → TTS narration SCRIPT.md · sN.txt · pf_dora · sN.wav SCRIPT dot M D 6–9 scenes assets/txt/ sN dot txt expanded speech Kokoro pf_dora PT-BR --speed 0.98 audio/ sN dot wav ~12s/scene ① SCRIPT ② TEXT ③ TTS ④ WAV 📏 ffprobe — measure duration -show_entries format=duration · noprint_wrappers=1:nokey=1
1

🎬 SCRIPT dot M D with 6–9 scenes

The SCRIPT dot M D is the heart of your video. It describes each scene in text — what appears on screen and what is spoken. A well-structured arc ensures maximum retention in short videos.

Main Concept

The SCRIPT dot M D follows an 8-step dramatic arc: hook → first principle → mechanics → key concept → application → advanced → real example → close → CTA. Each step is a scene. The arc is there to hold attention from the beginning through the CTA for inema ponto club.

Videos with ~100 seconds of speech (≈ 1:50 of video) perform better for retention and fit as Shorts. Each scene should have 1–3 sentences, never a monologue.

8–9-scene arc
1
Hook — The Promise

Opening line that immediately grabs attention. A question, bold statement, or surprising fact. Example: "What if you could create professional videos for free?"

2
First principle — foundation

Explains the most basic concept. No jargon. One paragraph anyone can understand. This is where understanding begins.

3
Mechanics — how it works

Explain the inner workings in detail: Chrome captures frames, FFmpeg encodes, Kokoro speaks. Technical but direct.

4
Key concept — the insight

The idea that changes how you see the problem. Usually a memorable sentence. Example: "The browser is a movie camera."

5
Application—in practice

First command or concrete step. The viewer sees how to use it. Example: "Run npx hyperframes init and the project is already configured."

6
Advanced — higher level

A resource that separates beginners from advanced users. Hooks, flags, alternative voices, two formats. Sparks curiosity and adds value.

7
Real example — social proof

This video was made with the tool itself. Or a link, screenshot, concrete result. It breaks skepticism.

8
Closing — summary

A sentence that closes the hook's loop. Reinforces the transformation the viewer will experience. Short and firm.

9
CTA — inema dot club

Point to the full course at inema dot club. Always the last scene. Direct, with a clear action: "Go to inema dot club now."

SCRIPT structure, point M D
# SCRIPT.md — example for a video about HyperFrames Skills
## Scene 1 — Hook
**On screen:** title "Skills in Claude Code" fadeIn
**Spoken:** "What are Skills in Claude Code, really?"

## Scene 2 — First principle
**On screen:** folder + SKILL.md diagram
**Spoken:** "A Skill is just a folder with a file called SKILL dot M D."

## ... (scenes 3–8)

## Scene 9 — CTA
**On screen:** INEMA.CLUB logo + URL
**Spoken:** "Full course at inema dot club."
✓ Scriptwriting best practices
  • ✓ Each scene has a single focus — don’t mix two concepts
  • ✓ 1–3 sentences of narration per scene (≤15 seconds)
  • ✓ Total spoken content ≈ 100s for a ~1:50 video
  • ✓ Always end with an explicit CTA
✗ Script errors
  • ✗ Scenes that are too long — attention drops after 15s
  • ✗ Hook Without Tension — Doesn't Promise a Transformation
  • ✗ Skip from the fundamentals straight to advanced
  • ✗ Forget the CTA — without it, the video won't convert
Key concepts
🎯
Dramatic arc
8–9 steps
⏱️
~100s of speech
≈ 1:50 video
📌
1–3 sentences/scene
Maximum retention
📣
Final CTA
inema.club
2

✂️ Short narration per scene

Each scene gets 1 to 3 narration sentences. About 100 seconds of speech produces approximately 1 minute and 50 seconds of video — ideal for retention and compatible with Shorts.

100-second rule

People retain more information in short videos. With ~100s of speech in total, each scene lasts ≈11–17 seconds—long enough to absorb an idea, and short enough to avoid boredom.

8 scenes
×12.5s average
~100s
of total speech
≈1:50
of the final video
Real example — 8 scenes from narration-template.sh
# assets/narration.sh — excerpt with write() for each scene
write s1 "What are Skills in Claude Code, really?"
write s2 "A Skill is just a folder with a file called SKILL dot M D."
write s3 "Every SKILL dot M D starts with two lines: name and description."
write s4 "Progressive disclosure: Claude loads only what it needs, when it needs it."
write s5 "Skills live in dot claude slash skills in your project or in the global folder."
write s6 "At the advanced level, a Skill includes scripts, palettes, and complete templates."
write s7 "This video was made with the HyperFrames Skill. One skill, one workflow, one result."
write s8 "Keep it simple. Full course at inema dot club."
💡
Short sentence = natural speech

Kokoro TTS performs best with simple, direct sentences. Avoid excessive use of gerunds, long subordinate clauses, or spoken lists. If you need a pause, split the text into two sentences with a period.

Comparison: time distribution by scene
Scene 1
~8s
Scenes 2–7
~72s
Scene 8
~6s
Final CTA
~5s
Narration
CTA/Closing
Key concepts
✂️
1–3 sentences
per scene
📐
~12s/scene
ideal average
⏱️
100s total
accumulated speech
📱
Fits in Shorts
≤60s on screen
3

🗣️ Expand acronyms for speech

Kokoro TTS reads the text literally. Acronyms, file extensions, and URLs need to be written as they're spoken — otherwise, the TTS will pronounce them strangely or incomprehensibly.

Expansion Rule

The file text sN.txt is written for ears, not for the eyes. Any symbol, acronym, or path that isn't a pronounceable word needs to be replaced with its exact pronunciation.

Table of required expansions
Written text Speech in sN.txt Reason
SKILL.md SKILL point M D file extension
.claude/skills dot claude slash skills path with symbols
inema.club inema dot club URL / domain
build-index.mjs build hyphen index dot M J S file name
s1.txt S one point T X T file name
--speed 0.98 speed zero point ninety-eight CLI flag with a number
pf_dora P F underscore dora voice identifier
npx hyperframes N P X hyperframes acronym + command
Example: s2.txt — expanded text for TTS
# assets/txt/s2.txt — version for TTS to read
"Start with the essentials. A Skill is just a folder with a file called
SKILL point M D. Inside it, Markdown instructions that teach Claude
for doing something specific: creating videos, reviewing code, designing interfaces.
It's packaged knowledge."

# DO NOT write: "SKILL.md" → the TTS would read it as "squill dot md" or get it wrong
# DO NOT write: ".claude/skills" → it would be read incorrectly as "dot claude slash skills"
✓ Correct expansions
  • ✓ "dot claude slash skills" for .claude/skills
  • ✓ "inema dot club" for inema.club
  • ✓ "SKILL dot M D" for SKILL.md
  • ✓ Test the pronunciation aloud before saving
✗ Expansion errors
  • ✗ Leave SKILL.md without expanding
  • ✗ Use a literal URL https://inema.club
  • ✗ Writing bash commands like npx --help
  • ✗ Use hyphenated lists — TTS reads the hyphen aloud
💡
Tip: two files, two purposes

Keep the SCRIPT dot M D with the “visual” text (with acronyms, paths, and URLs as usual) for human reference. The file sN.txt is the "for the ears" version — text already expanded that goes to TTS. They are different documents with different purposes.

Key concepts
👁️
Visual text
SCRIPT.md
👂
Spoken text
sN.txt
🔤
Expansion
dot, slash, hyphen
🎤
Spoken test
Read aloud
4

🔊 Generate WAV with Kokoro

With the files sN.txt ready and expanded, the command npx hyperframes tts generates the WAVs locally. The first run automatically downloads ~340 MB of the Kokoro model.

Exact command — npx hyperframes tts
# Generates narration for scene 1
npx hyperframes tts "assets/txt/s1.txt" \
--voice pf_dora \
--speed 0.98 \
--output assets/audio/s1.wav

# Loop through all scenes (narration.sh)
for i in 1 2 3 4 5 6 7 8; do
npx hyperframes tts "assets/txt/s$i.txt" \
--voice pf_dora --speed 0.98 \
--output "assets/audio/s$i.wav"
done
⚠️
First run: download of ~340 MB

The first time you run it npx hyperframes tts, Kokoro downloads the voice model (~340 MB) automatically. No key, no espeak-ng, no config. After the download, subsequent runs are instant. Make sure you have an internet connection the first time.

Why --speed 0.98?

Kokoro's default speed sounds slightly too fast for technical narration in Portuguese. With --speed 0.98 the voice sounds natural without sounding slow. Don’t go above 1.05 — the voice sounds metallic.

0.80
Too slow
0.98 ✓
Ideal for Brazilian Portuguese
1.05+
Metallic
💡
Generate a test WAV before the full loop

Run only scene 1 first. Listen to the result. If the pronunciation of any expansion sounds strange, correct the s1.txt before generating the other 7 files. Reworking scenes one at a time is much faster than redoing everything.

Key concepts
🔊
npx hyperframes tts
generator command
🎙️
pf_dora
default PT-BR voice
⚡
--speed 0.98
natural and fluent
📦
~340 MB
single download
5

📏 Measure durations with ffprobe

Before assembling the scenes in the build-index.mjs, you need to know exactly how many seconds each narration lasts. The ffprobe returns the WAV duration in seconds on one line.

Why measure before assembling?

O build-index.mjs defines how long each scene stays on screen via LEAD, TAIL and the audio duration. If you don't know the exact WAV duration, the text will disappear before the speech ends—or stay on screen too long.

ffprobe commands — WAV duration and full loop
# Duration of a single file
ffprobe -v error \
-show_entries format=duration \
-of default=noprint_wrappers=1:nokey=1 \
assets/audio/s1.wav
# Output: 12.384000

# Loop — measure all scenes at once
for i in 1 2 3 4 5 6 7 8; do
d=$(ffprobe -v error -show_entries format=duration \
-of default=noprint_wrappers=1:nokey=1 \
"assets/audio/s$i.wav" 2>/dev/null)
echo "s$i: ${d}s"
done
📊 Reference values (real narration-template)
s1
~8–10s
short hook
s2–s6
~12–18s
development
s7
~8–12s
real example
s8
~5–8s
Quick CTA
✓ Correct workflow
  • ✓ Generate all WAVs first, then measure
  • ✓ Note the durations in SCRIPT point M D
  • ✓ Use the durations in build-index to set the timing
  • ✓ LEAD=0.5 before speech + TAIL=0.9 after speech
✗ Timing errors
  • ✗ Use estimated duration — always measure the actual WAV
  • ✗ Cut the scene before the audio ends
  • ✗ Don't let the LEAD start before the narration
  • ✗ Forget FADE=0.45 at the end of the scene
💡
Pipeline default timing values

O build-index.mjs uses by default: LEAD=0.5 (silence before speech), TAIL=0.9 (silence after speech) and FADE=0.45 (exit fade-out). The total scene duration = LEAD + wav_duration + TAIL.

Key concepts
📏
ffprobe
measures WAV
⏩
LEAD=0.5
before the speech
⏸️
TAIL=0.9
after the speech
🌅
FADE=0.45
final fade-out
6

🎚️ Available PT-BR voices

Kokoro has three PT-BR voices ready to use: pf_dora (female, recommended default), pm_alex e pm_santa. Each voice has a distinct timbre—choose one based on the video's tone.

🎙️
pf_dora
PT-BR · Female

Clear, natural female voice in Brazilian Portuguese. It’s the default voice for narration-template.sh and recommended for all videos in the HyperFrames pipeline.

--voice pf_dora --speed 0.98
🎤
pm_alex
PT-BR · Male

Deep male voice, good for more serious or technical content. An alternative for variety in long series or when the tone calls for more authority.

--voice pm_alex --speed 0.98
🔈
pm_santa
PT-BR · Alternative

Third PT-BR option with a distinct timbre. Use it to test whether the specific content sounds more natural in this voice or to A/B test retention.

--voice pm_santa --speed 0.98
✓ Voice best practices
  • ✓ Use pf_dora as the default — it's the most tested
  • ✓ Keep the same voice throughout the video
  • ✓ Test the voice with the most complex scene first
  • ✓ Speed 0.95–1.00 for technical narration
✗ Voice errors
  • ✗ Mixing voices in the same video
  • ✗ Speed above 1.05 — sounds robotic
  • ✗ Speed below 0.85 — drags too much
  • ✗ Try installing espeak-ng — Kokoro doesn't need it
📊 Technical comparison of the voices
Voice Gender Timbre Best for
pf_dora Female Clear, natural Courses, tutorials, standards
pm_alex Feminine Record, authoritative Serious tech, demos
pm_santa Feminine Distinctive A/B testing, variety
💡
No espeak-ng, no key, no account

Kokoro has a native PT-BR phonemizer — it doesn't need espeak-ng, which other open-source TTS engines require. No API key, no platform account. Installing espeak-ng may even conflict with Kokoro's phonetics, so avoid it.

Key concepts
🏆
pf_dora
recommended default
3️⃣
3 Brazilian Portuguese voices
dora, alex, santa
🚫
No espeak-ng
native phonetics
🔑
No key
local and free

📋 Module 2.3 Summary

What you learned
  • ✓ SCRIPT point M D with an 8–9-scene arc: hook → principle → mechanics → insight → application → advanced → example → close → CTA
  • ✓ 1–3 sentences of narration per scene; ~100s of total speech ≈ 1:50 of video
  • ✓ Acronym expansion: "SKILL.md" → "SKILL point M D"; ".claude/skills" → "point claude slash skills"; "inema.club" → "inema dot club"
  • ✓ Generate WAV: npx hyperframes tts "assets/txt/s1.txt" --voice pf_dora --speed 0.98 --output assets/audio/s1.wav
  • ✓ Measure duration: ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/s1.wav
  • ✓ PT-BR voices: pf_dora (default), pm_alex, pm_santa — no espeak-ng, no key, ~340 MB one-time download
Next module
2.4
🎞️ Scene composition
Build the scenes in the build-index.mjs using the WAV durations. Define LEAD, TAIL, FADE, and the timings for each animation to generate the index.html final, ready to render.
Go to module 2.4 →