🎬 SCRIPT dot M D with 6–9 scenes
The SCRIPT dot M D is the heart of your video. It describes each scene in text — what appears on screen and what is spoken. A well-structured arc ensures maximum retention in short videos.
The SCRIPT dot M D follows an 8-step dramatic arc: hook → first principle → mechanics → key concept → application → advanced → real example → close → CTA. Each step is a scene. The arc is there to hold attention from the beginning through the CTA for inema ponto club.
Videos with ~100 seconds of speech (≈ 1:50 of video) perform better for retention and fit as Shorts. Each scene should have 1–3 sentences, never a monologue.
Opening line that immediately grabs attention. A question, bold statement, or surprising fact. Example: "What if you could create professional videos for free?"
Explains the most basic concept. No jargon. One paragraph anyone can understand. This is where understanding begins.
Explain the inner workings in detail: Chrome captures frames, FFmpeg encodes, Kokoro speaks. Technical but direct.
The idea that changes how you see the problem. Usually a memorable sentence. Example: "The browser is a movie camera."
First command or concrete step. The viewer sees how to use it. Example: "Run npx hyperframes init and the project is already configured."
A resource that separates beginners from advanced users. Hooks, flags, alternative voices, two formats. Sparks curiosity and adds value.
This video was made with the tool itself. Or a link, screenshot, concrete result. It breaks skepticism.
A sentence that closes the hook's loop. Reinforces the transformation the viewer will experience. Short and firm.
Point to the full course at inema dot club. Always the last scene. Direct, with a clear action: "Go to inema dot club now."
- ✓ Each scene has a single focus — don’t mix two concepts
- ✓ 1–3 sentences of narration per scene (≤15 seconds)
- ✓ Total spoken content ≈ 100s for a ~1:50 video
- ✓ Always end with an explicit CTA
- ✗ Scenes that are too long — attention drops after 15s
- ✗ Hook Without Tension — Doesn't Promise a Transformation
- ✗ Skip from the fundamentals straight to advanced
- ✗ Forget the CTA — without it, the video won't convert
✂️ Short narration per scene
Each scene gets 1 to 3 narration sentences. About 100 seconds of speech produces approximately 1 minute and 50 seconds of video — ideal for retention and compatible with Shorts.
People retain more information in short videos. With ~100s of speech in total, each scene lasts ≈11–17 seconds—long enough to absorb an idea, and short enough to avoid boredom.
Kokoro TTS performs best with simple, direct sentences. Avoid excessive use of gerunds, long subordinate clauses, or spoken lists. If you need a pause, split the text into two sentences with a period.
🗣️ Expand acronyms for speech
Kokoro TTS reads the text literally. Acronyms, file extensions, and URLs need to be written as they're spoken — otherwise, the TTS will pronounce them strangely or incomprehensibly.
The file text sN.txt is written for ears, not for the eyes. Any symbol, acronym, or path that isn't a pronounceable word needs to be replaced with its exact pronunciation.
| Written text | Speech in sN.txt | Reason |
|---|---|---|
| SKILL.md | SKILL point M D | file extension |
| .claude/skills | dot claude slash skills | path with symbols |
| inema.club | inema dot club | URL / domain |
| build-index.mjs | build hyphen index dot M J S | file name |
| s1.txt | S one point T X T | file name |
| --speed 0.98 | speed zero point ninety-eight | CLI flag with a number |
| pf_dora | P F underscore dora | voice identifier |
| npx hyperframes | N P X hyperframes | acronym + command |
- ✓ "dot claude slash skills" for
.claude/skills - ✓ "inema dot club" for
inema.club - ✓ "SKILL dot M D" for
SKILL.md - ✓ Test the pronunciation aloud before saving
- ✗ Leave
SKILL.mdwithout expanding - ✗ Use a literal URL
https://inema.club - ✗ Writing bash commands like
npx --help - ✗ Use hyphenated lists — TTS reads the hyphen aloud
Keep the SCRIPT dot M D with the “visual” text (with acronyms, paths, and URLs as usual) for human reference. The file sN.txt is the "for the ears" version — text already expanded that goes to TTS. They are different documents with different purposes.
🔊 Generate WAV with Kokoro
With the files sN.txt ready and expanded, the command npx hyperframes tts generates the WAVs locally. The first run automatically downloads ~340 MB of the Kokoro model.
The first time you run it npx hyperframes tts, Kokoro downloads the voice model (~340 MB) automatically. No key, no espeak-ng, no config. After the download, subsequent runs are instant. Make sure you have an internet connection the first time.
Kokoro's default speed sounds slightly too fast for technical narration in Portuguese. With --speed 0.98 the voice sounds natural without sounding slow. Don’t go above 1.05 — the voice sounds metallic.
Run only scene 1 first. Listen to the result. If the pronunciation of any expansion sounds strange, correct the s1.txt before generating the other 7 files. Reworking scenes one at a time is much faster than redoing everything.
📏 Measure durations with ffprobe
Before assembling the scenes in the build-index.mjs, you need to know exactly how many seconds each narration lasts. The ffprobe returns the WAV duration in seconds on one line.
O build-index.mjs defines how long each scene stays on screen via LEAD, TAIL and the audio duration. If you don't know the exact WAV duration, the text will disappear before the speech ends—or stay on screen too long.
- ✓ Generate all WAVs first, then measure
- ✓ Note the durations in SCRIPT point M D
- ✓ Use the durations in build-index to set the timing
- ✓ LEAD=0.5 before speech + TAIL=0.9 after speech
- ✗ Use estimated duration — always measure the actual WAV
- ✗ Cut the scene before the audio ends
- ✗ Don't let the LEAD start before the narration
- ✗ Forget FADE=0.45 at the end of the scene
O build-index.mjs uses by default: LEAD=0.5 (silence before speech), TAIL=0.9 (silence after speech) and FADE=0.45 (exit fade-out). The total scene duration = LEAD + wav_duration + TAIL.
🎚️ Available PT-BR voices
Kokoro has three PT-BR voices ready to use: pf_dora (female, recommended default), pm_alex e pm_santa. Each voice has a distinct timbre—choose one based on the video's tone.
Clear, natural female voice in Brazilian Portuguese. It’s the default voice for narration-template.sh and recommended for all videos in the HyperFrames pipeline.
Deep male voice, good for more serious or technical content. An alternative for variety in long series or when the tone calls for more authority.
Third PT-BR option with a distinct timbre. Use it to test whether the specific content sounds more natural in this voice or to A/B test retention.
- ✓ Use
pf_doraas the default — it's the most tested - ✓ Keep the same voice throughout the video
- ✓ Test the voice with the most complex scene first
- ✓ Speed
0.95–1.00for technical narration
- ✗ Mixing voices in the same video
- ✗ Speed above 1.05 — sounds robotic
- ✗ Speed below 0.85 — drags too much
- ✗ Try installing espeak-ng — Kokoro doesn't need it
| Voice | Gender | Timbre | Best for |
|---|---|---|---|
| pf_dora | Female | Clear, natural | Courses, tutorials, standards |
| pm_alex | Feminine | Record, authoritative | Serious tech, demos |
| pm_santa | Feminine | Distinctive | A/B testing, variety |
Kokoro has a native PT-BR phonemizer — it doesn't need espeak-ng, which other open-source TTS engines require. No API key, no platform account. Installing espeak-ng may even conflict with Kokoro's phonetics, so avoid it.
📋 Module 2.3 Summary
- ✓ SCRIPT point M D with an 8–9-scene arc: hook → principle → mechanics → insight → application → advanced → example → close → CTA
- ✓ 1–3 sentences of narration per scene; ~100s of total speech ≈ 1:50 of video
- ✓ Acronym expansion: "SKILL.md" → "SKILL point M D"; ".claude/skills" → "point claude slash skills"; "inema.club" → "inema dot club"
- ✓ Generate WAV:
npx hyperframes tts "assets/txt/s1.txt" --voice pf_dora --speed 0.98 --output assets/audio/s1.wav - ✓ Measure duration:
ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/s1.wav - ✓ PT-BR voices:
pf_dora(default),pm_alex,pm_santa— no espeak-ng, no key, ~340 MB one-time download
build-index.mjs using the WAV durations. Define LEAD, TAIL, FADE, and the timings for each animation to generate the index.html final, ready to render.