🎙️ Local Kokoro TTS (voice pf_dora)
PT-BR narration generated 100% on the machine, free and without an API key — which keeps the skill self-contained.
Kokoro is an ONNX TTS that runs locally (pip install kokoro-onnx soundfile), with the voice pf_dora (PT-BR) and --speed 0.98. The first run downloads ~340MB of the model; after that, it works offline. No API key and no per-minute cost.
npx -y hyperframes tts "assets/txt/s1.txt" \
--voice pf_dora \
--speed 0.98 \
--output "assets/audio/s1.wav"
"512" → "five hundred twelve"; "URL" → spell it out; "inema.club" → "inema dot club". TTS reads literally—spelling it out avoids strange pronunciations.
The voice sounds good, but lacks performance. Always ask the user to validate the narration—they’re the one who listens before the final render.
📝 Generate the WAVs from steps.json
O narration-template.sh extracts the dialogue from steps.json and generates one WAV per step + CTA.
# extrai steps[].narration + ctaNarration -> assets/txt/sN.txt
node -e '
const d=JSON.parse(fs.readFileSync("steps.json","utf8"));
const lines=d.steps.map(s=>s.narration||"");
lines.push(d.ctaNarration);
lines.forEach((t,i)=>fs.writeFileSync(`assets/txt/s${i+1}.txt`, t));
'
# gera 1 WAV por txt (passos + CTA)
for i in $(seq 1 "$N"); do
npx -y hyperframes tts "assets/txt/s$i.txt" \
--voice pf_dora --speed 0.98 \
--output "assets/audio/s$i.wav"
done
- ✓Write the lines in the
steps.json - ✓Ensure 1 narration per step + the
ctaNarration - ✓Run from the project root (where steps.json and assets/ are located)
- ✗Keep the narration text separate from steps.json (they’ll drift out of sync)
- ✗Don’t forget the CTA — it’s the last WAV
- ✗Numbering incorrectly: the WAVs need to be s1..sN in step order
📏 Measure the actual WAV durations
Timing is the single source of truth: the generator measures each WAV with ffprobe — there is no manually entered array of timings.
const NA = STEPS.length + 1; // +1 da CTA
const AUDIO = [];
for (let i = 1; i <= NA; i++) {
const d = parseFloat(execSync(
`ffprobe -v error -show_entries format=duration \
-of default=noprint_wrappers=1:nokey=1 \
"assets/audio/s${i}.wav"`).toString().trim());
AUDIO.push(d); // duração REAL, não estimada
}
A “10-second” line rarely lasts exactly 10 seconds. Measuring the actual WAV and deriving the scene timing from it ensures the animation ends with the voice, without manual adjustments.
In git-bash environments, use ffmpeg -nostdin in the calls to extract frames—without the flag, the process may hang while reading interactive stdin.
⚙️ The composition: build-demo.mjs
The central generator: reads steps.json, measures the WAVs, and produces the index.html with frame, cursor, highlight, zoom, and CTA ready to go.
1. lê steps.jsonLoad viewport, window, the steps (shot + target + caption + narration) and the ctaNarration.
2. mede WAVs + calcula geometriaBuilds AUDIO[] with ffprobe and maps each bbox from screenshot space to the canvas (mapBox, center).
3. escreve index.html (16:9)Scenes (1 per step) + CTA, frame, global cursor, highlights, zoom, and GSAP timeline — all in a single renderable HTML file.
node build-demo.mjs # -> index.html (16:9), pronto para lint/render
Any adjustments go in the steps.json or in the build-demo.mjs, then run it again. Editing the HTML directly gets lost in the next build.
🔗 Sync audio ↔ animation using S[]
The timeline and audio files read the same array of times derived from AUDIO[] — synchronized by design.
const LEAD = 0.5, TAIL = 0.7, FADE = 0.4;
let t = 0;
const S = AUDIO.map((a, i) => {
const dur = LEAD + a + TAIL;
const o = { i: i+1, start: t, dur,
audioStart: t + LEAD, audioDur: a, end: t + dur };
t += dur; return o;
});
Scenes and captions go on alternating tracks (1/3 and 2/4): one track ends while the other prepares the next audio, preventing overlap during FADE transitions.
💬 Captions in the footer
Each step has a caption that becomes a translucent caption synced with the scene.
O caption for each step (written to steps.json) is rendered in the footer in Inter 600 over a translucent background, entering and exiting with the scene, on a track alternating with the scenes.
Narration is what the voice says (numbers can be spelled out); the caption is the short text on screen. Both live in the same step in steps.json, but serve different purposes—silent feeds display the caption.
Keep the caption to one sentence. Long text in the footer competes with the frame and the result — the caption reinforces the narration; it doesn’t replace it.
📣 INEMA.CLUB’s final CTA
The final scene is the brand callout—it’s already included in the template, with default narration.
"CONTINUES IN" + INEMA.CLUB (INEMA cream, .CLUB amber with glow) + 🌐 inema.club. Default narration: "This is INEMA ponto CLUB content. Visit: inema ponto club." In the CTA, the cursor disappears (opacity:0).
"ctaNarration": "Isso é conteúdo do INEMA ponto CLUB. Acesse: inema ponto club."
🎯 What you learned
- ✓Kokoro generates the PT-BR narration locally, for free, voice pf_dora --speed 0.98
- ✓narration-template.sh gets the narration from steps.json (steps + CTA)
- ✓build-demo.mjs measures the WAVs with ffprobe — the single source of truth for timing
- ✓S[] (start/dur/audioStart/end) syncs audio and animation
- ✓Captions and the INEMA.CLUB CTA are ready to use
Next: Module 3.3 — build the HyperFrames project, validate it with lint/inspect, and render the final MP4.