PTENES
Skill · Claude Code · HTML→MP4

From the subject to the narrated video, no API key.

You provide a topic. The skill writes the script, records the voiceover, animates the scenes, and renders MP4s in 16:9 and 9:16 — all on your machine.

video-explicativo project cover
What it is

A skill that packages the entire explainer video pipeline

It’s not an AI video generator. It’s a deterministic pipeline: animated HTML + headless Chrome + FFmpeg, orchestrated by HyperFrames, with local TTS narration. The result is always PT-BR, premium dark amber, and ends with the INEMA.CLUB CTA.

🔒 100% local, no API key

Headless Chrome + FFmpeg render the video; the narration comes from local inemavox (engine chatterbox, voice nei) with Kokoro pf_dora as a fallback. Nothing leaves the machine.

📐 Data-driven, dynamic number of scenes

The generator reads the array SCENES[]: one entry per beat in the script. A healthy range is 4–12 content scenes—the subject decides, not a fixed template.

📱 16:9 and 9:16 from the same base

The same generator writes both formats (--vertical: vertical version of the explainer—the real reel is made with makeshorts + explicavideos 9:16), with app UI safe zones and a persistent 2-line title in vertical format.

How it works

Eight steps, always in this order

Timing is the single source of truth: each scene declares the REAL duration of its WAV (measured with ffprobe), and the generator derives both the data-start/duration as well as tween timing — audio and animation never drift out of sync.

Script→ Text review→ Project→ Sources→ Narration→ Composition→ Validate→ Render

✍️ Text (1–2)

Script with a hook in scene 1 and a review in two forms per sentence: screen (PT-BR with accents, English in its original spelling) and speaks (acronyms expanded, English rendered phonetically — deploy → “deploy”).

🎙️ Audio (3–5)

Project created in ~/projetos/output/<nome>/, fonts downloaded as .woff2 locally, and one WAV per scene generated by narration.sh.

🎞️ Video (6–8)

Composition using the motion vocabulary M.*, lint + layout inspection, render in draft to check and high to deliver.

Prerequisites

What needs to be on the machine

Everything runs locally. If something fails, npx hyperframes doctor points out what’s missing.

Node 22+ and FFmpeg

FFmpeg at C:\ffmpeg\bin. In git-bash, always use -nostdin, otherwise the command exits with code 0 without generating a file.

# check the version
node -v
ffmpeg -nostdin -version

Headless Chrome (HyperFrames)

HyperFrames renders the HTML in its own headless Chrome — download it once.

# download the HyperFrames browser
npx hyperframes browser ensure

# general diagnostics
npx hyperframes doctor

TTS: local inemavox (Nei’s voice)

House default — requires GPU. python3 of the system runs the tts_direct.py; the conda env chatterbox is called internally for timbre transfer.

# optional fallback (Kokoro pf_dora)
pip install kokoro-onnx soundfile

No GPU? You can still run everything this way

The render (headless Chrome + FFmpeg) never needed a GPU. Only two parts of the default workflow require extra hardware — and both have direct substitutes.

🎙️ Narration → edge-tts

Synthesizes in Microsoft’s cloud, with no key or local model. List and listen to the voices first before generating the full video — changing voices midway means rework.

# install and view the PT-BR voices
pip install edge-tts
edge-tts --list-voices | grep pt-BR

# sample 1 sentence to listen to
edge-tts --voice pt-BR-FranciscaNeural \
  --text "Teste de voz." --write-media amostra.mp3

🔌 Offline narration → Kokoro

Small ONNX model, runs on CPU, no internet. PT voices: pf_dora (F), pm_alex e pm_santa (M). Same rule: generate one sentence with each candidate and choose by listening.

pip install kokoro-onnx soundfile

# the video timing comes from the actual audio
ffprobe -v error -show_entries format=duration \
  -of csv=p=0 assets/audio/s1.wav

🖼️ Images → Agnes AI

Instead of flux2-klein (local diffusion, GPU): the CLI imagens-agnes calls the Agnes API — US$ 0, no credits, nothing runs on your machine. Prompt in English (PT gets caught by the filter), at most 2 references, and download them right away (the URL expires).

cd ~/projetos/imagens-agnes
python3 gerar.py "dark premium abstract \
  data landscape, amber accent" \
  --ratio 16:9 --size 2K -o s3.png

Without Agnes or a GPU, the SVG fallback still applies: the house style is already vector/CSS, images are an enhancement and not a requirement. Full details in references/sem-gpu.md.

User guide · step by step

From the subject to the MP4

The commands below are the skill’s actual commands. All content — project, assets, audio, index.html and the final MP4s — all live in one folder: ~/projetos/output/<nome>/.

1

Write the script (SCRIPT.md)

One beat per scene, 1–3 short sentences (~8–15s of voice each). Scene 1 opens directly in the hook — a sharp question, shocking number, concrete promise, or common mistake. No logo, no “hey everyone”: the first ~3s determine retention. Default duration when nobody specifies one: ~1:40–2:00.

# reference arc (expand or split beats according to the topic)
# hook → first principle → mechanics → key concept →
# application → advanced → real example → closing → INEMA.CLUB CTA
2

Review the text before generating audio and slides

Each sentence has two forms. Screen (caption + literals in html(p)): PT-BR with accents checked word by word, English terms in their original spelling. Speech (txt/sN.txt): acronyms and URLs expanded, English rewritten phonetically — TTS phonemizes based on the written spelling.

# screen  → Toda skill começa com o SKILL.md
# speech  → Toda skiu começa com o SKILL ponto M D

# pronunciation guide: deploy→"duh-ploy" · design→"dih-zine" · framework→"fraym-work"
3

Create the project in a single folder

HyperFrames init with the example blank (the skill has its own house style). Copy the design.md from the house style reference to the project root.

cd ~/projetos/output
npx hyperframes init <nome> --example blank --non-interactive
4

Download the fonts locally

Sora, Inter, and JetBrains Mono as .woff2 (subset latin) + fonts.css. Never use Google Fonts via CDN: they disappear from the render.

# copy scripts/fetch-fonts.mjs into the project and run it
node fetch-fonts.mjs   # → assets/fonts/*.woff2 + fonts.css
5

Generate the narration (nei voice, local)

Copy scripts/narration-template.sh as assets/narration.sh, write the txt/sN.txt in forma-fala and run it. It iterates over all sN.txt if any exist, tries the voice nei in local inemavox and falls back to Kokoro per scene if it fails.

bash assets/narration.sh   # → assets/audio/sN.wav

# measure the REAL duration of each WAV (goes in the scene's `audio` field)
ffprobe -v error -show_entries format=duration \
  -of default=noprint_wrappers=1:nokey=1 assets/audio/s1.wav
6

Compose the scenes

Copy scripts/composition-template.mjs as build-index.mjs and edit the array SCENES[] — one entry per scene, each one { audio, caption, html(p), anim(at,p) }. The INEMA.CLUB CTA is already attached as the final scene. In 9:16, set TITLE with persona + hook.

const TITLE = { l1: "PROFISSIONAL <b>LIBERAL</b>",
                l2: "por que ainda faz tudo <b>sozinho?</b>" };

// the CTA is always last — do not remove
const ALL = [...SCENES, CTA];
7

Validate before rendering

Lint must report 0 errors and inspect must report 0 layout issues. Don’t leave two index*.html in the root — lint flags multiple_root_compositions.

npx hyperframes lint                  # 0 errors
npx hyperframes inspect --samples 16  # 0 layout issues
8

Render both formats

Render right after generating each mode — the generator always writes index.html. One file per format, and only that: the render is the deliverable, with no compressed copy alongside it. Check frames in draft before the high.

node build-index.mjs           && npx hyperframes render --quality high --output <nome>-16x9.mp4
node build-index.mjs --vertical && npx hyperframes render --quality high --output <nome>-9x16.mp4

# extract a frame for review
ffmpeg -nostdin -y -ss 12 -i <nome>-16x9.mp4 -vframes 1 -update 1 frame.png
Examples

What the skill has already produced

Real cases that validated each pipeline feature—and the course published about the skill itself.

🎓 Course: Skills in Claude Code

The example that comes with the narration-template.sh: 8 scenes + CTA explaining what Skills are, from SKILL.md to progressive disclosure. The complete course on this skill is published at skill-explainer-video.

💼 “Liberal v1” pilot

Case that validated the persistent 2-line title in 9:16 (v1.10.3): l1 = "INDEPENDENT PROFESSIONAL" (persona), l2 = "why are you still doing everything yourself?" (hook). Fixed at the top for the entire video, disappearing at the CTA.

📊 hormozi-12-tips

Origin of the 9:16 safe zones (v1.5.0/1.5.1): message in the middle, media as a top band entering from the right, caption hidden in vertical — clear space for the app UI.

🎞️ Variations (2–3 versions)

A variation is another angle on the same topic — educational vs. real case vs. contrarian vs. list — with its own script, hook, and narration. Each one generates both formats: <nome>-v1-16x9.mp4, <nome>-v2-16x9.mp4…

Roadmap

How the skill reached 1.12.3

Versioning v1.yy.xxx — yy = feature, xxx = correction. Full history in CHANGELOG.md.

1.0.0
Initial releaseHTML→MP4 pipeline (HyperFrames + local Kokoro TTS), dark premium amber house style, 16:9 and 9:16.
1.2.0
Data-driven scenes + mid-scene activityThe number of scenes now comes from the script (SCENES[]), CTA attached automatically and Ken Burns camera motion in every scene — the end of the slideshow.
1.3.0
Motion languageVocabulary M.* (reveal/sweep/type/float/pulse/glow/countUp) instead of improvised tweens, with smaller movements in 9:16.
1.4.0
Transitions between scenesMap TRANS in GSAP: fade (default), push, slideUp, zoom, wipe, fadeBlack — special effects only at 2–3 key moments.
1.5.x
Safe zones, SVG fallback, and single outputValidated vertical layout, raster images are optional (SVG covers it), the tail end changes, and there’s one MP4 per format — no duplicate -FINAL.
1.6.3
Text review + English pronunciationNew step before the WAVs: two forms per sentence (on-screen vs. spoken) and a phonetic English lexicon.
1.7.3
Hook, audio under the voice, and readabilityFour behaviors become the contract: opening hook, low music (MUSIC_VOL ~0.14), scrim/blur/panel beneath text, and what counts as a variation.
1.8.3
bella voice as defaultNarration now uses the house voice via inemavox (chatterbox-vc); Kokoro pf_dora automatically becomes the fallback for each scene.
1.10.3
9:16 title in 2 linesTITLE = { l1, l2 } — persona + curiosity, fixed at the top of the vertical frame and disappearing at the CTA, closing the retention loop.
1.12.3
Explainer, not a reelCurrent version. Reel and short now go through makeshorts with explicavideos 9:16; 100% local narration with Nei’s voice (no Edge TTS); narration checked with local transcription; one prototype approved before any batch.
1.11.3
Run without a GPUDocuments alternatives for those without a GPU: narration via edge-tts or Kokoro (with required voice validation) and images via Agnes AI (US$ 0) instead of flux2-klein.