You provide a topic. The skill writes the script, records the voiceover, animates the scenes, and renders MP4s in 16:9 and 9:16 — all on your machine.

It’s not an AI video generator. It’s a deterministic pipeline: animated HTML + headless Chrome + FFmpeg, orchestrated by HyperFrames, with local TTS narration. The result is always PT-BR, premium dark amber, and ends with the INEMA.CLUB CTA.
Headless Chrome + FFmpeg render the video; the narration comes from local inemavox (engine chatterbox, voice nei) with Kokoro pf_dora as a fallback. Nothing leaves the machine.
The generator reads the array SCENES[]: one entry per beat in the script. A healthy range is 4–12 content scenes—the subject decides, not a fixed template.
The same generator writes both formats (--vertical: vertical version of the explainer—the real reel is made with makeshorts + explicavideos 9:16), with app UI safe zones and a persistent 2-line title in vertical format.
Timing is the single source of truth: each scene declares the REAL duration of its WAV (measured with ffprobe), and the generator derives both the data-start/duration as well as tween timing — audio and animation never drift out of sync.
Script with a hook in scene 1 and a review in two forms per sentence: screen (PT-BR with accents, English in its original spelling) and speaks (acronyms expanded, English rendered phonetically — deploy → “deploy”).
Project created in ~/projetos/output/<nome>/, fonts downloaded as .woff2 locally, and one WAV per scene generated by narration.sh.
Composition using the motion vocabulary M.*, lint + layout inspection, render in draft to check and high to deliver.
Everything runs locally. If something fails, npx hyperframes doctor points out what’s missing.
FFmpeg at C:\ffmpeg\bin. In git-bash, always use -nostdin, otherwise the command exits with code 0 without generating a file.
# check the version node -v ffmpeg -nostdin -version
HyperFrames renders the HTML in its own headless Chrome — download it once.
# download the HyperFrames browser npx hyperframes browser ensure # general diagnostics npx hyperframes doctor
House default — requires GPU. python3 of the system runs the tts_direct.py; the conda env chatterbox is called internally for timbre transfer.
# optional fallback (Kokoro pf_dora) pip install kokoro-onnx soundfile
The render (headless Chrome + FFmpeg) never needed a GPU. Only two parts of the default workflow require extra hardware — and both have direct substitutes.
Synthesizes in Microsoft’s cloud, with no key or local model. List and listen to the voices first before generating the full video — changing voices midway means rework.
# install and view the PT-BR voices pip install edge-tts edge-tts --list-voices | grep pt-BR # sample 1 sentence to listen to edge-tts --voice pt-BR-FranciscaNeural \ --text "Teste de voz." --write-media amostra.mp3
Small ONNX model, runs on CPU, no internet. PT voices: pf_dora (F), pm_alex e pm_santa (M). Same rule: generate one sentence with each candidate and choose by listening.
pip install kokoro-onnx soundfile # the video timing comes from the actual audio ffprobe -v error -show_entries format=duration \ -of csv=p=0 assets/audio/s1.wav
Instead of flux2-klein (local diffusion, GPU): the CLI imagens-agnes calls the Agnes API — US$ 0, no credits, nothing runs on your machine. Prompt in English (PT gets caught by the filter), at most 2 references, and download them right away (the URL expires).
cd ~/projetos/imagens-agnes python3 gerar.py "dark premium abstract \ data landscape, amber accent" \ --ratio 16:9 --size 2K -o s3.png
Without Agnes or a GPU, the SVG fallback still applies: the house style is already vector/CSS, images are an enhancement and not a requirement. Full details in references/sem-gpu.md.
The commands below are the skill’s actual commands. All content — project, assets, audio, index.html and the final MP4s — all live in one folder: ~/projetos/output/<nome>/.
One beat per scene, 1–3 short sentences (~8–15s of voice each). Scene 1 opens directly in the hook — a sharp question, shocking number, concrete promise, or common mistake. No logo, no “hey everyone”: the first ~3s determine retention. Default duration when nobody specifies one: ~1:40–2:00.
# reference arc (expand or split beats according to the topic) # hook → first principle → mechanics → key concept → # application → advanced → real example → closing → INEMA.CLUB CTA
Each sentence has two forms. Screen (caption + literals in html(p)): PT-BR with accents checked word by word, English terms in their original spelling. Speech (txt/sN.txt): acronyms and URLs expanded, English rewritten phonetically — TTS phonemizes based on the written spelling.
# screen → Toda skill começa com o SKILL.md # speech → Toda skiu começa com o SKILL ponto M D # pronunciation guide: deploy→"duh-ploy" · design→"dih-zine" · framework→"fraym-work"
HyperFrames init with the example blank (the skill has its own house style). Copy the design.md from the house style reference to the project root.
cd ~/projetos/output npx hyperframes init <nome> --example blank --non-interactive
Sora, Inter, and JetBrains Mono as .woff2 (subset latin) + fonts.css. Never use Google Fonts via CDN: they disappear from the render.
# copy scripts/fetch-fonts.mjs into the project and run it node fetch-fonts.mjs # → assets/fonts/*.woff2 + fonts.css
Copy scripts/narration-template.sh as assets/narration.sh, write the txt/sN.txt in forma-fala and run it. It iterates over all sN.txt if any exist, tries the voice nei in local inemavox and falls back to Kokoro per scene if it fails.
bash assets/narration.sh # → assets/audio/sN.wav # measure the REAL duration of each WAV (goes in the scene's `audio` field) ffprobe -v error -show_entries format=duration \ -of default=noprint_wrappers=1:nokey=1 assets/audio/s1.wav
Copy scripts/composition-template.mjs as build-index.mjs and edit the array SCENES[] — one entry per scene, each one { audio, caption, html(p), anim(at,p) }. The INEMA.CLUB CTA is already attached as the final scene. In 9:16, set TITLE with persona + hook.
const TITLE = { l1: "PROFISSIONAL <b>LIBERAL</b>", l2: "por que ainda faz tudo <b>sozinho?</b>" }; // the CTA is always last — do not remove const ALL = [...SCENES, CTA];
Lint must report 0 errors and inspect must report 0 layout issues. Don’t leave two index*.html in the root — lint flags multiple_root_compositions.
npx hyperframes lint # 0 errors npx hyperframes inspect --samples 16 # 0 layout issues
Render right after generating each mode — the generator always writes index.html. One file per format, and only that: the render is the deliverable, with no compressed copy alongside it. Check frames in draft before the high.
node build-index.mjs && npx hyperframes render --quality high --output <nome>-16x9.mp4 node build-index.mjs --vertical && npx hyperframes render --quality high --output <nome>-9x16.mp4 # extract a frame for review ffmpeg -nostdin -y -ss 12 -i <nome>-16x9.mp4 -vframes 1 -update 1 frame.png
Real cases that validated each pipeline feature—and the course published about the skill itself.
The example that comes with the narration-template.sh: 8 scenes + CTA explaining what Skills are, from SKILL.md to progressive disclosure. The complete course on this skill is published at skill-explainer-video.
Case that validated the persistent 2-line title in 9:16 (v1.10.3): l1 = "INDEPENDENT PROFESSIONAL" (persona), l2 = "why are you still doing everything yourself?" (hook). Fixed at the top for the entire video, disappearing at the CTA.
Origin of the 9:16 safe zones (v1.5.0/1.5.1): message in the middle, media as a top band entering from the right, caption hidden in vertical — clear space for the app UI.
A variation is another angle on the same topic — educational vs. real case vs. contrarian vs. list — with its own script, hook, and narration. Each one generates both formats: <nome>-v1-16x9.mp4, <nome>-v2-16x9.mp4…
Versioning v1.yy.xxx — yy = feature, xxx = correction. Full history in CHANGELOG.md.
SCENES[]), CTA attached automatically and Ken Burns camera motion in every scene — the end of the slideshow.M.* (reveal/sweep/type/float/pulse/glow/countUp) instead of improvised tweens, with smaller movements in 9:16.TRANS in GSAP: fade (default), push, slideUp, zoom, wipe, fadeBlack — special effects only at 2–3 key moments.-FINAL.MUSIC_VOL ~0.14), scrim/blur/panel beneath text, and what counts as a variation.chatterbox-vc); Kokoro pf_dora automatically becomes the fallback for each scene.TITLE = { l1, l2 } — persona + curiosity, fixed at the top of the vertical frame and disappearing at the CTA, closing the retention loop.edge-tts or Kokoro (with required voice validation) and images via Agnes AI (US$ 0) instead of flux2-klein.