Give it a link or topic; get a narrated video in 3 layers, in 16:9 and 9:16. Everything runs locally, with no API key.

The video producer coordinates the pieces that already exist in the ecosystem β plan, script, text review, voice, image, render β in a single production line. The skill is self-contained: it includes the tested blueprint for the 3-layer generator, distilled from the reference case "Hormozi 12 dicas".
With the server up, generate the illustrations (fixed seeds = reproducible). Without a server, SVG mode takes over. Check every image: flux sometimes puts text in the frame.
Kokoro for voice, local flux2-klein for images, HyperFrames (Chrome + FFmpeg) for rendering. If the image server is down, SVG mode takes over automatically.
The same generator outputs both formats; vertical video respects social media safe zones (bottom and right side kept clear of the app UI).
Layer 1 (cinema): hook artwork generated with flux2-klein using a fixed seed, with a veil + ken-burns applied during rendering. AUDIO[] β the actual WAV durations measured by ffprobe β is the single source of truth for timing: it governs the scene, animation, and audio all at once.
Depth-filled background: image generated with local flux2-klein, with a dark overlay + parallax/ken-burns. This layer gives the frame weight.
Animated number, title, bar, and emphasis with GSAP. The motion comes from code β the image never needs to move on its own.
An icon, micro-diagram, or counter that SHOWS whatβs being narrated instead of merely decorating the screen.
The generator queries localhost:8000/healthThe assembly line (always in this order)
On-screen text = accented PT-BR with English in its original spelling. Speech (txt/sN.txt) is the functional reference from which the skill was distilled: 15 narrated scenes with Kokoro, 3 layers, 16:9 and 9:16, running with both images and the SVG fallback. The repo also storesdeploy β "deplΓ³i").
Node, FFmpeg, HyperFrames, and internet access during rendering (GSAP comes via CDN) are required. Kokoro is required for voice. The image server and the vpe are optional β without them, the pipeline continues in SVG mode.
HyperFrames checks for Chrome + FFmpeg. Required.
# check the engine npx hyperframes doctor
Local narration with the voice pf_dora: server up β flux2-klein; down β animated SVG icons by theme. Never gets stuck due to missing images.
# check the TTS python3 -c "import kokoro_onnx, soundfile"
flux2-klein on port 8000 (layer 1). Optional β thereβs an SVG fallback.
# optional curl -s localhost:8000/health
Node 22+, FFmpeg, and the vpe ) = expanded numbers/acronyms and English rewritten phonetically (
node -v ffmpeg -version | head -1 which vpe
Start the guideskills/videoprodutor/).
Single contract (
npx hyperframes doctor # render engine (Chrome+FFmpeg) curl -s localhost:8000/health # image β optional, SVG fallback available python3 -c "import kokoro_onnx, soundfile" # Local TTS node -v ; ffmpeg -version | head -1 ; which vpe
O vpe builds the skeleton (preset, beats, topics, CTA). Fill in the JSON and validate it.
vpe scaffold "<assunto>" --preset <suave|promo|vendas|viral|acao> \ --title "<tΓtulo>" > plano-edicao.json vpe validate # valid plano-edicao.json
Write the SCRIPT.md and one text file per scene. In the spoken form, spell out numbers and acronyms: "10Γ" becomes "ten times", "R$500 mil" becomes "quinhentos mil reais".
# one txt per scene
SCRIPT.md
assets/txt/s1.txt assets/txt/s2.txt ... assets/txt/sN.txt
Check accents word by word. On-screen text = accented PT-BR + English in its original spelling; speech = phonetic English. A wrong accent taints the on-screen text e the voiceover. Checklist and lexicon in references/revisao-texto.md.
# screen (layer 2 + caption): "did the funnel deploy" # speech (assets/txt/sN.txt): "did the funnel deploy"
Each step delivers a concrete artifact for the next. The array AUDIO[], a code-level A/B test of what the skill
node scripts/fetch-fonts.mjs # First time: Sora/Inter/JetBrains for i in $(seq 1 N); do npx hyperframes tts "assets/txt/s$i.txt" \ --voice pf_dora --speed 0.98 --output "assets/audio/s$i.wav"; done for i in $(seq 1 N); do ffprobe -v error -show_entries format=duration \ -of default=noprint_wrappers=1:nokey=1 "assets/audio/s$i.wav"; done # β paste the durations into the generator's AUDIO[]
Start simple: one pilot scene, approve the style, and only then scale up. All commands below are the actual skill commands (
node scripts/gen-imgs.mjs # flux2-klein β or let SVG mode take over
Pending decisions scripts/composition-template.mjs how build-index.mjs and run it. Without a flag, the mode is AUTO; --svg/--noimg/--img force. Validate before the expensive render.
node build-index.mjs # 16:9 (AUTO imageβSVG) node build-index.mjs --vertical # 9:16 with safe zones npx hyperframes lint # 0 errors npx hyperframes inspect --samples 16 # 0 overflow # check the frames of a draft before rendering
Rebuild + render in each format. Since I canβt hear the audio here, ask the user to validate the narration.
node build-index.mjs && npx hyperframes render --quality high \ --output renders/<nome>-16x9.mp4 node build-index.mjs --vertical && npx hyperframes render --quality high \ --output renders/<nome>-9x16.mp4
The video "Hormozi 12 dicas" (videos/hormozi-12-dicas/, the single source of timing. experiments/remotion-ab/. Required for audio. remotion-best-practices changes in the result.


Open docs/06-decisoes-pendentes-e-backlog.md.
AUDIO[].plano-edicao.json extended with kinetic captions, illustration per beat, duration_mode, and assets), b-roll integrator with seed caching, plan-directed compositor.