PTENES
The Producer Β· end-to-end video

From link to professional video β€” on a single production line.

Give it a link or topic; get a narrated video in 3 layers, in 16:9 and 9:16. Everything runs locally, with no API key.

Cinema (image with veil + parallax/ken-burns), kinetic text (GSAP), and topic illustration (icon/diagram/counter). That's what separates professional video from a slideshow.
What it is

An orchestrator, not another video generator

The video producer coordinates the pieces that already exist in the ecosystem β€” plan, script, text review, voice, image, render β€” in a single production line. The skill is self-contained: it includes the tested blueprint for the 3-layer generator, distilled from the reference case "Hormozi 12 dicas".

🎬 3 layers of truth

With the server up, generate the illustrations (fixed seeds = reproducible). Without a server, SVG mode takes over. Check every image: flux sometimes puts text in the frame.

πŸ”Œ Everything local, no API key

Kokoro for voice, local flux2-klein for images, HyperFrames (Chrome + FFmpeg) for rendering. If the image server is down, SVG mode takes over automatically.

πŸ“± 16:9 and 9:16 in the same build

The same generator outputs both formats; vertical video respects social media safe zones (bottom and right side kept clear of the app UI).

Compose 3 layers

The skill is at v0.2.0 and already produces end-to-end video using the tested template. The complete backlog, with decisions still open, lives in

Layer 1 (cinema): hook artwork generated with flux2-klein using a fixed seed, with a veil + ken-burns applied during rendering. AUDIO[] β€” the actual WAV durations measured by ffprobe β€” is the single source of truth for timing: it governs the scene, animation, and audio all at once.

Link / Topic→ Plan (vpe)→ Script→ Text review→ Voice + timing→ Layer 1 artwork→ Compose the 3 layers and validate→ Validate→ Render 16:9 + 9:16
1

Video producer cover β€” premium dark amber artwork for the INEMA video orchestrator

Depth-filled background: image generated with local flux2-klein, with a dark overlay + parallax/ken-burns. This layer gives the frame weight.

2

Kinetic text

Animated number, title, bar, and emphasis with GSAP. The motion comes from code β€” the image never needs to move on its own.

3

Topic illustration

An icon, micro-diagram, or counter that SHOWS what’s being narrated instead of merely decorating the screen.

πŸ–ΌοΈ Image is the default, SVG is the fallback

The generator queries localhost:8000/healthThe assembly line (always in this order)

πŸ—£οΈ Two forms per sentence

On-screen text = accented PT-BR with English in its original spelling. Speech (txt/sN.txt) is the functional reference from which the skill was distilled: 15 narrated scenes with Kokoro, 3 layers, 16:9 and 9:16, running with both images and the SVG fallback. The repo also storesdeploy β†’ "deplΓ³i").

Prerequisites

The pre-flight check before any video

Node, FFmpeg, HyperFrames, and internet access during rendering (GSAP comes via CDN) are required. Kokoro is required for voice. The image server and the vpe are optional β€” without them, the pipeline continues in SVG mode.

🎞️ Render engine

HyperFrames checks for Chrome + FFmpeg. Required.

# check the engine
npx hyperframes doctor

πŸ—£οΈ Local TTS (Kokoro)

Local narration with the voice pf_dora: server up β†’ flux2-klein; down β†’ animated SVG icons by theme. Never gets stuck due to missing images.

# check the TTS
python3 -c "import kokoro_onnx, soundfile"

πŸ–ΌοΈ Image server

flux2-klein on port 8000 (layer 1). Optional β€” there’s an SVG fallback.

# optional
curl -s localhost:8000/health

βš™οΈ Runtime and utilities

Node 22+, FFmpeg, and the vpe ) = expanded numbers/acronyms and English rewritten phonetically (

node -v
ffmpeg -version | head -1
which vpe
User guide Β· step by step

From topic to MP4 in eight steps

Start the guideskills/videoprodutor/).

1

Pre-flight

Single contract (

npx hyperframes doctor                      # render engine (Chrome+FFmpeg)
curl -s localhost:8000/health               # image β€” optional, SVG fallback available
python3 -c "import kokoro_onnx, soundfile"  # Local TTS
node -v ; ffmpeg -version | head -1 ; which vpe
2

Edit plan

O vpe builds the skeleton (preset, beats, topics, CTA). Fill in the JSON and validate it.

vpe scaffold "<assunto>" --preset <suave|promo|vendas|viral|acao> \
    --title "<tΓ­tulo>" > plano-edicao.json
vpe validate                                   # valid plano-edicao.json
3

Script

Write the SCRIPT.md and one text file per scene. In the spoken form, spell out numbers and acronyms: "10Γ—" becomes "ten times", "R$500 mil" becomes "quinhentos mil reais".

# one txt per scene
SCRIPT.md
assets/txt/s1.txt  assets/txt/s2.txt  ...  assets/txt/sN.txt
4

Text review (before voiceover)

Check accents word by word. On-screen text = accented PT-BR + English in its original spelling; speech = phonetic English. A wrong accent taints the on-screen text e the voiceover. Checklist and lexicon in references/revisao-texto.md.

# screen  (layer 2 + caption):  "did the funnel deploy"
# speech  (assets/txt/sN.txt):   "did the funnel deploy"
5

Sources, voice, and timing

Each step delivers a concrete artifact for the next. The array AUDIO[], a code-level A/B test of what the skill

node scripts/fetch-fonts.mjs          # First time: Sora/Inter/JetBrains

for i in $(seq 1 N); do npx hyperframes tts "assets/txt/s$i.txt" \
  --voice pf_dora --speed 0.98 --output "assets/audio/s$i.wav"; done

for i in $(seq 1 N); do ffprobe -v error -show_entries format=duration \
  -of default=noprint_wrappers=1:nokey=1 "assets/audio/s$i.wav"; done
# β†’ paste the durations into the generator's AUDIO[]
6

Download the .woff2 files the first time, generate the WAVs with Kokoro, and measure the durations β€” they become the array

Start simple: one pilot scene, approve the style, and only then scale up. All commands below are the actual skill commands (

node scripts/gen-imgs.mjs   # flux2-klein β€” or let SVG mode take over
7

Confirm the dependencies before spending time on the script.

Pending decisions scripts/composition-template.mjs how build-index.mjs and run it. Without a flag, the mode is AUTO; --svg/--noimg/--img force. Validate before the expensive render.

node build-index.mjs                  # 16:9 (AUTO image→SVG)
node build-index.mjs --vertical       # 9:16 with safe zones

npx hyperframes lint                  # 0 errors
npx hyperframes inspect --samples 16  # 0 overflow
# check the frames of a draft before rendering
8

Final render in both formats

Rebuild + render in each format. Since I can’t hear the audio here, ask the user to validate the narration.

node build-index.mjs            && npx hyperframes render --quality high \
  --output renders/<nome>-16x9.mp4
node build-index.mjs --vertical && npx hyperframes render --quality high \
  --output renders/<nome>-9x16.mp4
Examples

The case that became the blueprint

The video "Hormozi 12 dicas" (videos/hormozi-12-dicas/, the single source of timing. experiments/remotion-ab/. Required for audio. remotion-best-practices changes in the result.

Opening frame of the Hormozi 12 tips video β€” cinematic layer with a dark overlay
Layer 3 (topic illustration): artwork for one of the 12 tips β€” an image faithful to the narrated concept, checked one by one before inclusion.
Topic illustration from the Hormozi 12 tips video
Cinema layer
Roadmap

Where it is and where it’s headed

Open docs/06-decisoes-pendentes-e-backlog.md.

Done
v0.1.0 β€” the executable templateSelf-contained link/topic β†’ video orchestrator: 3 layers, default image with automatic SVG fallback, safe zones in 9:16, single-source timing via AUDIO[].
Done
v0.2.0 β€” text reviewReview step between script and voice: a two-form contract for each sentence (on-screen Γ— spoken) and pronunciation of English terms, closing the gap where raw text went straight to TTS.
P0
Factory blueprintCopyplano-edicao.json extended with kinetic captions, illustration per beat, duration_mode, and assets), b-roll integrator with seed caching, plan-directed compositor.
P1
Quality of layers 2 and 3More kinetic text components (typewriter, char-stagger, count-up), micro-diagrams and spotlight, plus a unified design system for tokens/eases/icons.
P2
Switch themeA variety of transitions, light leaks, and real LUTs, musical beat sync, 1:1 format, post-production, and OpenTimelineIO export.
Finishing touches
From topic to MP4 in eight stepsLong-term rendering engine (HyperFrames today Γ— unified Remotion), word-level timing (forced alignment Γ— proportional estimate), and default voice.