Track map
Detailed content
🎬 What HyperFrames Is
First principles: you write animated HTML, HyperFrames captures it frame by frame with headless Chrome and assembles the MP4 with FFmpeg. No API key, all local.
HyperFrames renders an HTML page scene by scene, captures each frame as a PNG image, and uses FFmpeg to assemble the final video — with synchronized WAV audio.
Understanding the end-to-end workflow prevents surprises: each step produces a concrete artifact (HTML, WAV, MP4) that you can inspect.
HTML → PNG frames → FFmpeg → MP4; each scene is a page state.
Chrome headless, managed by HyperFrames, renders each animation frame. FFmpeg merges the frames with the WAV audio files and generates the final MP4 file.
Knowing these are two separate processes helps with troubleshooting: visual issues → Chrome; audio/timing issues → FFmpeg.
Puppeteer/Chrome, FFmpeg, two-stage pipeline.
Kokoro is a TTS model that runs locally via Python (kokoro-onnx). Generates high-quality WAV files without any calls to external services.
Eliminates variable costs and dependence on the internet. The model (~340 MB) downloads once and stays on your machine.
kokoro-onnx, pf_dora voice, ONNX runtime, local WAV.
The entire pipeline (TTS, render, encode) runs locally. After the initial setup, you can produce videos without internet access or per-use costs.
Removes the fear of escalating costs and ensures the project can be reproduced on any machine with the prerequisites installed.
Offline-first, zero cost per video, reproducible.
HyperFrames supports both main formats. The same script automatically generates a 16:9 video (YouTube, horizontal) and a 9:16 video (Shorts, TikTok, Reels).
One production, two deliverables. The scene's CSS adapts the layout for each format via a media query or variable.
1920×1080, 1080×1920, multi-format, reuse.
Ideal for animated explainer videos with narration (tutorials, onboarding, launches). Not for live screen recordings, interviews, or face-cam videos.
Knowing the right scope prevents frustration. HyperFrames shines with motion graphics content + narration; it doesn’t replace live recording.
Motion graphics, TTS narration, structured technical content.
🛠️ Setup & prerequisites
Install Node 22+, FFmpeg, managed Chrome, and Kokoro — run npx hyperframes doctor and make sure everything is green before you start.
Node 22 LTS is the minimum required. On Windows, FFmpeg goes in C:\ffmpeg\bin and, in git-bash, always use ffmpeg -nostdin to avoid stdin hangs.
Older versions of Node can break the CLI. The flag -nostdin is a classic Windows gotcha that silently stalls rendering.
Node 22+, FFmpeg on PATH, -nostdin in git-bash.
HyperFrames uses an isolated Chrome, downloaded via npx hyperframes browser ensure. It stays in the local cache and doesn't interfere with the Chrome you use every day.
Using your personal Chrome profile can cause conflicts. The isolated browser ensures a clean, reproducible environment.
npx hyperframes browser ensure, local cache, isolated Puppeteer.
Install with pip install kokoro-onnx soundfile. On the first run, the model (~340 MB) is downloaded automatically and cached. Subsequent runs are offline.
Without Kokoro installed, narration generation fails. The download takes a while the first time—plan for it during setup.
pip install kokoro-onnx soundfile, one-time download ~340 MB, local cache.
The command npx hyperframes doctor checks Node, FFmpeg, Chrome, and Kokoro all at once and prints the status of each dependency. Everything green = ready to create.
Saves debugging time: instead of discovering the failure midway through rendering, you identify the problem before you start.
npx hyperframes doctor, dependency checklist, quick diagnosis.
Create the scaffolding with npx hyperframes init <nome> --example blank. Generates the folders audio/, frames/, the scripts and input HTML already have the correct structure.
Starting from the right template prevents structural errors that only show up at render time.
npx hyperframes init, --example blank, folders audio/ e frames/.
O design.md defines your channel's color palette, typography, and visual rules. Copy it from the skill reference into the project and reference it in the instructions to Claude.
Without a design.md, each video may look different. Having the file ensures visual identity consistency.
design.md, house style, #0D1321 palette, visual identity.
📝 Script & TTS narration
Write SCRIPT.md with 6–9 scenes in a hook→principle→advanced→CTA arc, generate the WAVs with Kokoro, and measure the durations before composing.
SCRIPT.md is a Markdown file with one section per scene. The ideal arc: hook (why watch) → principle (core concept) → advanced (practical detail) → CTA (next step).
A well-structured script before coding prevents animation rework. The narrative guides the visuals, not the other way around.
SCRIPT.md, 6–9 scenes, hook→CTA arc, one idea per scene.
With 6–9 scenes and ~100 seconds of total narration, the video runs ~1:50 — ideal for Shorts and short YouTube videos.
Long text in a scene creates long audio that stretches the animation beyond what's supported. Each scene should have no more than 3–4 short sentences.
~100s total, 3–4 sentences per scene, attention-friendly pacing.
TTS reads the text literally. Write "SKILL dot M D" instead of "SKILL.md", "M J S" instead of ".mjs", and "N P X" instead of "npx" — otherwise, the result sounds strange.
This is one of the most common gotchas. Hearing "SKILL dot MD" naturally versus the model trying to pronounce "skill-dot-md" literally makes a huge difference in perceived quality.
Phonetic text, acronym expansion, review the narration aloud.
Use the voice pf_dora with --speed 0.98 for more natural, slightly slower speech. The result is one WAV file per scene in audio/.
The default speed (1.0) can sound rushed. The 0.98 setting is subtle but improves clarity without losing pace.
voice pf_dora, --speed 0.98, WAV per scene, folder audio/.
After generating the WAVs, use ffprobe -show_entries format=duration in each file to get the exact duration in seconds. These values populate the array AUDIO[] in build-index.
Estimated durations produce videos with cut-off audio or silence at the end. ffprobe gives the exact number HyperFrames needs.
ffprobe -show_entries format=duration, actual duration in seconds, AUDIO[] array.
Beyond the default voice pf_dora, Kokoro offers pm_alex (neutral male) and pm_santa (deeper male) for variety or channel customization.
Choosing the voice before recording everything prevents rework. Test all three with 2–3 lines from the script and decide before generating all the WAVs.
pf_dora, pm_alex, pm_santa, test before generating everything.
🎞️ Scene composition
build-index.mjs is the heart of it: AUDIO[] array with actual durations, sceneN() HTML functions, GSAP animations with anim(i,t) and CAPTIONS[] subtitles — all aligned to the same timing.
O build-index.mjs is the generator: it reads AUDIO[], calls the scene functions, and produces the index.html final that HyperFrames will render. Copy from the template and edit it.
Understanding the generator's structure lets you customize it with confidence. Each part has a clear responsibility: data, scene HTML, animation.
build-index.mjs, generator, copied and edited template.
The array AUDIO[] map each scene to its WAV file and exact duration (in seconds, with decimals) obtained from ffprobe. E.g.: { file: 'audio/scene1.wav', dur: 12.34 }.
This is the single source of truth for timing. All animation calculations derive from the actual AUDIO[] durations — never estimate.
AUDIO[], actual durations, single source of truth, ffprobe.
Each scene is a function sceneN() that returns absolutely positioned HTML inside the video container. The elements start out invisible, and GSAP animates them.
Separating scene HTML into functions keeps the code organized and makes it easier to edit one scene without affecting the others.
sceneN(), absolute positioning, initial opacity 0, GSAP animates.
The function anim(i, t) receives the scene index and GSAP timeline and adds the animations. GSAP interpolates the values frame by frame, ensuring Chrome captures smooth motion.
GSAP is the default in HyperFrames because it guarantees deterministic timing — the same frame always looks the same, which is essential for headless capture.
anim(i, t), GSAP timeline, deterministic timing, frame-perfect.
The array CAPTIONS[] defines the caption text for each scene, displayed in the lower band of the video. It can be the full narration or a bulleted summary.
Captions improve accessibility and increase retention on platforms where videos play without sound by default.
CAPTIONS[], lower third, accessibility, video without sound.
Three constants control the timing of every animation: LEAD=0.5 (pause before narration), TAIL=0.9 (hold after the speech ends) and FADE=0.45 (fade duration between scenes). Changing this here affects everything uniformly.
Having a single source of timing is what keeps audio and animation synchronized at all times. Never hard-code seconds in scene functions.
LEAD=0.5, TAIL=0.9, FADE=0.45, a single source of truth, audio+visual sync.
✅ Validate & render
Before the final render: lint, inspect, draft to check the visuals, validate with the user—and only then render at high quality at 30fps for both versions.
The command npx hyperframes lint checks the generated HTML: durations, audio references, timeline structure, and syntax errors. Must return 0 errors before proceeding.
Rendering with lint errors can produce silent, cut-off videos or missing scenes. Cheap linting now avoids expensive renders later.
npx hyperframes lint, 0 errors, pre-render validation.
O npx hyperframes inspect --samples 16 opens Chrome headless, captures 16 frames distributed throughout the video, and displays thumbnails for quick visual inspection of the layout.
Clipped text, off-screen elements, or overlaps are only visible visually. Inspect catches them before you waste minutes rendering.
npx hyperframes inspect --samples 16, thumbnails, layout issues.
The mode --quality draft renders at a lower resolution and higher speed for a quick preview of the full video before the final render.
The draft lets you check the sequence, transitions, and timing without waiting for the full render. Make corrections in the draft, not in the high.
--quality draft, quick preview, inexpensive iteration.
Extract frames from the draft with FFmpeg (ffmpeg -i draft.mp4 -vf fps=1 frames/%04d.png) and share it with the user for visual validation. Claude can't hear the audio.
Human validation before the final render catches subjective issues (text that's too small, colors that are off, incorrect order) that lint can't detect.
ffmpeg -vf fps=1, PNG frames, human validation; Claude doesn't listen.
After approving the draft, run --quality high --fps 30 for the full render. Generates the final MP4 at full resolution, ready to upload.
The high render takes longer and shouldn’t be repeated. Approving the draft first ensures the rendering effort goes into the right file.
--quality high --fps 30, final render, don't repeat without validation.
Run the render twice with different format parameters (or use the flag --both if available in your project). The result is the horizontal (16:9) and vertical (9:16) files, ready for distribution.
YouTube and Shorts have different reach. Generating both versions in the same render maximizes distribution without reworking the content.
16:9 YouTube, 9:16 Shorts, two versions, cross-platform distribution.