PTENES
TRACK 2

🎬 Pipeline: HTML → MP4

The hands-on path. From zero to rendered video: how HyperFrames turns animated HTML + TTS narration into an MP4 via headless Chrome and FFmpeg—all on your machine, with no API key.

5
Modules
30
Topics
~2.5h
Duration
Practical
Level
Animated HTML scenes + GSAP TTS narration Kokoro pf_dora Timeline LEAD=0.5 TAIL=0.9 HyperFrames Headless Chrome + FFmpeg MP4 16:9 YouTube / horizontal MP4 9:16 Shorts / vertical 100% local without an API key

HyperFrames pipeline — HTML + narration + timeline → MP4 16:9 and 9:16

Track map

Detailed content

2.1~30 min

🎬 What HyperFrames Is

First principles: you write animated HTML, HyperFrames captures it frame by frame with headless Chrome and assembles the MP4 with FFmpeg. No API key, all local.

What it is:

HyperFrames renders an HTML page scene by scene, captures each frame as a PNG image, and uses FFmpeg to assemble the final video — with synchronized WAV audio.

Why learn:

Understanding the end-to-end workflow prevents surprises: each step produces a concrete artifact (HTML, WAV, MP4) that you can inspect.

Key concepts:

HTML → PNG frames → FFmpeg → MP4; each scene is a page state.

What it is:

Chrome headless, managed by HyperFrames, renders each animation frame. FFmpeg merges the frames with the WAV audio files and generates the final MP4 file.

Why learn:

Knowing these are two separate processes helps with troubleshooting: visual issues → Chrome; audio/timing issues → FFmpeg.

Key concepts:

Puppeteer/Chrome, FFmpeg, two-stage pipeline.

What it is:

Kokoro is a TTS model that runs locally via Python (kokoro-onnx). Generates high-quality WAV files without any calls to external services.

Why learn:

Eliminates variable costs and dependence on the internet. The model (~340 MB) downloads once and stays on your machine.

Key concepts:

kokoro-onnx, pf_dora voice, ONNX runtime, local WAV.

What it is:

The entire pipeline (TTS, render, encode) runs locally. After the initial setup, you can produce videos without internet access or per-use costs.

Why learn:

Removes the fear of escalating costs and ensures the project can be reproduced on any machine with the prerequisites installed.

Key concepts:

Offline-first, zero cost per video, reproducible.

What it is:

HyperFrames supports both main formats. The same script automatically generates a 16:9 video (YouTube, horizontal) and a 9:16 video (Shorts, TikTok, Reels).

Why learn:

One production, two deliverables. The scene's CSS adapts the layout for each format via a media query or variable.

Key concepts:

1920×1080, 1080×1920, multi-format, reuse.

What it is:

Ideal for animated explainer videos with narration (tutorials, onboarding, launches). Not for live screen recordings, interviews, or face-cam videos.

Why learn:

Knowing the right scope prevents frustration. HyperFrames shines with motion graphics content + narration; it doesn’t replace live recording.

Key concepts:

Motion graphics, TTS narration, structured technical content.

View Full
2.2~30 min

🛠️ Setup & prerequisites

Install Node 22+, FFmpeg, managed Chrome, and Kokoro — run npx hyperframes doctor and make sure everything is green before you start.

What it is:

Node 22 LTS is the minimum required. On Windows, FFmpeg goes in C:\ffmpeg\bin and, in git-bash, always use ffmpeg -nostdin to avoid stdin hangs.

Why learn:

Older versions of Node can break the CLI. The flag -nostdin is a classic Windows gotcha that silently stalls rendering.

Key concepts:

Node 22+, FFmpeg on PATH, -nostdin in git-bash.

What it is:

HyperFrames uses an isolated Chrome, downloaded via npx hyperframes browser ensure. It stays in the local cache and doesn't interfere with the Chrome you use every day.

Why learn:

Using your personal Chrome profile can cause conflicts. The isolated browser ensures a clean, reproducible environment.

Key concepts:

npx hyperframes browser ensure, local cache, isolated Puppeteer.

What it is:

Install with pip install kokoro-onnx soundfile. On the first run, the model (~340 MB) is downloaded automatically and cached. Subsequent runs are offline.

Why learn:

Without Kokoro installed, narration generation fails. The download takes a while the first time—plan for it during setup.

Key concepts:

pip install kokoro-onnx soundfile, one-time download ~340 MB, local cache.

What it is:

The command npx hyperframes doctor checks Node, FFmpeg, Chrome, and Kokoro all at once and prints the status of each dependency. Everything green = ready to create.

Why learn:

Saves debugging time: instead of discovering the failure midway through rendering, you identify the problem before you start.

Key concepts:

npx hyperframes doctor, dependency checklist, quick diagnosis.

What it is:

Create the scaffolding with npx hyperframes init <nome> --example blank. Generates the folders audio/, frames/, the scripts and input HTML already have the correct structure.

Why learn:

Starting from the right template prevents structural errors that only show up at render time.

Key concepts:

npx hyperframes init, --example blank, folders audio/ e frames/.

What it is:

O design.md defines your channel's color palette, typography, and visual rules. Copy it from the skill reference into the project and reference it in the instructions to Claude.

Why learn:

Without a design.md, each video may look different. Having the file ensures visual identity consistency.

Key concepts:

design.md, house style, #0D1321 palette, visual identity.

View Full
2.3~30 min

📝 Script & TTS narration

Write SCRIPT.md with 6–9 scenes in a hook→principle→advanced→CTA arc, generate the WAVs with Kokoro, and measure the durations before composing.

What it is:

SCRIPT.md is a Markdown file with one section per scene. The ideal arc: hook (why watch) → principle (core concept) → advanced (practical detail) → CTA (next step).

Why learn:

A well-structured script before coding prevents animation rework. The narrative guides the visuals, not the other way around.

Key concepts:

SCRIPT.md, 6–9 scenes, hook→CTA arc, one idea per scene.

What it is:

With 6–9 scenes and ~100 seconds of total narration, the video runs ~1:50 — ideal for Shorts and short YouTube videos.

Why learn:

Long text in a scene creates long audio that stretches the animation beyond what's supported. Each scene should have no more than 3–4 short sentences.

Key concepts:

~100s total, 3–4 sentences per scene, attention-friendly pacing.

What it is:

TTS reads the text literally. Write "SKILL dot M D" instead of "SKILL.md", "M J S" instead of ".mjs", and "N P X" instead of "npx" — otherwise, the result sounds strange.

Why learn:

This is one of the most common gotchas. Hearing "SKILL dot MD" naturally versus the model trying to pronounce "skill-dot-md" literally makes a huge difference in perceived quality.

Key concepts:

Phonetic text, acronym expansion, review the narration aloud.

What it is:

Use the voice pf_dora with --speed 0.98 for more natural, slightly slower speech. The result is one WAV file per scene in audio/.

Why learn:

The default speed (1.0) can sound rushed. The 0.98 setting is subtle but improves clarity without losing pace.

Key concepts:

voice pf_dora, --speed 0.98, WAV per scene, folder audio/.

What it is:

After generating the WAVs, use ffprobe -show_entries format=duration in each file to get the exact duration in seconds. These values populate the array AUDIO[] in build-index.

Why learn:

Estimated durations produce videos with cut-off audio or silence at the end. ffprobe gives the exact number HyperFrames needs.

Key concepts:

ffprobe -show_entries format=duration, actual duration in seconds, AUDIO[] array.

What it is:

Beyond the default voice pf_dora, Kokoro offers pm_alex (neutral male) and pm_santa (deeper male) for variety or channel customization.

Why learn:

Choosing the voice before recording everything prevents rework. Test all three with 2–3 lines from the script and decide before generating all the WAVs.

Key concepts:

pf_dora, pm_alex, pm_santa, test before generating everything.

View Full
2.4~30 min

🎞️ Scene composition

build-index.mjs is the heart of it: AUDIO[] array with actual durations, sceneN() HTML functions, GSAP animations with anim(i,t) and CAPTIONS[] subtitles — all aligned to the same timing.

What it is:

O build-index.mjs is the generator: it reads AUDIO[], calls the scene functions, and produces the index.html final that HyperFrames will render. Copy from the template and edit it.

Why learn:

Understanding the generator's structure lets you customize it with confidence. Each part has a clear responsibility: data, scene HTML, animation.

Key concepts:

build-index.mjs, generator, copied and edited template.

What it is:

The array AUDIO[] map each scene to its WAV file and exact duration (in seconds, with decimals) obtained from ffprobe. E.g.: { file: 'audio/scene1.wav', dur: 12.34 }.

Why learn:

This is the single source of truth for timing. All animation calculations derive from the actual AUDIO[] durations — never estimate.

Key concepts:

AUDIO[], actual durations, single source of truth, ffprobe.

What it is:

Each scene is a function sceneN() that returns absolutely positioned HTML inside the video container. The elements start out invisible, and GSAP animates them.

Why learn:

Separating scene HTML into functions keeps the code organized and makes it easier to edit one scene without affecting the others.

Key concepts:

sceneN(), absolute positioning, initial opacity 0, GSAP animates.

What it is:

The function anim(i, t) receives the scene index and GSAP timeline and adds the animations. GSAP interpolates the values frame by frame, ensuring Chrome captures smooth motion.

Why learn:

GSAP is the default in HyperFrames because it guarantees deterministic timing — the same frame always looks the same, which is essential for headless capture.

Key concepts:

anim(i, t), GSAP timeline, deterministic timing, frame-perfect.

What it is:

The array CAPTIONS[] defines the caption text for each scene, displayed in the lower band of the video. It can be the full narration or a bulleted summary.

Why learn:

Captions improve accessibility and increase retention on platforms where videos play without sound by default.

Key concepts:

CAPTIONS[], lower third, accessibility, video without sound.

What it is:

Three constants control the timing of every animation: LEAD=0.5 (pause before narration), TAIL=0.9 (hold after the speech ends) and FADE=0.45 (fade duration between scenes). Changing this here affects everything uniformly.

Why learn:

Having a single source of timing is what keeps audio and animation synchronized at all times. Never hard-code seconds in scene functions.

Key concepts:

LEAD=0.5, TAIL=0.9, FADE=0.45, a single source of truth, audio+visual sync.

View Full
2.5~30 min

✅ Validate & render

Before the final render: lint, inspect, draft to check the visuals, validate with the user—and only then render at high quality at 30fps for both versions.

What it is:

The command npx hyperframes lint checks the generated HTML: durations, audio references, timeline structure, and syntax errors. Must return 0 errors before proceeding.

Why learn:

Rendering with lint errors can produce silent, cut-off videos or missing scenes. Cheap linting now avoids expensive renders later.

Key concepts:

npx hyperframes lint, 0 errors, pre-render validation.

What it is:

O npx hyperframes inspect --samples 16 opens Chrome headless, captures 16 frames distributed throughout the video, and displays thumbnails for quick visual inspection of the layout.

Why learn:

Clipped text, off-screen elements, or overlaps are only visible visually. Inspect catches them before you waste minutes rendering.

Key concepts:

npx hyperframes inspect --samples 16, thumbnails, layout issues.

What it is:

The mode --quality draft renders at a lower resolution and higher speed for a quick preview of the full video before the final render.

Why learn:

The draft lets you check the sequence, transitions, and timing without waiting for the full render. Make corrections in the draft, not in the high.

Key concepts:

--quality draft, quick preview, inexpensive iteration.

What it is:

Extract frames from the draft with FFmpeg (ffmpeg -i draft.mp4 -vf fps=1 frames/%04d.png) and share it with the user for visual validation. Claude can't hear the audio.

Why learn:

Human validation before the final render catches subjective issues (text that's too small, colors that are off, incorrect order) that lint can't detect.

Key concepts:

ffmpeg -vf fps=1, PNG frames, human validation; Claude doesn't listen.

What it is:

After approving the draft, run --quality high --fps 30 for the full render. Generates the final MP4 at full resolution, ready to upload.

Why learn:

The high render takes longer and shouldn’t be repeated. Approving the draft first ensures the rendering effort goes into the right file.

Key concepts:

--quality high --fps 30, final render, don't repeat without validation.

What it is:

Run the render twice with different format parameters (or use the flag --both if available in your project). The result is the horizontal (16:9) and vertical (9:16) files, ready for distribution.

Why learn:

YouTube and Shorts have different reach. Generating both versions in the same render maximizes distribution without reworking the content.

Key concepts:

16:9 YouTube, 9:16 Shorts, two versions, cross-platform distribution.

View Full
← All tracks Track 3: Under the Hood →