🚀 npx hyperframes init <nome>
The starting point: create an isolated HyperFrames project where the video will be assembled and rendered.
npx hyperframes init <nome> --example blank --non-interactive
Node 22+ and FFmpeg; HyperFrames' Chrome (npx hyperframes browser ensure); Kokoro (pip install kokoro-onnx soundfile) e o agent-browser on the PATH. The target app needs to be running (e.g., localhost:8000).
📦 Copy scripts + assets/fonts
The skill is self-contained: you copy the scripts and fonts into the project, and you’re ready to run.
capture.mjs ← dirige o agent-browser
composition-template.mjs → build-demo.mjs ← renomear
narration-template.sh ← gera os WAVs
assets/fonts/ ← Sora/Inter/JetBrains (.woff2 + fonts.css)
# ou, em vez de copiar fonts:
node fetch-fonts.mjs ← baixa as fontes
The skill includes its own fonts and house style. It doesn’t depend on any other project, repository, or skill to work—it generates everything from scratch in any project.
📸 node capture.mjs actions.json
The capture generates the shots and steps.json—the link to Track 2 and the composition input.
node capture.mjs actions.json # -> assets/shots/*.png + steps.json
bash narration-template.sh # -> assets/audio/sN.wav (voz pf_dora)
The bounding boxes and screenshots must come from the SAME viewport (width ≤ ~1280 to fit in 16:9)—otherwise the cursor misses its target. Action types: fill, click, clickText, setValue, wait.
For login or dynamic states, you can drive agent-browser manually (snapshot, fill, screenshot, and get the bbox via eval) and build the steps.json—the result is the same.
🧱 build-demo + lint + inspect
Generate index.html and validate it before spending time rendering — lint and inspect are cheap.
node build-demo.mjsReads steps.json, measures the WAVs, and writes the index.html (16:9).
npx hyperframes lint0 errorsFetches fonts via CDN, animation in .clip wrong and framework rules broken.
npx hyperframes inspect --samples 140 issuesSamples frames and detects layout issues (crop, off-canvas without data-layout-ignore).
🎬 render --quality high
Start with a draft to review, then render in high quality—always validating frames and narration with the user.
# 1) conferir rápido
npx hyperframes render --quality draft
# extrair 1 frame por passo e mostrar ao usuário
ffmpeg -nostdin -y -ss <t> -i video.mp4 -vframes 1 -update 1 frame.png
# 2) render final
npx hyperframes render --quality high --fps 30 \
--output renders/<nome>-16x9.mp4
Always review the frames with the user and ask them to validate the voice-over before the final render—you can’t listen to the audio on your end.
🧯 Final gotchas
The non-negotiable golden rules that save hours of debugging in the inspector.
- ✓Capture first, animate after (deterministic render)
- ✓Keep the viewport fixed = coordinate space
- ✓Animate the
.scene-inner, never the.clip - ✓Decorative elements and frame with
data-layout-ignore
- ✗Load the live site inside the video (no fetch during rendering)
- ✗Use CDN fonts — only the local ones from
assets/fonts/ - ✗Type a
AUDIO[]manual—the timing is measured - ✗Edit the
index.htmlby hand
Dynamic state (animations, video, live data) becomes a static screenshot. Apps with login require test credentials (or capture only public screens). The natural format is 16:9 — app screens are landscape.
🗺️ Roadmap (future)
What the skill doesn't do that yet — evolution items, clearly marked as future work.
- ①Scroll during capture — long pages today require manually driving agent-browser; the idea is one action
scroll+scrollIntoViewper step. - ②Inputs controlled by React/Vue — use the
fillnative to Playwright instead of setting.value. - ③Wait for a condition (
waitFor) instead of a fixed time, for asynchronous flows. - ④9:16 / Shorts — app screens are landscape; this would require reframing (zoom/pan on the active region).
- ⑤v3 — real screen recording (
agent-browser record) with narration/zoom overlaid, for apps with lots of movement. - ⑥Cursor polish — curved path (nonlinear) and "typing" (typewriter) when filling in fields.
Static screenshots + animated cursor/zoom, 16:9, local narration. Everything above is on the roadmap—don’t promise what the skill doesn’t deliver yet.
🎯 What you learned
- ✓project init + copy the skill's self-contained scripts and assets/fonts
- ✓capture.mjs generates shots + steps.json; narration generates the WAVs
- ✓build-demo.mjs generates index.html; lint + inspect validate (0/0)
- ✓render draft → high; check frames and validate the voice-over
- ✓non-negotiable gotchas and what's on the roadmap (9:16, v3, curved cursor)
You completed Track 3. From capture to narrated MP4 with an INEMA.CLUB CTA — the complete walkthrough.