🌐 agent-browser (Playwright)
The first piece is the capture. agent-browser drives a browser via Playwright: it opens the URL in a fixed viewport, navigates the app, performs actions, and takes real screenshots with their bounding boxes.
It's agent-browser that produces the video's raw materials: assets/shots/*.png (the states) and the steps.json (boxes + captions + narration). Without this step, there's nothing to animate.
agent-browser navigates the actual app — so the localhost:8000 (or URL) needs to be responding during capture. The details of the capture.mjs and of the actions.json stay in Track 2.
🎞️ HyperFrames (HTML→MP4)
The second piece is the render. HyperFrames converts an animated HTML page into an MP4 video, using headless Chrome to generate the frames and FFmpeg to assemble them.
The animated page (frame + shots + cursor + zoom + CTA), generated by build-demo.mjs.
Renders frame by frame, deterministically, with no visible interface.
Combines the frames and audio into a final 16:9 MP4.
Because the video is created from HTML rendered in headless Chrome, deterministic rendering (without network access) is the natural choice — exactly the principle from Module 1.2. Rendering details are covered in Track 3.
🗣️ Local Kokoro TTS
The third piece is the narration. Kokoro generates speech in PT-BR locally—one WAV per step (plus the CTA), with the voice pf_dora a --speed 0.98.
Kokoro runs as a local model (kokoro-onnx). The first run downloads ~340MB; after that, each video generates the WAVs without any external service. The generator measures these WAVs with ffprobe to sync the timing automatically.
- ✓ "512" → "five hundred twelve"
- ✓ "inema.club" → "inema dot club"
- ✓ Acronyms and spelled-out URLs
- ✓ 1 short sentence per step
- ✗ Leave numbers raw (they read incorrectly)
- ✗ Wait for dramatic delivery (the voice is good, but neutral)
- ✗ Long sentences that overflow the step
- ✗ English narration (always PT-BR)
Because audio cannot be heard during the generation workflow, the default is to show the user frames and ask them to confirm the narration before the final render.
🔒 No API key
All three components run locally. Capture, rendering, and narration don’t depend on any paid service or API key—which means zero cost per video and no data leaving the machine.
Node 22+ and FFmpeg; HyperFrames' Chrome (npx hyperframes browser ensure); Kokoro TTS (pip install kokoro-onnx soundfile); e o agent-browser on the PATH. All local software — no cloud credentials.
It includes its own fonts (Sora/Inter/JetBrains Mono in assets/fonts/) and the house style. It doesn't depend on any other project, repository, or service to work.
📥 Install the skill
Because a skill is just a folder, installing it is simple. There are three ways: clone the repo da skill, copy the folder, or create a symlink for development.
Whichever of the three methods you use, restart the session so Claude Code recognizes the newly installed skill.
📂 Where the skill lives
Skills live in .claude/skills. The location you choose determines the scope: global (any project) or project-specific (only in that repository).
Available in any project. Ideal for a skill you use all the time across multiple repositories.
Available only in that project and versioned with the code—good for teams.
A personal skill you use constantly → global. A skill specific to a product and shared with the team → in .claude/skills of the repository.
⚡ How to trigger the skill
Once installed, the skill is triggered by everyday phrases. Knowing these phrases ensures the demo (not the explainer) kicks in.
- • "demo video"
- • "app / system demo"
- • "walkthrough"
- • "show step by step how to use the app"
- • Provide a URL and ask for a video showing how to use it
- • Provide a
localhost:8000 - • "record the system screen"
- • "video showing how to use X"
"Make a video about X" (without an app) tends to trigger the video-explicativo. For the demo, make it clear there's an app to show — or provide the link.
Because the skill takes the app link as input, offering the URL right away communicates your intent and activates the right workflow.
📋 Module 1.3 summary
- ✓ agent-browser (Playwright) captures the shots + boxes
- ✓ HyperFrames renders HTML→MP4 (headless Chrome + FFmpeg)
- ✓ Local Kokoro TTS provides the narration (voice pf_dora, --speed 0.98)
- ✓ Everything runs on your machine — no API key, zero cost
- ✓ Install: zip, copy the folder, or symlink; global vs. project
- ✓ Trigger it with: "demo," "walkthrough," or by providing the link