You provide the URL. Claude Code navigates the actual app, captures real screens step by step, and creates a video with a browser frame, animated cursor, and narration. All on your machine, with no API key.

This isn't motion graphics explaining a concept (that's the skill video-explicativo): here are the video screens are your app, actually captured with an automated browser. The result is a 35–50s, 16:9 walkthrough in PT-BR that ends with the INEMA.CLUB CTA.
During capture, the skill takes the real bounding box of each target with getBoundingClientRect. The animated cursor lands in the center of the control—with a pulse and ripple on click. That’s what makes it look like a professional recording, not a screenshot with an arrow over it.
Capture via agent-browser (Playwright), HTML→MP4 rendering via HyperFrames, PT-BR narration via TTS Kokoro (voice pf_dora) running on the machine itself. No paid calls, no external services.
The generator reads the steps.json and measures the duration of each WAV with ffprobe. There is no hand-built timing table: audio and animation are created in sync, and the ambience loops prevent a silent tail at the end.
HyperFrames rendering is deterministic — with no network access during rendering. That's why the site is never loaded live inside the video: first, real screenshots are captured, then animation is added on top. The fixed capture viewport becomes the coordinate space the cursor aims at.
5–8 steps + CTA (≈35–50s). Each sentence has two ways: screen (accented PT-BR, app buttons in the original spelling — Generate, Upload) e speech (expanded numbers and English rewritten phonetically — upload → “upload”). Review comes before capture and narration.
O capture.mjs read a actions.json, directs the agent-browser through the app, take 1 screenshot per state and record each target's bounding box. Output: assets/shots/*.png + steps.json. You can do it manually when the app is unpredictable (login, dynamic states).
O build-demo.mjs read the steps.json, measures the WAVs, and assembles a browser frame + global cursor + highlight + zoom on the result + CTA. Then: lint, inspect, a draft to check the frames and final render at high / 30fps.
The skill is self-contained (it already includes the sources in assets/fonts/) and does not depend on any other project. It only needs the local runtime and the target app running.
HyperFrames foundation (HTML→MP4 rendering) and ffprobe, which measures the narration durations.
node --version # needs to be 22+ ffmpeg -version
Headless browser used for rendering. Downloads once and stays cached.
npx hyperframes browser ensureLocal narration, voice pf_dora. The first run downloads ~340MB of model data.
pip install kokoro-onnx soundfileNavigation skill (Playwright) that performs the actions and takes real screenshots.
agent-browser set viewport 1280 800The URL needs to be available during capture — localhost or public. An app with login requires test credentials.
# e.g.: your app serving at http://localhost:8000/
Copy skills/video-demonstrativo/ from this repo to your Claude Code skills.
cp -r skills/video-demonstrativo \ ~/.claude/skills/
In practice, you ask Claude Code for the video, and it guides the workflow. Below is what happens behind the scenes — the actual commands, in the order the skill runs them.
The list of actions to demonstrate, with 1 narration sentence per step. Arc: open the app → action 1 → action 2 → … → result → CTA. The 1st step is the home screen (intro:true); the last content item is the result (zoom:true).
# 5–8 steps + CTA ≈ 35–50s of video # e.g.: write prompt → choose 512² → adjust height → Generate → save
Check PT-BR accents word by word. Set the two forms for each sentence: screen (caption + labels, English in the original spelling) and speech (txt/sN.txt, English phonetically). Kokoro phonemizes based on the written spelling—an incorrect accent affects both the screen and narration.
# screen: "2 · Choose the size — 512²" (Generate, Upload in the original spelling) # speech: "Then, choose the size. Let's go with five hundred and twelve." # English→PT lexicon: upload→âploud · deploy→deplói · Generate→djenereit
Describe the URL, viewport, and steps in a actions.json; the script opens the app, performs each action, takes a screenshot of the state, and gets the target's actual bounding box. Actions: fill, click, clickText, setValue, wait.
node capture.mjs actions.json # -> assets/shots/*.png + steps.json # or manually, when the app is unpredictable (login, dynamic state): agent-browser set viewport 1280 800 agent-browser open http://localhost:8000/ agent-browser snapshot -i # discover refs @e1, @e2... agent-browser fill @e6 "texto" agent-browser screenshot assets/shots/01-prompt.png
Everything lives in a single folder at ~/projetos/output/<nome>/: project, captures, audio files, index.html and the final MP4. Copy the fonts bundled with the skill (or run the fetch-fonts.mjs) — no CDN in the render.
cd ~/projetos/output npx hyperframes init <nome> --example blank --non-interactive # fonts: copy assets/fonts/ from the skill, or: node fetch-fonts.mjs
The template reads the steps.json, writes assets/txt/sN.txt (1 per step + CTA, already in the revised spoken form) and generates the WAVs. At the end, it prints the duration of each track.
bash narration-template.sh # -> assets/txt/sN.txt + assets/audio/sN.wav # under the hood, track by track: npx hyperframes tts assets/txt/s1.txt --voice pf_dora --speed 0.98 --output assets/audio/s1.wav
Copy scripts/composition-template.mjs how build-demo.mjs. It reads the steps.json and measures the WAVs with ffprobe — browser frame, animated global cursor, highlight, zoom on the result, and INEMA.CLUB CTA are ready to use.
cp ~/.claude/skills/video-demonstrativo/scripts/composition-template.mjs build-demo.mjs node build-demo.mjs # -> index.html (16:9). Do not edit by hand.
Goal: 0 lint errors and 0 issues in inspect. Animate the .scene-inner (never the wrapper .clip), scenes and captions on alternating tracks, decorative elements and a frame with data-layout-ignore.
npx hyperframes lint # 0 errors npx hyperframes inspect --samples 14 # 0 issues
First, make a draft: extract 1 frame per step and show it to the user — you can't hear the audio, so they validate the narration. Once approved, move on to the final render. The MP4 is saved in the project's own root directory.
# check npx hyperframes render --quality draft ffmpeg -nostdin -y -ss 12 -i video.mp4 -vframes 1 -update 1 frame.png # final npx hyperframes render --quality high --fps 30 --output <nome>-16x9.mp4
O actions.json is the capture input; the steps.json is what comes out of it and feeds the composition. Complete examples at scripts/actions.example.json e scripts/steps.example.json — the reference example is a walkthrough of inemaimg (image-generation playground) creating an image from scratch.
You describe the action and the cursor target. The script handles the rest.
{
"url": "http://localhost:8000/",
"viewport": [1280, 800],
"window": { "urlLabel": "localhost:8000" },
"steps": [{
"id": "size",
"do": { "type": "clickText", "tag": "button", "text": "512²" },
"target": { "tag": "button", "text": "512²" },
"click": true,
"caption": "2 · Escolha o tamanho — 512²",
"narration": "Depois, escolha o tamanho. Vamos de quinhentos e doze."
}]
}The target is now a real coordinate in the screenshot's space. That's what the cursor follows.
{
"viewport": [1280, 800],
"steps": [
{ "shot": "00-home.png", "intro": true, "target": null },
{ "shot": "02-size.png", "click": true,
"target": { "x": 179, "y": 501, "w": 44, "h": 26 } },
{ "shot": "04-result.png", "zoom": true,
"target": { "x": 668, "y": 128, "w": 484, "h": 324 } }
]
}There is an INEMA.CLUB course about this skill — 3 tracks, 10 modules, from the principle "capture first, animate later" to the final render: inematds.github.io/skill-video-demonstrativo.
Be honest before rendering: if your case falls into this category, it's better to know now.
The app's dynamic state (animations, video, live data) becomes static screenshot. Real motion would require recording the screen video—a different approach, with harder narration sync.
An app with authentication needs test credentials. Without them, capture only the public screens.
Natural format, since app screens are landscape. 9:16 would require cropping and reframing each shot.
Kokoro is good, but it doesn't act. And you can't hear the result — the user always validates the narration.
Versioning v1.yy.xxx — yy = feature, xxx = fix. Full history in the CHANGELOG.
agent-browser + HyperFrames + TTS Kokoro, no API key needed. Browser frame, global cursor targeting the actual bounding box, highlight/zoom, and INEMA.CLUB CTA. The principle is "capture first, animate later." Single output at ~/projetos/output/<nome>/, with no silent tail (ambientRepeat).revisao-texto.md.agent-browser record instead of screenshots, to capture animations and live data. The challenge is syncing the narration.