PTENES
Skill · Video / AI · v1.1.0

From the app link to narrated walkthrough.

You provide the URL. Claude Code navigates the actual app, captures real screens step by step, and creates a video with a browser frame, animated cursor, and narration. All on your machine, with no API key.

video-demonstrativo project cover — web application demo video
What it is

A skill that shows the real app in use

This isn't motion graphics explaining a concept (that's the skill video-explicativo): here are the video screens are your app, actually captured with an automated browser. The result is a 35–50s, 16:9 walkthrough in PT-BR that ends with the INEMA.CLUB CTA.

🖱️ Cursor that hits the button

During capture, the skill takes the real bounding box of each target with getBoundingClientRect. The animated cursor lands in the center of the control—with a pulse and ripple on click. That’s what makes it look like a professional recording, not a screenshot with an arrow over it.

🔒 100% local, no API key

Capture via agent-browser (Playwright), HTML→MP4 rendering via HyperFrames, PT-BR narration via TTS Kokoro (voice pf_dora) running on the machine itself. No paid calls, no external services.

⏱️ Single source of timing

The generator reads the steps.json and measures the duration of each WAV with ffprobe. There is no hand-built timing table: audio and animation are created in sync, and the ambience loops prevent a silent tail at the end.

How it works

Capture first, animate later

HyperFrames rendering is deterministic — with no network access during rendering. That's why the site is never loaded live inside the video: first, real screenshots are captured, then animation is added on top. The fixed capture viewport becomes the coordinate space the cursor aims at.

1 · Script (STEPS.md)→ 2 · Text review→ 3 · Capture→ 4 · Project→ 5 · Narration→ 6 · Composition→ 7 · Validate→ 8 · Render

Script and text

5–8 steps + CTA (≈35–50s). Each sentence has two ways: screen (accented PT-BR, app buttons in the original spelling — Generate, Upload) e speech (expanded numbers and English rewritten phonetically — upload → “upload”). Review comes before capture and narration.

Capture and coordinates

O capture.mjs read a actions.json, directs the agent-browser through the app, take 1 screenshot per state and record each target's bounding box. Output: assets/shots/*.png + steps.json. You can do it manually when the app is unpredictable (login, dynamic states).

Composition and rendering

O build-demo.mjs read the steps.json, measures the WAVs, and assembles a browser frame + global cursor + highlight + zoom on the result + CTA. Then: lint, inspect, a draft to check the frames and final render at high / 30fps.

Prerequisites

What needs to be on the machine

The skill is self-contained (it already includes the sources in assets/fonts/) and does not depend on any other project. It only needs the local runtime and the target app running.

Node 22+ and FFmpeg

HyperFrames foundation (HTML→MP4 rendering) and ffprobe, which measures the narration durations.

node --version   # needs to be 22+
ffmpeg -version

HyperFrames Chrome

Headless browser used for rendering. Downloads once and stays cached.

npx hyperframes browser ensure

Kokoro TTS (PT-BR)

Local narration, voice pf_dora. The first run downloads ~340MB of model data.

pip install kokoro-onnx soundfile

agent-browser in PATH

Navigation skill (Playwright) that performs the actions and takes real screenshots.

agent-browser set viewport 1280 800

The target app running

The URL needs to be available during capture — localhost or public. An app with login requires test credentials.

# e.g.: your app serving at
http://localhost:8000/

The installed skill

Copy skills/video-demonstrativo/ from this repo to your Claude Code skills.

cp -r skills/video-demonstrativo \
  ~/.claude/skills/
User guide · step by step

From STEPS.md to MP4

In practice, you ask Claude Code for the video, and it guides the workflow. Below is what happens behind the scenes — the actual commands, in the order the skill runs them.

1

Write the step-by-step script (STEPS.md)

The list of actions to demonstrate, with 1 narration sentence per step. Arc: open the app → action 1 → action 2 → … → result → CTA. The 1st step is the home screen (intro:true); the last content item is the result (zoom:true).

# 5–8 steps + CTA ≈ 35–50s of video
# e.g.: write prompt → choose 512² → adjust height → Generate → save
2

Review the text before capturing and narrating

Check PT-BR accents word by word. Set the two forms for each sentence: screen (caption + labels, English in the original spelling) and speech (txt/sN.txt, English phonetically). Kokoro phonemizes based on the written spelling—an incorrect accent affects both the screen and narration.

# screen:  "2 · Choose the size — 512²"     (Generate, Upload in the original spelling)
# speech:  "Then, choose the size. Let's go with five hundred and twelve."
# English→PT lexicon: upload→âploud · deploy→deplói · Generate→djenereit
3

Capture the real app

Describe the URL, viewport, and steps in a actions.json; the script opens the app, performs each action, takes a screenshot of the state, and gets the target's actual bounding box. Actions: fill, click, clickText, setValue, wait.

node capture.mjs actions.json   # -> assets/shots/*.png + steps.json

# or manually, when the app is unpredictable (login, dynamic state):
agent-browser set viewport 1280 800
agent-browser open http://localhost:8000/
agent-browser snapshot -i                 # discover refs @e1, @e2...
agent-browser fill @e6 "texto"
agent-browser screenshot assets/shots/01-prompt.png
4

Create the video project

Everything lives in a single folder at ~/projetos/output/<nome>/: project, captures, audio files, index.html and the final MP4. Copy the fonts bundled with the skill (or run the fetch-fonts.mjs) — no CDN in the render.

cd ~/projetos/output
npx hyperframes init <nome> --example blank --non-interactive

# fonts: copy assets/fonts/ from the skill, or:
node fetch-fonts.mjs
5

Generate the narration with Kokoro

The template reads the steps.json, writes assets/txt/sN.txt (1 per step + CTA, already in the revised spoken form) and generates the WAVs. At the end, it prints the duration of each track.

bash narration-template.sh   # -> assets/txt/sN.txt + assets/audio/sN.wav

# under the hood, track by track:
npx hyperframes tts assets/txt/s1.txt --voice pf_dora --speed 0.98 --output assets/audio/s1.wav
6

Compose the video

Copy scripts/composition-template.mjs how build-demo.mjs. It reads the steps.json and measures the WAVs with ffprobe — browser frame, animated global cursor, highlight, zoom on the result, and INEMA.CLUB CTA are ready to use.

cp ~/.claude/skills/video-demonstrativo/scripts/composition-template.mjs build-demo.mjs
node build-demo.mjs   # -> index.html (16:9). Do not edit by hand.
7

Validate before rendering

Goal: 0 lint errors and 0 issues in inspect. Animate the .scene-inner (never the wrapper .clip), scenes and captions on alternating tracks, decorative elements and a frame with data-layout-ignore.

npx hyperframes lint                  # 0 errors
npx hyperframes inspect --samples 14  # 0 issues
8

Render (draft → high)

First, make a draft: extract 1 frame per step and show it to the user — you can't hear the audio, so they validate the narration. Once approved, move on to the final render. The MP4 is saved in the project's own root directory.

# check
npx hyperframes render --quality draft
ffmpeg -nostdin -y -ss 12 -i video.mp4 -vframes 1 -update 1 frame.png

# final
npx hyperframes render --quality high --fps 30 --output <nome>-16x9.mp4
Examples

How a step is described

O actions.json is the capture input; the steps.json is what comes out of it and feeds the composition. Complete examples at scripts/actions.example.json e scripts/steps.example.json — the reference example is a walkthrough of inemaimg (image-generation playground) creating an image from scratch.

Input · actions.json

You describe the action and the cursor target. The script handles the rest.

{
  "url": "http://localhost:8000/",
  "viewport": [1280, 800],
  "window": { "urlLabel": "localhost:8000" },
  "steps": [{
    "id": "size",
    "do": { "type": "clickText", "tag": "button", "text": "512²" },
    "target": { "tag": "button", "text": "512²" },
    "click": true,
    "caption": "2 · Escolha o tamanho — 512²",
    "narration": "Depois, escolha o tamanho. Vamos de quinhentos e doze."
  }]
}

Output · steps.json

The target is now a real coordinate in the screenshot's space. That's what the cursor follows.

{
  "viewport": [1280, 800],
  "steps": [
    { "shot": "00-home.png", "intro": true, "target": null },
    { "shot": "02-size.png", "click": true,
      "target": { "x": 179, "y": 501, "w": 44, "h": 26 } },
    { "shot": "04-result.png", "zoom": true,
      "target": { "x": 668, "y": 128, "w": 484, "h": 324 } }
  ]
}

🎓 Want the full course?

There is an INEMA.CLUB course about this skill — 3 tracks, 10 modules, from the principle "capture first, animate later" to the final render: inematds.github.io/skill-video-demonstrativo.

Known limitations

What this skill doesn't do (yet)

Be honest before rendering: if your case falls into this category, it's better to know now.

Static screen

The app's dynamic state (animations, video, live data) becomes static screenshot. Real motion would require recording the screen video—a different approach, with harder narration sync.

Login

An app with authentication needs test credentials. Without them, capture only the public screens.

16:9 only

Natural format, since app screens are landscape. 9:16 would require cropping and reframing each shot.

Unacted voice

Kokoro is good, but it doesn't act. And you can't hear the result — the user always validates the narration.

Roadmap

Where it came from and where it’s going

Versioning v1.yy.xxx — yy = feature, xxx = fix. Full history in the CHANGELOG.

1.0.0
Initial releaseNarrated walkthrough of a web app: real capture with agent-browser + HyperFrames + TTS Kokoro, no API key needed. Browser frame, global cursor targeting the actual bounding box, highlight/zoom, and INEMA.CLUB CTA. The principle is "capture first, animate later." Single output at ~/projetos/output/<nome>/, with no silent tail (ambientRepeat).
1.1.0 · current
Text review + English pronunciationA new step before capture and narration, closing the gap where text went straight to the screen and TTS without review. A contract with two forms per sentence (on-screen vs. spoken), an English→PT glossary, and the reference revisao-texto.md.
v3 · future
Real motion (screen recording)Alternative path already mapped in the limitations: record the screen with agent-browser record instead of screenshots, to capture animations and live data. The challenge is syncing the narration.