PTENES
MODULE 1.3

🧰 The API-key-free stack

Three local components implement the principle: agent-browser (Playwright) captures, HyperFrames renders HTML→MP4, and Kokoro narrate with local TTS. None of this depends on an API key—everything runs on your machine. Here you’ll also learn how to install and run the skill.

7
Topics
~30
Minutes
Basic
Level
Practical
Type
LOCAL STACK · NO API KEY 🌐 agent-browser Playwright capture 🎞️ HyperFrames HTML → MP4 render 🗣️ Kokoro pf_dora narration 💻 everything runs on your machine · zero cost · private 🔒 zero API keys
1

🌐 agent-browser (Playwright)

The first piece is the capture. agent-browser drives a browser via Playwright: it opens the URL in a fixed viewport, navigates the app, performs actions, and takes real screenshots with their bounding boxes.

Main Concept

It's agent-browser that produces the video's raw materials: assets/shots/*.png (the states) and the steps.json (boxes + captions + narration). Without this step, there's nothing to animate.

Driving agent-browser manually
# the manual capture pattern
agent-browser set viewport 1280 800
agent-browser open http://localhost:8000/
agent-browser snapshot -i # refs @e1, @e2…
agent-browser fill @e6 "text"
agent-browser screenshot assets/shots/01-prompt.png
💡
The target app needs to be live

agent-browser navigates the actual app — so the localhost:8000 (or URL) needs to be responding during capture. The details of the capture.mjs and of the actions.json stay in Track 2.

2

🎞️ HyperFrames (HTML→MP4)

The second piece is the render. HyperFrames converts an animated HTML page into an MP4 video, using headless Chrome to generate the frames and FFmpeg to assemble them.

How HTML becomes an MP4
📄
index.html

The animated page (frame + shots + cursor + zoom + CTA), generated by build-demo.mjs.

🖥️
Headless Chrome

Renders frame by frame, deterministically, with no visible interface.

🎬
FFmpeg → MP4

Combines the frames and audio into a final 16:9 MP4.

Validate and render
# lint and inspect before the final render
npx hyperframes lint # 0 errors
npx hyperframes inspect --samples 14 # 0 issues
npx hyperframes ... --quality high --fps 30 --output renders/demo-16x9.mp4
💡
HTML→MP4 explains determinism

Because the video is created from HTML rendered in headless Chrome, deterministic rendering (without network access) is the natural choice — exactly the principle from Module 1.2. Rendering details are covered in Track 3.

3

🗣️ Local Kokoro TTS

The third piece is the narration. Kokoro generates speech in PT-BR locally—one WAV per step (plus the CTA), with the voice pf_dora a --speed 0.98.

Main Concept

Kokoro runs as a local model (kokoro-onnx). The first run downloads ~340MB; after that, each video generates the WAVs without any external service. The generator measures these WAVs with ffprobe to sync the timing automatically.

Narration best practices
✓ Expand for speech
  • ✓ "512" → "five hundred twelve"
  • ✓ "inema.club" → "inema dot club"
  • ✓ Acronyms and spelled-out URLs
  • ✓ 1 short sentence per step
✗ What to avoid
  • ✗ Leave numbers raw (they read incorrectly)
  • ✗ Wait for dramatic delivery (the voice is good, but neutral)
  • ✗ Long sentences that overflow the step
  • ✗ English narration (always PT-BR)
🔎
You don’t listen — the user validates

Because audio cannot be heard during the generation workflow, the default is to show the user frames and ask them to confirm the narration before the final render.

4

🔒 No API key

All three components run locally. Capture, rendering, and narration don’t depend on any paid service or API key—which means zero cost per video and no data leaving the machine.

Prerequisites (local)

Node 22+ and FFmpeg; HyperFrames' Chrome (npx hyperframes browser ensure); Kokoro TTS (pip install kokoro-onnx soundfile); e o agent-browser on the PATH. All local software — no cloud credentials.

What “no API” guarantees
💸
Zero cost
By video
🔐
Private
Nothing leaves the machine
✈️
Offline-friendly
After setup
📦
Self-contained
Skill provides sources
💡
The skill is self-contained

It includes its own fonts (Sora/Inter/JetBrains Mono in assets/fonts/) and the house style. It doesn't depend on any other project, repository, or service to work.

5

📥 Install the skill

Because a skill is just a folder, installing it is simple. There are three ways: clone the repo da skill, copy the folder, or create a symlink for development.

Three ways to install
# A) clone the skill repo
git clone https://github.com/inematds/video-demonstrativo

# B) copy the folder (the skill lives in skills/ inside the repo)
cp -r video-demonstrativo/skills/video-demonstrativo ~/.claude/skills/

# C) symlink (dev — edits in the repo take effect immediately)
ln -s "$(pwd)/video-demonstrativo/skills/video-demonstrativo" ~/.claude/skills/video-demonstrativo
📊 Which path to choose
clone
Want the latest version, straight from the source.
copy folder
You cloned the repo and want a stable installed copy.
symlink
You’re going to edit the skill and want to see the changes right away.
💡
Restart the session afterward

Whichever of the three methods you use, restart the session so Claude Code recognizes the newly installed skill.

6

📂 Where the skill lives

Skills live in .claude/skills. The location you choose determines the scope: global (any project) or project-specific (only in that repository).

🌐 Global
~/.claude/skills/video-demonstrativo/

Available in any project. Ideal for a skill you use all the time across multiple repositories.

📁 Project
<repo>/.claude/skills/video-demonstrativo/

Available only in that project and versioned with the code—good for teams.

💡
Choose the right scope

A personal skill you use constantly → global. A skill specific to a product and shared with the team → in .claude/skills of the repository.

7

⚡ How to trigger the skill

Once installed, the skill is triggered by everyday phrases. Knowing these phrases ensures the demo (not the explainer) kicks in.

Trigger phrases
🗣️ Direct requests
  • • "demo video"
  • • "app / system demo"
  • • "walkthrough"
  • • "show step by step how to use the app"
🔗 Or simply the link
  • • Provide a URL and ask for a video showing how to use it
  • • Provide a localhost:8000
  • • "record the system screen"
  • • "video showing how to use X"
⚠️
Be careful not to trigger the explanatory one

"Make a video about X" (without an app) tends to trigger the video-explicativo. For the demo, make it clear there's an app to show — or provide the link.

💡
Providing the link is the strongest trigger

Because the skill takes the app link as input, offering the URL right away communicates your intent and activates the right workflow.

📋 Module 1.3 summary

What you learned
  • ✓ agent-browser (Playwright) captures the shots + boxes
  • ✓ HyperFrames renders HTML→MP4 (headless Chrome + FFmpeg)
  • ✓ Local Kokoro TTS provides the narration (voice pf_dora, --speed 0.98)
  • ✓ Everything runs on your machine — no API key, zero cost
  • ✓ Install: zip, copy the folder, or symlink; global vs. project
  • ✓ Trigger it with: "demo," "walkthrough," or by providing the link
Next track
T2
📸 Capture
With the fundamentals in place, Track 2 moves into hands-on capture: actions.json, capture.mjs, steps.json, login, dynamic states, and saving the result.
Go to Track 2 →