Trail map
Detailed content
🎯 Demo vs. Explainer
The skill video-demonstrativo SHOWS a real app in use — it doesn’t explain a concept with motion graphics. Here you’ll learn what a walkthrough is, when to use one, and what the skill delivers.
A walkthrough is a video that goes through a web application step by step, showing real app screens in use — clicking, filling in fields, generating, seeing the result.
It's a format that demonstrates how something works in practice, without the user having to open the app themselves.
Real app, step by step, screen by screen, narrated.
video-explicativo explain a concept with motion graphics (animated scenes). video-demonstrativo shows a real app being navigated for real.
Choosing the wrong skill costs time. Abstract concept → explainer; concrete tool → demo.
Showing vs. explaining; real screen vs. motion graphics.
If there’s a screen to show (an app, a localhost, or a URL), it’s a demo. If it’s an idea with no interface, it’s an explainer.
A single criterion removes uncertainty and points you to the right workflow right from the start.
"Is there a screen?" = demo; "Is it a concept?" = explainer.
Product onboarding, a tutorial for a new feature, a launch video, and support material are natural uses for the walkthrough.
Recognizing the use case helps you define the right sequence of steps.
Onboarding, feature, launch, support.
Real screenshots go inside a browser frame (bar + URL), with an animated cursor that clicks the controls, zooms in on the result, and TTS narration.
These four elements turn static screenshots into a video that looks like a professional screen recording.
Browser frame, cursor, highlight/zoom, narration.
The output is an MP4 in 16:9 (app screens are landscape), with narration in PT-BR and the INEMA.CLUB CTA in the final scene.
Knowing the natural format avoids asking for 9:16 (which would require cropping) and sets expectations.
16:9, MP4, narrated, final CTA.
The starting point is simply the app link — a public URL or a localhost:8000. The app must be live so the capture can navigate it.
Understanding that the input is the link (not a script file) changes how you prepare the work.
App link, localhost, app live.
🧠 Capture first, animate later
The principle behind the whole skill. HyperFrames rendering is deterministic — with no network access at render time. That’s why real screenshots are captured beforehand, and the fixed capture viewport becomes the cursor’s coordinate space.
HyperFrames renders the HTML into frames deterministically: no fetch or network requests during rendering.
It's the technical constraint from which everything else in the principle stems.
Deterministic, offline, frame by frame.
Because rendering does not access the network, the app is never loaded live inside the video (not even in an iframe) — that would be inconsistent and would break.
Prevents the wrong instinct to “embed the site” and points you toward capturing it.
No live content, no app iframe, just an image.
Before assembling the video, navigate the actual app and take a real screenshot for each state (initial screen, after filling it in, after the result…).
These are the screenshots that go inside the frame—the visual foundation of everything.
1 shot per state, before rendering, real screens.
The fixed viewport used for capture (e.g., 1280×800) becomes the coordinate system used to position everything afterward.
If the viewport changes between the shot and the box measurement, the cursor misses the target.
Fixed viewport, same size, ≤ ~1280 wide.
The position of each target element comes from getBoundingClientRect — coordinates relative to the viewport, not "by eye".
It's the actual bounding box that makes the cursor land exactly on the control, giving it a professional look.
getBoundingClientRect, viewport-relative, actual box.
The cursor animates toward the center of each bounding box. To get it right, the screenshot and box must come from the same viewport and scroll position.
Shot↔coordinate consistency is what keeps the cursor aligned with what appears on screen.
Cursor targets the box, same viewport, same scroll position.
🧰 The API-key-free stack
Three local components do it all: agent-browser (Playwright) to capture, HyperFrames (HTML→MP4) to render, and Kokoro TTS to narrate. No API key — everything runs on your machine.
agent-browser drives a browser through Playwright: it opens the URL, navigates the app, performs actions, and takes real screenshots.
It's the tool that produces the video's raw materials (shots + bounding boxes).
Playwright, automated navigation, screenshots.
HyperFrames converts an animated HTML page into an MP4, using headless Chrome to generate frames and FFmpeg to assemble the video.
It's the output engine—understanding that it's HTML→MP4 explains why the render is deterministic.
HTML→MP4, headless Chrome, FFmpeg.
Kokoro generates the narration locally in PT-BR (voice pf_dora, --speed 0.98)—one WAV per step, with no external service.
It's the audio source; the generator measures these WAVs to sync the timing.
Local TTS, pf_dora, PT-BR, WAV per step.
All three components run locally—capture, rendering, and narration don’t depend on any paid service or API key.
This means zero cost per video and no data leaving your machine.
Local, no API, zero cost, private.
There are three ways: clone the skill repo, copy the folder skills/video-demonstrativo/ to ~/.claude/skills/, or create a symlink (for development).
Because a skill is just a folder, installing it is a matter of copying and pasting — no installer.
unzip, copy the folder, symlink.
~/.claude/skills/ makes the skill global (every project); .claude/skills/ in the repository restricts it to that project.
Choosing the right scope prevents duplicating the skill or failing to find it.
Global, project, scope.
Requests like "demo video," "app demo," "walkthrough," or simply providing a link/localhost trigger the skill.
Knowing the triggers ensures the right skill kicks in (and not the explanatory one).
Trigger phrases, provide the link, walkthrough.