PTENES
TRACK 1

🧭 Fundamentals

Before capturing a screen or animating a cursor, you need to understand when the demo video is the right choice — and the principle that guides this entire skill: capture the real screen first, animate over it afterward. Here you’ll learn the difference between show e explain, why the render is deterministic, and the local stack without an API key.

3
Modules
~20
Topics
~1,5h
Duration
Basic
Level
🔗 app link localhost or URL 📸 capture real screenshots 🖱️ animate on top cursor · zoom · narration 🎬 MP4 16:9 HyperFrames · no API CAPTURE FIRST · ANIMATE AFTER rendering is deterministic — no live website inside the video

The demo video workflow — illustrative diagram

Trail map

Detailed content

1.1~30 min

🎯 Demo vs. Explainer

The skill video-demonstrativo SHOWS a real app in use — it doesn’t explain a concept with motion graphics. Here you’ll learn what a walkthrough is, when to use one, and what the skill delivers.

What it is:

A walkthrough is a video that goes through a web application step by step, showing real app screens in use — clicking, filling in fields, generating, seeing the result.

Why learn:

It's a format that demonstrates how something works in practice, without the user having to open the app themselves.

Key concepts:

Real app, step by step, screen by screen, narrated.

What it is:

video-explicativo explain a concept with motion graphics (animated scenes). video-demonstrativo shows a real app being navigated for real.

Why learn:

Choosing the wrong skill costs time. Abstract concept → explainer; concrete tool → demo.

Key concepts:

Showing vs. explaining; real screen vs. motion graphics.

What it is:

If there’s a screen to show (an app, a localhost, or a URL), it’s a demo. If it’s an idea with no interface, it’s an explainer.

Why learn:

A single criterion removes uncertainty and points you to the right workflow right from the start.

Key concepts:

"Is there a screen?" = demo; "Is it a concept?" = explainer.

What it is:

Product onboarding, a tutorial for a new feature, a launch video, and support material are natural uses for the walkthrough.

Why learn:

Recognizing the use case helps you define the right sequence of steps.

Key concepts:

Onboarding, feature, launch, support.

What it is:

Real screenshots go inside a browser frame (bar + URL), with an animated cursor that clicks the controls, zooms in on the result, and TTS narration.

Why learn:

These four elements turn static screenshots into a video that looks like a professional screen recording.

Key concepts:

Browser frame, cursor, highlight/zoom, narration.

What it is:

The output is an MP4 in 16:9 (app screens are landscape), with narration in PT-BR and the INEMA.CLUB CTA in the final scene.

Why learn:

Knowing the natural format avoids asking for 9:16 (which would require cropping) and sets expectations.

Key concepts:

16:9, MP4, narrated, final CTA.

What it is:

The starting point is simply the app link — a public URL or a localhost:8000. The app must be live so the capture can navigate it.

Why learn:

Understanding that the input is the link (not a script file) changes how you prepare the work.

Key concepts:

App link, localhost, app live.

View Full
1.2~30 min

🧠 Capture first, animate later

The principle behind the whole skill. HyperFrames rendering is deterministic — with no network access at render time. That’s why real screenshots are captured beforehand, and the fixed capture viewport becomes the cursor’s coordinate space.

What it is:

HyperFrames renders the HTML into frames deterministically: no fetch or network requests during rendering.

Why learn:

It's the technical constraint from which everything else in the principle stems.

Key concepts:

Deterministic, offline, frame by frame.

What it is:

Because rendering does not access the network, the app is never loaded live inside the video (not even in an iframe) — that would be inconsistent and would break.

Why learn:

Prevents the wrong instinct to “embed the site” and points you toward capturing it.

Key concepts:

No live content, no app iframe, just an image.

What it is:

Before assembling the video, navigate the actual app and take a real screenshot for each state (initial screen, after filling it in, after the result…).

Why learn:

These are the screenshots that go inside the frame—the visual foundation of everything.

Key concepts:

1 shot per state, before rendering, real screens.

What it is:

The fixed viewport used for capture (e.g., 1280×800) becomes the coordinate system used to position everything afterward.

Why learn:

If the viewport changes between the shot and the box measurement, the cursor misses the target.

Key concepts:

Fixed viewport, same size, ≤ ~1280 wide.

What it is:

The position of each target element comes from getBoundingClientRect — coordinates relative to the viewport, not "by eye".

Why learn:

It's the actual bounding box that makes the cursor land exactly on the control, giving it a professional look.

Key concepts:

getBoundingClientRect, viewport-relative, actual box.

What it is:

The cursor animates toward the center of each bounding box. To get it right, the screenshot and box must come from the same viewport and scroll position.

Why learn:

Shot↔coordinate consistency is what keeps the cursor aligned with what appears on screen.

Key concepts:

Cursor targets the box, same viewport, same scroll position.

View Full
1.3~30 min

🧰 The API-key-free stack

Three local components do it all: agent-browser (Playwright) to capture, HyperFrames (HTML→MP4) to render, and Kokoro TTS to narrate. No API key — everything runs on your machine.

What it is:

agent-browser drives a browser through Playwright: it opens the URL, navigates the app, performs actions, and takes real screenshots.

Why learn:

It's the tool that produces the video's raw materials (shots + bounding boxes).

Key concepts:

Playwright, automated navigation, screenshots.

What it is:

HyperFrames converts an animated HTML page into an MP4, using headless Chrome to generate frames and FFmpeg to assemble the video.

Why learn:

It's the output engine—understanding that it's HTML→MP4 explains why the render is deterministic.

Key concepts:

HTML→MP4, headless Chrome, FFmpeg.

What it is:

Kokoro generates the narration locally in PT-BR (voice pf_dora, --speed 0.98)—one WAV per step, with no external service.

Why learn:

It's the audio source; the generator measures these WAVs to sync the timing.

Key concepts:

Local TTS, pf_dora, PT-BR, WAV per step.

What it is:

All three components run locally—capture, rendering, and narration don’t depend on any paid service or API key.

Why learn:

This means zero cost per video and no data leaving your machine.

Key concepts:

Local, no API, zero cost, private.

What it is:

There are three ways: clone the skill repo, copy the folder skills/video-demonstrativo/ to ~/.claude/skills/, or create a symlink (for development).

Why learn:

Because a skill is just a folder, installing it is a matter of copying and pasting — no installer.

Key concepts:

unzip, copy the folder, symlink.

What it is:

~/.claude/skills/ makes the skill global (every project); .claude/skills/ in the repository restricts it to that project.

Why learn:

Choosing the right scope prevents duplicating the skill or failing to find it.

Key concepts:

Global, project, scope.

What it is:

Requests like "demo video," "app demo," "walkthrough," or simply providing a link/localhost trigger the skill.

Why learn:

Knowing the triggers ensures the right skill kicks in (and not the explanatory one).

Key concepts:

Trigger phrases, provide the link, walkthrough.

View Full
← All tracks Track 2: Capture →