PTENES
TRACK 3 · ADVANCED

🎬 Composition & Render

The screens have already been captured. Now you turn these real screenshots into a video: browser frame, animated cursor that points to the bounding box real, zoom on the result, local narration with Kokoro, audio↔animation sync, and the final MP4 render via HyperFrames — ending with the CTA for INEMA.CLUB.

3
Modules
~22
Topics
~1,5h
Duration
Advanced
Level
shots/*.png steps.json screens + bboxes + lines frame + cursor + zoom animates over the shot Kokoro pf_dora narration WAVs measured with ffprobe build-demo.mjs builds index.html HyperFrames HTML → MP4 (render) MP4 16:9 INEMA.CLUB CTA

From capture to MP4: animate on top, narrate locally, and render with unified timing

Trail map

Detailed content

3.1~30 min

🖱️ Frame, cursor & zoom

The animation layer that makes the screenshot look professional: the browser frame around the shot, the SVG cursor that glides to the bounding box real, the click with pulse + ripple, the zoom on the result, and the premium dark amber house style.

What it is:

The persistent window #appwin stays in the bg-layer with data-layout-ignore: border, radius, large shadow, title bar with 3 dots, and a mono URL pill (lock + window.urlLabel). The screenshots go inside it at (WIN_L, SHOT_T).

Why learn:

The frame keeps the premium look and disguises the fact that the content is a static screenshot—without it, it looks like a loose screenshot.

Key concepts:

# persistent appwin; data-layout-ignore; bar + URL pill; shot inside the window.

What it is:

A single element #cursor (SVG arrow) animated on the main timeline—not by scene—slides continuously. The hotspot stays at the tip at ~(6,3), so the tween uses x = alvoX - 6, y = alvoY - 3 so the tip lands in the center of the bbox. Movement duration:.7, ease:"power3.inOut".

Why learn:

Targeting the actual bounding box (not the eye) with curved easing is what makes the cursor look professional instead of robotic.

Key concepts:

Global cursor; hotspot at the tip; real bbox; power3.inOut.

What it is:

At the moment of the click (~start + 1.15s), the cursor scale:.82 in yoyo (pulse) and a #ripple (amber circle) expands and disappears, positioned at the center of the target.

Why learn:

The visual click feedback communicates the action to the viewer — without it, the cursor just seems to pass over the button.

Key concepts:

pulse scale .82 yoyo; amber #ripple; timing ~start+1.15s.

What it is:

The highlight .hlbox is just a rectangle with box-shadow (amber ring + glow) on the mapped bbox with ~6px of clearance, coming in with back.out. In step zoom:true, o <img> does scale 1 → 1.12 with transformOrigin in the center of the target.

Why learn:

Draws the eye to the active control without covering anything and gives it a cinematic push-in at the “ta-da” moment (the result).

Key concepts:

.hlbox box-shadow; back.out; zoom:true; transformOrigin on the target.

What it is:

The fixed palette: background #0D1321, panel #1D2D44, border #3E5C76, text #F0EBD8, amber accent #FFC300 and aqua-green code #2EC4B6. Amber is the ONLY dominant highlight — the ripple, glow, and frame use it.

Why learn:

Keeping a single accent and the same background in every scene gives the channel its consistent premium identity.

Key concepts:

6 fixed colors; amber only; constant background; teal only for values.

What it is:

Sora 700–800 for titles, Inter for body text/captions, JetBrains Mono for the URL pill and values. All three come as .woff2 local (Latin subset covers PT-BR) in assets/fonts/ — o build-demo.mjs injects the fonts.css.

Why learn:

Rendering is deterministic and offline — a font loaded via CDN would fall back and break the visuals.

Key concepts:

Sora/Inter/JetBrains; local woff2; no CDN; Latin subset.

What it is:

Recommended pacing: LEAD 0.5s (scene appears before the voice), TAIL 0.7s, FADE 0.4s between scenes. The cursor reaches the target ~0.7s before the main explanation; the click lands near start+1.15s.

Why learn:

This rhythm is what makes the animation seem to react to the narration — understanding the parameters lets you adjust the video without losing sync.

Key concepts:

LEAD 0.5s; TAIL 0.7s; FADE 0.4s; cursor before the narration.

View Full
3.2~30 min

🔊 Local narration & composition

The voice and assembly: free local Kokoro TTS (voice pf_dora), generating WAVs from the steps.json, automatic duration measurement, and the build-demo.mjs that syncs everything using the same timing array and ends with the INEMA.CLUB CTA.

What it is:

Kokoro is a TTS that runs 100% on your machine (pip install kokoro-onnx soundfile), voice pf_dora in PT-BR, with --speed 0.98. The first run downloads ~340MB of the model. No API key required.

Why learn:

Free, offline narration keeps the skill self-contained; the voice sounds good but lacks performance — so the user validates the voice-over afterward.

Key concepts:

Local Kokoro; pf_dora; speed 0.98; no API.

What it is:

O narration-template.sh read steps[].narration e ctaNarration of the steps.json, writes assets/txt/sN.txt (1 per step + CTA) and calls TTS to generate assets/audio/sN.wav.

Why learn:

steps.json is the source of the narration — generating the text from it prevents discrepancies between what’s written and what’s spoken.

Key concepts:

steps[].narration; sN.txt → sN.wav; CTA as the last one.

What it is:

O build-demo.mjs runs ffprobe in each sN.wav and builds the array AUDIO[] with real durations. There’s no manually typed array of times.

Why learn:

Timing is the single source of truth: measuring the actual audio eliminates the risk of the video and voice getting out of sync.

Key concepts:

ffprobe; AUDIO[] derived; single source of truth for timing.

What it is:

O composition-template.mjs (copied as build-demo.mjs) reads the steps.json (screens + bboxes + captions + narration), measures the WAVs, and generates the index.html with frame, cursor, highlight, zoom, and CTA — everything ready to go.

Why learn:

It's the heart of the render: a single node build-demo.mjs produces renderable HTML from the captured data.

Key concepts:

build-demo.mjs; reads steps.json; generates index.html; do not edit by hand.

What it is:

Starting from AUDIO[], the generator calculates S[] with start, dur, audioStart e end per scene (adding LEAD/TAIL). The GSAP timeline and the <audio> tracks read this same S[] — audio and animation synchronized by design.

Why learn:

Sharing the time source between voice and movement is what ensures synchronization without manual adjustment.

Key concepts:

S[] derived from AUDIO[]; start/dur/audioStart; alternating tracks.

What it is:

Each step has a caption no steps.json that becomes a translucent footer caption (Inter 600), appearing and disappearing with the scene, on a track alternating with the scenes.

Why learn:

Captions make the video readable without sound (muted feeds) and reinforce the speech—written once in steps.json.

Key concepts:

caption per step; translucent footer; synchronized with the scene.

What it is:

The final scene is the CTA: “CONTINUE AT” + INEMA.CLUB (INEMA cream, .CLUB amber with glow) + 🌐 inema.club. The default narration is "This is content from INEMA dot CLUB. Visit: inema dot club." and is already included in the template.

Why learn:

It's the channel's conversion point—included by default, it ensures no video goes out without the brand callout.

Key concepts:

Automatic CTA; INEMA.CLUB; ctaNarration; cursor disappears at the CTA.

View Full
3.3~30 min

🎞️ Render in HyperFrames

The final stretch: set up the HyperFrames project, run the capture and composition, validate with lint e inspect, render at high resolution — and what’s on the roadmap (9:16, real recording, curved cursor).

What it is:

npx hyperframes init <nome> --example blank --non-interactive creates the video project folder with the base HyperFrames structure.

Why learn:

It's the starting point for every render—the project stays isolated, and you copy the skill's scripts into it.

Key concepts:

init; --example blank; --non-interactive.

What it is:

Copy to the project: capture.mjs, composition-template.mjs (such as build-demo.mjs), narration-template.sh e assets/fonts/ — or run node fetch-fonts.mjs to download them.

Why learn:

The skill is self-contained: it includes its own fonts and house style, without depending on another project.

Key concepts:

capture/build-demo/narration; assets/fonts; fetch-fonts.mjs.

What it is:

node capture.mjs actions.json opens the app in a fixed viewport, performs each action, takes 1 screenshot per state, and gets the target's real bounding box — output: assets/shots/*.png + steps.json.

Why learn:

It's the link between Track 2 (capture) and this one—without steps.json, build-demo has nothing to assemble.

Key concepts:

actions.json; fixed viewport; real bbox; shots + steps.json.

What it is:

node build-demo.mjs generates the index.html. Afterward npx hyperframes lint (0 errors) and npx hyperframes inspect --samples 14 (0 issues) validate layout and rules before spending time on rendering.

Why learn:

Lint and inspect catch CDN fonts, incorrect clips, and off-canvas elements without data-layout-ignore — cheap compared to re-rendering.

Key concepts:

build-demo.mjs; lint 0 errors; inspect --samples 14.

What it is:

Render --quality draft to check (extract 1 frame per step and show it to the user), then --quality high --fps 30 --output renders/<nome>-16x9.mp4.

Why learn:

You don’t listen to the audio in the render — always review the frames with the user and ask them to validate the voiceover before the final version.

Key concepts:

draft → high; --fps 30; check frames; user validates voice.

What it is:

The golden rules: capture first, animate later (deterministic render, no live site); fixed viewport = coordinate space; animate .scene-inner never .clip; decorative elements and frame with data-layout-ignore; local fonts.

Why learn:

These are the mistakes that most often break the render—knowing them beforehand saves hours of debugging in inspect.

Key concepts:

capture first; fixed viewport; .scene-inner; data-layout-ignore.

What it is:

Not implemented yet: 9:16/Shorts (app screens are landscape, so they’d need reframing); v3 real screen recording (agent-browser record) for apps with lots of movement; polish the curved cursor and "typing" (typewriter) in the fields.

Why learn:

Knowing what’s on the roadmap avoids promising what the skill doesn’t do yet—currently, the natural format is 16:9 with animated static screenshots.

Key concepts:

9:16 future; v3 record; curved cursor; typewriter — all on the roadmap.

View Full
← Track 2: Capture All tracks