PTENES
TRACK 2

🎥 Capture of the real app

The workflow that actually drives the app. From the link, agent-browser (with Playwright under the hood) opens the app in a fixed viewport, performs each step, captures the real screens, and measures each target’s bounding box—generating the shots PNG + o steps.json that feed the render.

4
Modules
~28
Topics
~2h
Duration
Practical
Level
actions.json URL + viewport steps + targets agent-browser Playwright · fixed viewport click · fill · screenshot shots/*.png 1 real screen per state steps.json bboxes + captions real app no API key

Capture — actions.json → agent-browser navigates the actual app → PNG shots + steps.json with bounding boxes

Trail map

Detailed content

2.1~30 min

🌐 Drive the app with agent-browser

agent-browser (with Playwright under the hood) opens the actual app in a fixed viewport, navigates to the URL, finds elements, performs actions, and takes screenshots — it captures the real screens used in the video.

What it is:

It's a browsing skill that controls a real browser (using Playwright under the hood) through simple commands: open, snapshot, fill, click, screenshot, eval.

Why learn:

It's the capture engine. Without it, there are no real screens—and real screens are what set this skill apart from generic motion graphics.

Key concepts:

agent-browser, Playwright, target app running (e.g., localhost:8000), without an API key.

What it is:

You set the viewport (agent-browser set viewport 1280 800) and opens the URL (agent-browser open http://localhost:8000/). The session stays alive between commands.

Why learn:

Order matters: viewport BEFORE opening, so the bounding boxes come from the same space as the screenshots.

Key concepts:

set viewport, open, persistent session, fixed viewport.

What it is:

agent-browser snapshot -i lists the interactive elements and assigns refs (@e1, @e2…) that you use to target actions.

Why learn:

The refs change after navigation or a new DOM. Take a new snapshot after major changes and prefer stable selectors when the state changes.

Key concepts:

snapshot -i, @eN refs, volatile refs, re-snapshot.

What it is:

fill fills in a field, click click, screenshot saves the screen, eval runs JS. Scrolling is currently done manually (a roadmap item in capture.mjs).

Why learn:

These are the building blocks of each demo step. Each action changes the screen state—and each state becomes a screenshot.

Key concepts:

fill, click, screenshot, eval, manual scrolling (roadmap).

What it is:

After each action, agent-browser screenshot assets/shots/01-prompt.png records the real screen in that state as a PNG.

Why learn:

It's the golden rule: "capture first, animate later"—the render is deterministic, so no live sites, just real screenshots.

Key concepts:

1 shot per state, assets/shots/NN-id.png, capture first/animate later.

What it is:

Via eval, you read the getBoundingClientRect() of the target and stores {x,y,w,h}. It's the exact box the cursor will aim for in the video.

Why learn:

The actual bbox is what makes the cursor land in the center of the button, not "eyeballed." It's the detail that makes it look professional.

Key concepts:

getBoundingClientRect, {x,y,w,h}, cursor target, eval --json.

What it is:

The capture viewport (e.g., 1280×800) is the coordinate space of all bboxes and screenshots. Width ≤ ~1280 to fit the 16:9 canvas.

Why learn:

An inconsistent viewport makes the cursor miss the target. If you recapture, recapture everything at the same viewport size.

Key concepts:

fixed viewport, coordinate space, ≤ ~1280 wide, recapture everything.

View Full
2.2~30 min

🗺️ The actions.json

The entry file for the capture.mjs: describes the URL, viewport, and list of steps with their selectors and narration. Run node capture.mjs actions.json generates the PNG shots + the steps.json.

What it is:

One JSON with url, viewport, window, eyebrow, ctaNarration, ctaCaption and the array steps. It's the script that capture.mjs runs.

Why learn:

It's the automated (recommended) way to capture: one file describes the entire demo from start to finish.

Key concepts:

url, viewport, window, steps[], embedded CTA.

What it is:

"url": "http://localhost:8000/" e "viewport": [1280, 800]. capture.mjs does set viewport e open exactly with these values.

Why learn:

The viewport defined here is the coordinate space for the bounding boxes. Change it here and it changes throughout the video.

Key concepts:

url, viewport [W,H], coordinate space, ≤ ~1280.

What it is:

Each item in steps has id, optional do, target, flags (intro, click, zoom), caption e narration. The capture runs in order.

Why learn:

5–8 steps + CTA ≈ 35–50s. The 1st is usually the home screen (intro:true) and the last one the result (zoom:true).

Key concepts:

open→actions→result→CTA arc, intro, zoom, 5–8 steps.

What it is:

The action to perform before the screenshot: fill (CSS selector), click, clickText ({tag,text}), setValue (triggers input/change) and wait (ms).

Why learn:

Without do, the step only takes a screenshot of the current state. Choosing the right type avoids capturing the wrong screen.

Key concepts:

fill, click, clickText, setValue, wait.

What it is:

target is a CSS selector (string) or {tag,text}. capture.mjs reads the target's bbox — that's where the cursor goes in the video.

Why learn:

O do e o target can be different elements: you can fill in one field and target another button.

Key concepts:

target, CSS selector, {tag,text}, target bbox.

What it is:

Each step carries caption (on-screen caption) and narration (the speech). Numbers and acronyms are expanded: "512" → "five hundred twelve".

Why learn:

The narration goes into steps.json, and from there Kokoro generates the WAVs. Phonetic text = natural voice.

Key concepts:

caption, narration, expand numbers/acronyms, 1 sentence per step.

What it is:

capture.mjs generates assets/shots/NN-id.png e o steps.json (screens + bboxes + flags + captions + narration)—everything ready for the composition-template.

Why learn:

steps.json is the bridge between capture (T2) and render (T3). It’s the contract between the two halves of the pipeline.

Key concepts:

steps.json, shots/*.png, capture→render contract.

View Full
2.3~30 min

🎯 Coordinates, selectors & bounding boxes

What makes the cursor land exactly on the control: viewport-relative bounding boxes, robust selectors, targeting the center of the box, the zoom region, and how to check that it selected the right element.

What it is:

The bboxes are relative to the viewport (the window’s top-left corner). In the video, they become canvasX = WIN_L + sx, canvasY = SHOT_T + sy.

Why learn:

It's the mapping that places the cursor over the right element in the screenshot within the browser frame.

Key concepts:

viewport-relative, WIN_L, SHOT_T, coordinate mapping.

What it is:

For each step with target, capture.mjs reads getBoundingClientRect() and rounds to {x,y,w,h} in steps.json.

Why learn:

If the target isn’t found, the bbox is null and capture warns — that's the sign of a wrong selector.

Key concepts:

getBoundingClientRect, {x,y,w,h}, null target = warning.

What it is:

Prefer data-testid, role or visible text ({tag,text}) instead of refs @eN — which shift when the DOM changes.

Why learn:

After you click "Generate," a download link appears and the refs shift. Stable selectors survive state changes.

Key concepts:

data-testid, role, {tag,text}, avoid @eN in mutable state.

What it is:

The cursor is an SVG with the hotspot at its tip (~6,3 inside 42px). The tween uses x = alvoX - 6 so the tip lands in the center of the bbox.

Why learn:

Aiming for the center of the box (not the edge) is what makes the click feel precise and professional.

Key concepts:

hotspot at the tip, bbox center, cursor tween.

What it is:

The bbox also positions the highlight (amber ring with glow) and, in the step zoom:true, the push-in with transformOrigin in the center of the target.

Why learn:

Highlight and zoom reuse the same cursor coordinates—a correct bbox works for all three effects.

Key concepts:

highlight (.hlbox), zoom:true, transformOrigin on the target.

What it is:

If the target is below the fold, you need to scroll to it before taking the screenshot and measuring the bbox (currently done manually; capture.mjs doesn’t scroll yet).

Why learn:

Because bboxes are viewport-relative, they match the screenshot from that scroll position — but only if the screenshot and measurement are taken after the same scroll.

Key concepts:

scrollIntoView, target below the fold, scrolling = capture.mjs roadmap.

What it is:

Compare the bbox with the screenshot: the {x,y,w,h} should land on the visible control. capture.mjs prints each bbox in the log for a quick check.

Why learn:

Picking the wrong element only shows up in the render. Checking the capture saves a whole rerender cycle.

Key concepts:

check bbox against shot, capture log, validate before rendering.

View Full
2.4~30 min

⏳ Long pages, React inputs & multi-state

The tricky cases in real demos: scrolling long pages, handling React/Vue-controlled inputs, waiting for a condition instead of a fixed time, and capturing asynchronous multi-state flows. Here it’s clear what already works and what’s on the roadmap.

What it is:

On long pages, you need to scrollIntoView in the section before capturing. Today, this is done by manually driving agent-browser—capture.mjs still can’t scroll.

Why learn:

That was exactly the case with inemaVOX. It’s backlog item no. 1: add an action scroll/scrollTo to capture.mjs.

Key concepts:

scrollIntoView, long page, manual capture for now, scrolling is on the roadmap.

What it is:

Set .value via eval shows the text but doesn't trigger the framework state — the buttons remain disabled. Use the fill native to Playwright.

Why learn:

O setValue from capture.mjs triggers input+change and helps, but the ideal approach (roadmap item nº 2) is the fill @ref native.

Key concepts:

controlled input, don't set .value, setValue triggers events, fill native (roadmap).

What it is:

Today, capture uses wait (ms) to spare. Ideally, a waitFor text/selector — wait for the real condition instead of hoping the timing works.

Why learn:

Slow generation (flux2-klein took ~2.5 min in the POC) needs polling for a <img> real before the result screenshot.

Key concepts:

wait ms (today), waitFor (roadmap #3), polling by element.

What it is:

Long pipelines (analyze → approve → dub → complete) are captured as named substeps, with completion polling between them.

Why learn:

Tested the skill in inemaVOX: 14 steps, 2:08. Polling was done with a handwritten script; embedding it is roadmap item #4.

Key concepts:

named substeps, completion polling, 14-step demo.

What it is:

Some flows stop in waiting states (e.g., waiting_approval) until someone approves. The capture needs to recognize this state and continue.

Why learn:

One wait a fixed time doesn’t solve it: the state changes when there’s approval, not when the clock strikes. Hence the need for waitFor.

Key concepts:

waiting_approval, event-triggered state, status polling.

What it is:

Refs that shift, eval --json nested (data URL in data.result), screenshot within the canvas limits, inconsistent viewport.

Why learn:

These are the mistakes that cost the most time. Applying these fixes BEFORE rendering avoids re-renders over minor details.

Key concepts:

volatile refs, data.result, canvas limits, gotchas.md.

What it is:

Works today: actions.json with fill/click/clickText/setValue/wait, bboxes, steps.json. Roadmap: scroll, native fill, waitFor, built-in multi-state polling, 9:16, v3.

Why learn:

Knowing the boundary avoids promising what the skill doesn’t do yet—and shows where you still need to operate agent-browser manually.

Key concepts:

current v1, backlog, scroll/fill/waitFor/poll, v3 screen recording.

View Full
← Track 1: Fundamentals Track 3: Composition & Rendering →