Trail map
Detailed content
🌐 Drive the app with agent-browser
agent-browser (with Playwright under the hood) opens the actual app in a fixed viewport, navigates to the URL, finds elements, performs actions, and takes screenshots — it captures the real screens used in the video.
It's a browsing skill that controls a real browser (using Playwright under the hood) through simple commands: open, snapshot, fill, click, screenshot, eval.
It's the capture engine. Without it, there are no real screens—and real screens are what set this skill apart from generic motion graphics.
agent-browser, Playwright, target app running (e.g., localhost:8000), without an API key.
You set the viewport (agent-browser set viewport 1280 800) and opens the URL (agent-browser open http://localhost:8000/). The session stays alive between commands.
Order matters: viewport BEFORE opening, so the bounding boxes come from the same space as the screenshots.
set viewport, open, persistent session, fixed viewport.
agent-browser snapshot -i lists the interactive elements and assigns refs (@e1, @e2…) that you use to target actions.
The refs change after navigation or a new DOM. Take a new snapshot after major changes and prefer stable selectors when the state changes.
snapshot -i, @eN refs, volatile refs, re-snapshot.
fill fills in a field, click click, screenshot saves the screen, eval runs JS. Scrolling is currently done manually (a roadmap item in capture.mjs).
These are the building blocks of each demo step. Each action changes the screen state—and each state becomes a screenshot.
fill, click, screenshot, eval, manual scrolling (roadmap).
After each action, agent-browser screenshot assets/shots/01-prompt.png records the real screen in that state as a PNG.
It's the golden rule: "capture first, animate later"—the render is deterministic, so no live sites, just real screenshots.
1 shot per state, assets/shots/NN-id.png, capture first/animate later.
Via eval, you read the getBoundingClientRect() of the target and stores {x,y,w,h}. It's the exact box the cursor will aim for in the video.
The actual bbox is what makes the cursor land in the center of the button, not "eyeballed." It's the detail that makes it look professional.
getBoundingClientRect, {x,y,w,h}, cursor target, eval --json.
The capture viewport (e.g., 1280×800) is the coordinate space of all bboxes and screenshots. Width ≤ ~1280 to fit the 16:9 canvas.
An inconsistent viewport makes the cursor miss the target. If you recapture, recapture everything at the same viewport size.
fixed viewport, coordinate space, ≤ ~1280 wide, recapture everything.
🗺️ The actions.json
The entry file for the capture.mjs: describes the URL, viewport, and list of steps with their selectors and narration. Run node capture.mjs actions.json generates the PNG shots + the steps.json.
One JSON with url, viewport, window, eyebrow, ctaNarration, ctaCaption and the array steps. It's the script that capture.mjs runs.
It's the automated (recommended) way to capture: one file describes the entire demo from start to finish.
url, viewport, window, steps[], embedded CTA.
"url": "http://localhost:8000/" e "viewport": [1280, 800]. capture.mjs does set viewport e open exactly with these values.
The viewport defined here is the coordinate space for the bounding boxes. Change it here and it changes throughout the video.
url, viewport [W,H], coordinate space, ≤ ~1280.
Each item in steps has id, optional do, target, flags (intro, click, zoom), caption e narration. The capture runs in order.
5–8 steps + CTA ≈ 35–50s. The 1st is usually the home screen (intro:true) and the last one the result (zoom:true).
open→actions→result→CTA arc, intro, zoom, 5–8 steps.
The action to perform before the screenshot: fill (CSS selector), click, clickText ({tag,text}), setValue (triggers input/change) and wait (ms).
Without do, the step only takes a screenshot of the current state. Choosing the right type avoids capturing the wrong screen.
fill, click, clickText, setValue, wait.
target is a CSS selector (string) or {tag,text}. capture.mjs reads the target's bbox — that's where the cursor goes in the video.
O do e o target can be different elements: you can fill in one field and target another button.
target, CSS selector, {tag,text}, target bbox.
Each step carries caption (on-screen caption) and narration (the speech). Numbers and acronyms are expanded: "512" → "five hundred twelve".
The narration goes into steps.json, and from there Kokoro generates the WAVs. Phonetic text = natural voice.
caption, narration, expand numbers/acronyms, 1 sentence per step.
capture.mjs generates assets/shots/NN-id.png e o steps.json (screens + bboxes + flags + captions + narration)—everything ready for the composition-template.
steps.json is the bridge between capture (T2) and render (T3). It’s the contract between the two halves of the pipeline.
steps.json, shots/*.png, capture→render contract.
🎯 Coordinates, selectors & bounding boxes
What makes the cursor land exactly on the control: viewport-relative bounding boxes, robust selectors, targeting the center of the box, the zoom region, and how to check that it selected the right element.
The bboxes are relative to the viewport (the window’s top-left corner). In the video, they become canvasX = WIN_L + sx, canvasY = SHOT_T + sy.
It's the mapping that places the cursor over the right element in the screenshot within the browser frame.
viewport-relative, WIN_L, SHOT_T, coordinate mapping.
For each step with target, capture.mjs reads getBoundingClientRect() and rounds to {x,y,w,h} in steps.json.
If the target isn’t found, the bbox is null and capture warns — that's the sign of a wrong selector.
getBoundingClientRect, {x,y,w,h}, null target = warning.
Prefer data-testid, role or visible text ({tag,text}) instead of refs @eN — which shift when the DOM changes.
After you click "Generate," a download link appears and the refs shift. Stable selectors survive state changes.
data-testid, role, {tag,text}, avoid @eN in mutable state.
The cursor is an SVG with the hotspot at its tip (~6,3 inside 42px). The tween uses x = alvoX - 6 so the tip lands in the center of the bbox.
Aiming for the center of the box (not the edge) is what makes the click feel precise and professional.
hotspot at the tip, bbox center, cursor tween.
The bbox also positions the highlight (amber ring with glow) and, in the step zoom:true, the push-in with transformOrigin in the center of the target.
Highlight and zoom reuse the same cursor coordinates—a correct bbox works for all three effects.
highlight (.hlbox), zoom:true, transformOrigin on the target.
If the target is below the fold, you need to scroll to it before taking the screenshot and measuring the bbox (currently done manually; capture.mjs doesn’t scroll yet).
Because bboxes are viewport-relative, they match the screenshot from that scroll position — but only if the screenshot and measurement are taken after the same scroll.
scrollIntoView, target below the fold, scrolling = capture.mjs roadmap.
Compare the bbox with the screenshot: the {x,y,w,h} should land on the visible control. capture.mjs prints each bbox in the log for a quick check.
Picking the wrong element only shows up in the render. Checking the capture saves a whole rerender cycle.
check bbox against shot, capture log, validate before rendering.
⏳ Long pages, React inputs & multi-state
The tricky cases in real demos: scrolling long pages, handling React/Vue-controlled inputs, waiting for a condition instead of a fixed time, and capturing asynchronous multi-state flows. Here it’s clear what already works and what’s on the roadmap.
On long pages, you need to scrollIntoView in the section before capturing. Today, this is done by manually driving agent-browser—capture.mjs still can’t scroll.
That was exactly the case with inemaVOX. It’s backlog item no. 1: add an action scroll/scrollTo to capture.mjs.
scrollIntoView, long page, manual capture for now, scrolling is on the roadmap.
Set .value via eval shows the text but doesn't trigger the framework state — the buttons remain disabled. Use the fill native to Playwright.
O setValue from capture.mjs triggers input+change and helps, but the ideal approach (roadmap item nº 2) is the fill @ref native.
controlled input, don't set .value, setValue triggers events, fill native (roadmap).
Today, capture uses wait (ms) to spare. Ideally, a waitFor text/selector — wait for the real condition instead of hoping the timing works.
Slow generation (flux2-klein took ~2.5 min in the POC) needs polling for a <img> real before the result screenshot.
wait ms (today), waitFor (roadmap #3), polling by element.
Long pipelines (analyze → approve → dub → complete) are captured as named substeps, with completion polling between them.
Tested the skill in inemaVOX: 14 steps, 2:08. Polling was done with a handwritten script; embedding it is roadmap item #4.
named substeps, completion polling, 14-step demo.
Some flows stop in waiting states (e.g., waiting_approval) until someone approves. The capture needs to recognize this state and continue.
One wait a fixed time doesn’t solve it: the state changes when there’s approval, not when the clock strikes. Hence the need for waitFor.
waiting_approval, event-triggered state, status polling.
Refs that shift, eval --json nested (data URL in data.result), screenshot within the canvas limits, inconsistent viewport.
These are the mistakes that cost the most time. Applying these fixes BEFORE rendering avoids re-renders over minor details.
volatile refs, data.result, canvas limits, gotchas.md.
Works today: actions.json with fill/click/clickText/setValue/wait, bboxes, steps.json. Roadmap: scroll, native fill, waitFor, built-in multi-state polling, 9:16, v3.
Knowing the boundary avoids promising what the skill doesn’t do yet—and shows where you still need to operate agent-browser manually.
current v1, backlog, scroll/fill/waitFor/poll, v3 screen recording.