🧱 actions.json structure
One JSON with the global configuration and the list of steps. It's the script that capture.mjs runs from start to finish.
The actions.json describes the entire demo in one file: where to go, at what size, what to do in each step, and what to say. A node capture.mjs actions.json and the capture is ready.
📺 URL + viewport
Where and at what size. The viewport defined here is the coordinate space for all video bboxes.
agent-browser open <url>. The app needs to be running at this address.set viewport W H before open. Width ≤ ~1280 to fit in 16:9.Because the viewport is the coordinate space for bboxes, changing it after capture breaks cursor alignment. Decide on the viewport before generating the shots.
Above ~1280 pixels wide, the screenshot comes out compressed or extends beyond the edges of the 1920×1080 canvas. Reduce the viewport during capture instead of forcing it.
📋 The list of steps
The array steps in order: each item is a screen state. 5–8 steps + CTA ≈ 35–50s.
The 1st step introduces the app, with no action — just a screenshot of the initial state.
Each step runs a do, captures the screen and measures the bbox of the target.
The final content step gently zooms in on the result. The INEMA.CLUB CTA comes from the composition-template.
Also, video is tiring, and capture/editing becomes a major task. 5–8 steps + CTA make for a ~35–50s walkthrough.
⚙️ The do:{type} field
The action performed before the screenshot. Without do, the step only captures the current state.
| do.type | Fields | What it does |
|---|---|---|
| fill | selector, value | Fills in an input/textarea (CSS selector) |
| click | selector | Click (CSS selector) |
| clickText | tag, text | Click the <tag> whose text === text |
| setValue | selector, value | Sets the value and triggers input/change |
| wait | ms | Wait (e.g., slow generation) |
That was the case with the height field: fill wasn't enough. The setValue triggers input+change for the framework to respond. (Native Playwright fill is on the roadmap — module 2.4.)
🎯 The target / selector for each step
O target is what the cursor targets. It may differ from the element that the do triggers.
- ✓ CSS selector:
"textarea" - ✓ CSS by attribute:
"input[name=height]" - ✓ By text:
{tag:"button", text:"Gerar"} - ✓ capture reads the target's bbox via eval
- ✗ Target with no match → bbox
null+ warning in the log - ✗ Avoid volatile @eN on screens that change
- ✗ Text with accent/space must match exactly
- ✗ Don't confuse the click target with the cursor target
You can fill in a field (do) and targets the submit button (target). The cursor moves to the target; the action happens on the do.
🗣️ Narration text for each step
caption is the on-screen caption; narration is the spoken line. Expand numbers and acronyms so they sound natural when read aloud.
The text for each step is written in assets/txt/sN.txt and Kokoro (voice pf_dora, --speed 0.98) generates the WAVs. Phonetic text = natural voice.
📦 The output: steps.json + the PNG shots
capture.mjs generates the screenshots and steps.json — the bridge between capture (T2) and render (T3).
00-home.png, 01-prompt.png...).O composition-template.mjs (T3) reads steps.json and assembles frame + cursor + highlight + zoom + CTA, measuring the WAVs with ffprobe — unified timing.
🎯 Module summary
- ✓ actions.json = url + viewport + window + steps[] + CTA
- ✓ viewport defined here is the coordinate space for the bboxes
- ✓ do.type: fill / click / clickText / setValue / wait
- ✓ target = what the cursor points to (CSS selector or {tag,text})
- ✓ output: shots/*.png + steps.json (capture→render contract)