PTENES
MODULE 2.2

🗺️ The actions.json

The entry file for the capture.mjs: URL, viewport, and the list of steps with selectors and narration. Run node capture.mjs actions.json generates the PNG shots + the steps.json.

7
Topics
~30
Minutes
Practical
Level
Config
Type
actions.json URL · VIEWPORT · STEPS { } URL viewport window steps[ ] of · target caption · narration capture.mjs executes steps shots/*.png 1 per state steps.json bbox + caption node capture.mjs actions.json
1

🧱 actions.json structure

One JSON with the global configuration and the list of steps. It's the script that capture.mjs runs from start to finish.

Main Concept

The actions.json describes the entire demo in one file: where to go, at what size, what to do in each step, and what to say. A node capture.mjs actions.json and the capture is ready.

actions.json — real example (header)
{
"url": "http://localhost:8000/",
"viewport": [1280, 800],
"window": { "top": 96, "titleH": 52, "urlLabel": "localhost:8000" },
"eyebrow": "DEMONSTRAÇÃO · INEMAIMG",
"ctaNarration": "This is INEMA ponto CLUB content...",
"steps": [ ... ]
}
🔗
URL
where to go
📺
viewport
fixed size
🪟
window
frame
📋
steps[]
the steps
2

📺 URL + viewport

Where and at what size. The viewport defined here is the coordinate space for all video bboxes.

📊 What capture.mjs does with these fields
URL
Becomes agent-browser open <url>. The app needs to be running at this address.
viewport [W,H]
Becomes set viewport W H before open. Width ≤ ~1280 to fit in 16:9.
💡
Change the viewport here and it changes throughout the video

Because the viewport is the coordinate space for bboxes, changing it after capture breaks cursor alignment. Decide on the viewport before generating the shots.

⚠️
Very wide screens overflow the canvas

Above ~1280 pixels wide, the screenshot comes out compressed or extends beyond the edges of the 1920×1080 canvas. Reduce the viewport during capture instead of forcing it.

3

📋 The list of steps

The array steps in order: each item is a screen state. 5–8 steps + CTA ≈ 35–50s.

The arc of a demo
1
Initial screen (intro:true)

The 1st step introduces the app, with no action — just a screenshot of the initial state.

2
Actions 1, 2, 3...

Each step runs a do, captures the screen and measures the bbox of the target.

3
Result (zoom:true) + CTA

The final content step gently zooms in on the result. The INEMA.CLUB CTA comes from the composition-template.

💡
Keep it to 5–8 steps

Also, video is tiring, and capture/editing becomes a major task. 5–8 steps + CTA make for a ~35–50s walkthrough.

4

⚙️ The do:{type} field

The action performed before the screenshot. Without do, the step only captures the current state.

do.type Fields What it does
fillselector, valueFills in an input/textarea (CSS selector)
clickselectorClick (CSS selector)
clickTexttag, textClick the <tag> whose text === text
setValueselector, valueSets the value and triggers input/change
waitmsWait (e.g., slow generation)
Example: click by the button text
{
"id": "generate",
"of the": { "type": "clickText", "tag": "button", "text": "Generate" },
"target": { "tag": "button", "text": "Generate" }, "click": true
}
💡
setValue for JS-controlled fields

That was the case with the height field: fill wasn't enough. The setValue triggers input+change for the framework to respond. (Native Playwright fill is on the roadmap — module 2.4.)

5

🎯 The target / selector for each step

O target is what the cursor targets. It may differ from the element that the do triggers.

✓ Target types
  • ✓ CSS selector: "textarea"
  • ✓ CSS by attribute: "input[name=height]"
  • ✓ By text: {tag:"button", text:"Gerar"}
  • ✓ capture reads the target's bbox via eval
✗ Things to watch out for
  • ✗ Target with no match → bbox null + warning in the log
  • ✗ Avoid volatile @eN on screens that change
  • ✗ Text with accent/space must match exactly
  • ✗ Don't confuse the click target with the cursor target
💡
of ≠ target

You can fill in a field (do) and targets the submit button (target). The cursor moves to the target; the action happens on the do.

6

🗣️ Narration text for each step

caption is the on-screen caption; narration is the spoken line. Expand numbers and acronyms so they sound natural when read aloud.

Step with narration (real example)
{
"id": "size", "click": true,
"caption": "2 · Choose the size — 512²",
"narration": "Next, choose the size. Let's go with five hundred twelve."
}
📊 Expand it for speech
"512"
→ "five hundred twelve"
"768"
→ "seven hundred sixty-eight"
"inema.club"
→ "inema dot club"
💡
The narration goes into steps.json → Kokoro

The text for each step is written in assets/txt/sN.txt and Kokoro (voice pf_dora, --speed 0.98) generates the WAVs. Phonetic text = natural voice.

7

📦 The output: steps.json + the PNG shots

capture.mjs generates the screenshots and steps.json — the bridge between capture (T2) and render (T3).

steps.json — output (1 step)
{ "shot": "01-prompt.png",
"target": {"x":129,"y":252,"w":482,"h":96},
"click": false, "caption": "1 · Write a detailed prompt" }
📊 What the capture produces
assets/shots/*.png
1 real screenshot per state (00-home.png, 01-prompt.png...).
steps.json
viewport, window, eyebrow, CTA, and the steps with bbox + flags + caption + narration.
💡
It's the contract with the render

O composition-template.mjs (T3) reads steps.json and assembles frame + cursor + highlight + zoom + CTA, measuring the WAVs with ffprobe — unified timing.

🎯 Module summary

  • ✓ actions.json = url + viewport + window + steps[] + CTA
  • ✓ viewport defined here is the coordinate space for the bboxes
  • ✓ do.type: fill / click / clickText / setValue / wait
  • ✓ target = what the cursor points to (CSS selector or {tag,text})
  • ✓ output: shots/*.png + steps.json (capture→render contract)
Next module
2.3
🎯 Coordinates, selectors & bounding boxes
How the bbox becomes the cursor's target on the canvas, robust selectors, and the zoom region.
Go to module 2.3 →