📍 Viewport-relative coordinates
The bboxes are relative to the viewport. In the video, they become canvas coordinates by adding the window offset.
The capture runs in a fixed viewport (e.g., 1280×800). Every bbox comes from this space. On the 1920×1080 canvas, the screenshot is drawn inside the window and each point is mapped.
📐 Capture the bbox for each target
For each step with target, capture.mjs reads the rect and rounds to {x,y,w,h}.
If the selector doesn’t match, capture records target: null and prints ! alvo não encontrado no passo X. It's a clear sign the selector is wrong — fix it before rendering.
🧷 Robust selectors
data-testid, role, and visible text survive state changes. The refs @eN no.
- ✓
[data-testid="gerar"] - ✓
role+ accessible name - ✓ Visible text:
{tag:"button",text:"Gerar"} - ✓ Attributes:
input[name=height]
- ✗ Refs
@eNon screens that change - ✗ Generated classes (CSS-in-JS) that change during the build
- ✗ positional nth-child in a dynamic layout
- ✗ Text with accent/space that doesn't match exactly
In the POC, when the "download PNG" link appeared, all the refs shifted. So, in a mutable state, use CSS selectors or {tag,text} — and take a new snapshot when the DOM changes significantly.
🎯 The cursor aims at the center of the box
The cursor is a global SVG with the hotspot at its tip. The tween compensates for the offset so the tip lands in the center.
A single #cursor animated on the main timeline — moves continuously as the screenshots change.
The pointer sits at ~(6,3) inside the 42px SVG. The tween uses x = alvoX - 6, y = alvoY - 3.
scale:.82 yoyo + one #ripple amber that expands over the target at the moment of the click.
The movement uses duration:.7, ease:"power3.inOut", starting ~0.35s after the scene appears. The cursor reaches the target before the main explanation.
🔍 Zoom / highlight region
The same bbox positions the highlight ring and the zoom push-in. One correct coordinate serves all three effects.
transformOrigin in the center of the target.O .hlbox is just box-shadow (ring + glow): comes in with back.out and pulses, drawing the eye to the active control without covering anything.
📜 Handle a target that’s off-screen
If the target is below the fold, scroll to it before capturing and measuring—and do both after the same scroll.
capture.mjs doesn't scroll the page today. On long screens, control agent-browser manually (scrollIntoView) before the screenshot. Add scroll/scrollTo is backlog item #1.
Because the bbox is measured in the current viewport, it matches the screenshot from that scroll position — as long as the screenshot and measurement are taken after of the same scrollIntoView.
✅ Check that you selected the right element
Compare the bbox with the screenshot before rendering. Capture prints each bbox in the log for a quick check.
- ✓ Bounding box falls over the visible control in the shot
- ✓ No "target not found" warning
- ✓ plausible w/h (not 0×0 or the entire screen)
- ✗ target=— where there should have been a bbox
- ✗ Bounding box over the wrong element
- ✗ Coordinates outside the shot bounds
🎯 Module summary
- ✓ viewport-relative bbox → canvasX = WIN_L + sx, canvasY = SHOT_T + sy
- ✓ capture reads getBoundingClientRect; no target match = bbox null + warning
- ✓ prefer data-testid / role / visible text over refs @eN
- ✓ global cursor with hotspot at the tip; click = pulse + ripple
- ✓ the same bbox is used for the cursor, highlight, and zoom; check the log before rendering