PTENES
MODULE 2.1

🌐 Drive the app with agent-browser

agent-browser (with Playwright under the hood) opens the actual app in a fixed viewport, navigates to the URL, finds elements, performs actions, and captures each screen — it's the capture engine for this skill.

7
Topics
~30
Minutes
Practical
Level
Capture
Type
agent-browser PLAYWRIGHT · FIXED VIEWPORT localhost:8000 1280 × 800 target button getBoundingClientRect actions open · snapshot -i fill · click · screenshot shot.png real screen bbox x,y,w,h capture first · animate later · no API key
1

🤖 What is agent-browser

A browsing skill that controls a real browser (using Playwright under the hood) through command-line commands — it opens the actual app and captures the screens that become the video.

Main Concept

agent-browser doesn't simulate the screen — it uses the app. Each screenshot is a real capture of the DOM in that step's state, at the viewport you set.

That's why this skill shows a real app, instead of explaining a concept with motion graphics.

✓ What agent-browser does
  • ✓ Opens a real URL in a controlled browser
  • ✓ Lists interactive elements with snapshot -i
  • ✓ Fills in fields, clicks, and takes screenshots of each state
  • ✓ Runs JS via eval (e.g., reading the bbox)
✗ What it is NOT
  • ✗ Doesn't render the video (that's HyperFrames, in T3)
  • ✗ Doesn't record screen video (that would be v3 mode, on the roadmap)
  • ✗ No API key or cloud required
  • ✗ Doesn't animate anything — only captures the actual state
📦 Capture prerequisites
agent-browser in PATH
The navigation skill needs to be available on the command line.
Live app
The target app needs to be running, e.g.: localhost:8000.
Node 22+
O capture.mjs runs in Node and calls agent-browser.
2

🔗 Open a session and navigate to the URL

Set the viewport, open the URL, and establish the state. The order matters: viewport before to open.

Open the app (manual method)
# viewport BEFORE opening — becomes the coordinate space
agent-browser set viewport 1280 800
agent-browser open http://localhost:8000/
agent-browser snapshot -i # discovers refs @e1, @e2...
💡
The session stays alive between commands

Each command acts on the same open tab. You navigate, act, capture, and measure—step by step—without reopening the app for each call.

Equivalent in actions.json (automated method)
{
"url": "http://localhost:8000/",
"viewport": [1280, 800]
}
# capture.mjs sets the viewport + opens with these values
3

📑 @ref references and snapshot

O snapshot -i maps the elements and assigns refs (@e1, @e2…) — but they shift when the DOM changes.

Element discovery flow
1
snapshot -i lists the interactive elements

Each button, field, and link gets a ref (@e6 = textarea, for example).

2
You act based on the ref

agent-browser fill @e6 "texto" fills in that specific element.

3
Did the DOM change? Take a new snapshot

After “Generate,” a download link appears and the refs shift. Take a new snapshot or switch to a stable selector.

⚠️
Refs are volatile

Don't trust @eN when the state changes between steps. On screens that change a lot, prefer CSS selectors or {tag,text} — topic of module 2.3.

4

🖱️ Actions: click, fill, scroll, screenshot

The basic vocabulary of each step. Each action changes the screen state — and each state becomes a screenshot.

Action What it does Status
fillFills in an input/textarea (CSS selector)works
clickClick an element (CSS selector)works
screenshotSaves the current screen state as PNGworks
evalRuns JS in the browser (e.g., read the bbox)works
scrollScroll to the target before the screenshotmanual / roadmap
💡
Scrolling is still manual

O capture.mjs doesn't scroll the page on its own. On long pages, you currently drive agent-browser manually (scrollIntoView). It’s backlog item #1 — detailed in module 2.4.

5

📸 Capture a screenshot of the screen

A screenshot for each state, captured at assets/shots/NN-id.png. It's the basis of the golden rule: capture first, animate later.

Capture first, animate after

HyperFrames rendering is deterministic (no network access during rendering). That’s why the live site is never loaded in the video: we capture real screenshots beforehand and animate over them.

Capture a state
agent-browser fill @e6 "a horse galloping on the beach"
agent-browser screenshot assets/shots/01-prompt.png
# 1 shot per state → 02-size.png, 03-height.png, 04-resultado.png ...
🖼️
1 per state
one shot/step
📁
assets/shots/
NN-id.png
🎯
Deterministic
no live site
🪟
In the frame
browser look
6

📐 Measure an element’s bounding box

Via eval, read the getBoundingClientRect() of the target. It’s the exact box the cursor aims at in the video.

Read the target’s bbox
# bbox in screenshot space (viewport-relative)
agent-browser eval "(()=>{const r=
document.querySelector('textarea').getBoundingClientRect();
return{x:Math.round(r.x),y:Math.round(r.y),
w:Math.round(r.width),h:Math.round(r.height)}})()" --json
# result: {"x":129,"y":252,"w":482,"h":96}
⚠️
eval --json comes nested

The result goes in data.result (not in result). O capture.mjs it already handles that—but it’s useful to know how to drive it manually.

💡
This is what makes it look professional

The actual bbox (not "eyeballed") is what makes the cursor land exactly on the button/field. Mapping the bbox to the video canvas is the topic of module 2.3.

7

🖼️ The fixed viewport and why it matters

The capture viewport is the coordinate space for everything. Keep it the same for bounding boxes and screenshots — otherwise, the cursor will miss its target.

✓ Set up the viewport properly
  • ✓ The same viewport for the bbox and screenshot
  • ✓ Width ≤ ~1280 to fit the 16:9 canvas
  • ✓ If you recapture, recapture everything in the same viewport
  • ✓ set viewport before open
✗ Broken viewport
  • ✗ Capture bounding box and shot in different viewports
  • ✗ Very wide screens appear squeezed or overflow
  • ✗ Recapture just one screen and mix it with the old ones
  • ✗ Open the app and only then set the viewport
📊 Screenshot limitations on the canvas

With viewport 1280×800 and the window centered, the screenshot spans x320..1600 and y148..948 — within the canvas 1920×1080. Change the viewport and check that it doesn’t extend past the edges.

🎯 Module summary

  • ✓ agent-browser = Playwright on the command line; uses the real app
  • ✓ viewport BEFORE open; snapshot -i gives volatile @eN refs
  • ✓ actions: fill, click, screenshot, eval (scroll is still manual/roadmap)
  • ✓ 1 shot per state in assets/shots/NN-id.png
  • ✓ bbox via getBoundingClientRect; fixed viewport = coordinate space
Next module
2.2
🗺️ The actions.json
Describe the URL, viewport, and steps in a single file and let the capture.mjs generate the shots and the steps.json.
Go to module 2.2 →