🤖 What is agent-browser
A browsing skill that controls a real browser (using Playwright under the hood) through command-line commands — it opens the actual app and captures the screens that become the video.
agent-browser doesn't simulate the screen — it uses the app. Each screenshot is a real capture of the DOM in that step's state, at the viewport you set.
That's why this skill shows a real app, instead of explaining a concept with motion graphics.
- ✓ Opens a real URL in a controlled browser
- ✓ Lists interactive elements with
snapshot -i - ✓ Fills in fields, clicks, and takes screenshots of each state
- ✓ Runs JS via
eval(e.g., reading the bbox)
- ✗ Doesn't render the video (that's HyperFrames, in T3)
- ✗ Doesn't record screen video (that would be v3 mode, on the roadmap)
- ✗ No API key or cloud required
- ✗ Doesn't animate anything — only captures the actual state
localhost:8000.capture.mjs runs in Node and calls agent-browser.🔗 Open a session and navigate to the URL
Set the viewport, open the URL, and establish the state. The order matters: viewport before to open.
Each command acts on the same open tab. You navigate, act, capture, and measure—step by step—without reopening the app for each call.
📑 @ref references and snapshot
O snapshot -i maps the elements and assigns refs (@e1, @e2…) — but they shift when the DOM changes.
Each button, field, and link gets a ref (@e6 = textarea, for example).
agent-browser fill @e6 "texto" fills in that specific element.
After “Generate,” a download link appears and the refs shift. Take a new snapshot or switch to a stable selector.
Don't trust @eN when the state changes between steps. On screens that change a lot, prefer CSS selectors or {tag,text} — topic of module 2.3.
🖱️ Actions: click, fill, scroll, screenshot
The basic vocabulary of each step. Each action changes the screen state — and each state becomes a screenshot.
| Action | What it does | Status |
|---|---|---|
| fill | Fills in an input/textarea (CSS selector) | works |
| click | Click an element (CSS selector) | works |
| screenshot | Saves the current screen state as PNG | works |
| eval | Runs JS in the browser (e.g., read the bbox) | works |
| scroll | Scroll to the target before the screenshot | manual / roadmap |
O capture.mjs doesn't scroll the page on its own. On long pages, you currently drive agent-browser manually (scrollIntoView). It’s backlog item #1 — detailed in module 2.4.
📸 Capture a screenshot of the screen
A screenshot for each state, captured at assets/shots/NN-id.png. It's the basis of the golden rule: capture first, animate later.
HyperFrames rendering is deterministic (no network access during rendering). That’s why the live site is never loaded in the video: we capture real screenshots beforehand and animate over them.
📐 Measure an element’s bounding box
Via eval, read the getBoundingClientRect() of the target. It’s the exact box the cursor aims at in the video.
The result goes in data.result (not in result). O capture.mjs it already handles that—but it’s useful to know how to drive it manually.
The actual bbox (not "eyeballed") is what makes the cursor land exactly on the button/field. Mapping the bbox to the video canvas is the topic of module 2.3.
🖼️ The fixed viewport and why it matters
The capture viewport is the coordinate space for everything. Keep it the same for bounding boxes and screenshots — otherwise, the cursor will miss its target.
- ✓ The same viewport for the bbox and screenshot
- ✓ Width ≤ ~1280 to fit the 16:9 canvas
- ✓ If you recapture, recapture everything in the same viewport
- ✓
set viewportbeforeopen
- ✗ Capture bounding box and shot in different viewports
- ✗ Very wide screens appear squeezed or overflow
- ✗ Recapture just one screen and mix it with the old ones
- ✗ Open the app and only then set the viewport
With viewport 1280×800 and the window centered, the screenshot spans x320..1600 and y148..948 — within the canvas 1920×1080. Change the viewport and check that it doesn’t extend past the edges.
🎯 Module summary
- ✓ agent-browser = Playwright on the command line; uses the real app
- ✓ viewport BEFORE open;
snapshot -igives volatile @eN refs - ✓ actions: fill, click, screenshot, eval (scroll is still manual/roadmap)
- ✓ 1 shot per state in
assets/shots/NN-id.png - ✓ bbox via
getBoundingClientRect; fixed viewport = coordinate space
capture.mjs generate the shots and the steps.json.