PTENES
MODULE 1.2

🧠 Capture first, animate later

The principle behind the whole skill. How HyperFrames rendering is deterministic (offline), the live site is never loaded in the video: the actual screen is captured before and animates on top. The capture's fixed viewport becomes the cursor's coordinate space.

6
Topics
~30
Minutes
Basic
Level
Concept
Type
CAPTURE → MEASURE → ANIMATE ① CAPTURE real screenshot 1280×800 viewport ② BOUNDING BOX x,y,w,h getBounding... viewport-relative ③ CURSOR target target the box · click ⚙️ Deterministic rendering — no network during rendering the same viewport and scroll ensure shot ↔ coordinate consistency
1

⚙️ Deterministic rendering

HyperFrames renders the HTML into frames deterministically: given the same HTML and timeline, the video comes out identical. To do this, there is no network access during rendering.

Main Concept

Deterministic means the render doesn’t depend on anything that can vary—it doesn’t fetch, doesn’t load remote images or call APIs along the way. Each frame is a pure function of the HTML + time.

This is exactly what makes the video reproducible and ensures the animation matches the audio frame by frame.

⚠️
Network access during rendering = inconsistency

If a frame depended on a request, the result would change based on latency, server state, or the data at that moment. The render would lose predictability—and synchronization with the narration would break.

Key concepts
🔁
Reproducible
Same result
🚫
No network
No fetch
🎞️
Frame by frame
Controlled timing
🔗
Beat-matched audio
Exact synchronization
2

🚫 Never the live site

As a direct consequence of deterministic rendering, the app is never loaded live inside the video. No iframe pointing to the real URL during rendering.

✓ The right way
  • ✓ Real screenshots saved in assets/shots/
  • ✓ Local images referenced in the video HTML
  • ✓ Everything embedded before rendering starts
  • ✓ Animation over the static image
✗ What not to do
  • ✗ <iframe src="https://app…"> in the video
  • ✗ Load images via remote URL during rendering
  • ✗ Wait for the app to respond during rendering
  • ✗ Rely on fonts served via CDN (use local ones)
💡
The frame disguises the fact that it's a screenshot

The static screenshot is placed inside a browser frame (bar + URL) and gets a cursor + zoom overlay. To viewers, it looks like a live screen recording — but it’s an image plus animation.

3

📸 Capture real screenshots FIRST

The step that makes all the difference happens before rendering: the app is navigated for real with agent-browser, and a screenshot is taken for each state in the journey.

The pipeline order
1
Open the app in a fixed viewport

agent-browser set viewport 1280 800 and open the URL. The viewport is set once and kept fixed.

2
Run the step’s action

Fill in a field, click a button, wait for a result — the action demonstrated in that step.

3
Take 1 screenshot per state

Each state becomes a assets/shots/NN-id.png — the visual foundation for that step.

4
Get the target’s bounding box

Along with the shot, the actual box of the element the cursor will target is measured. All of this becomes steps.json.

💡
Two ways to capture

Automated (capture.mjs + actions.json) for predictable apps, or manually drive agent-browser step by step when there's a login or dynamic states. The details are in Track 2: Capture.

4

🖼️ The fixed viewport becomes the coordinate space

This is the most important concept in the module: the viewport used for capture defines the coordinate system everything is positioned in afterward. That’s why it’s sacred.

Main Concept

If the capture was made in a 1280×800 viewport, then the screenshot is 1280×800 and any coordinate (x, y) refers to that space. The cursor, zoom, and frame all use this same coordinate system.

The fixed viewport rule
# the SAME viewport for the shot and the box
viewport = 1280 × 800 # defined once
screenshot → 1280 × 800 px # same size
box.x, box.y → relative to 1280 × 800

# width ≤ ~1280 to fit in 16:9
⚠️
Changing the viewport breaks the alignment

If the screenshot comes from one viewport and the bounding box from another, the coordinates won’t match: the cursor lands in the wrong place. That’s why the viewport must be identical throughout the capture.

5

📦 Viewport-relative bounding boxes

The position of each target element isn’t estimated “by eye”: it comes from getBoundingClientRect, which returns actual coordinates relative to the viewport.

Measuring the target’s actual box
# eval in agent-browser gets the element's box
const r = el.getBoundingClientRect();
return {
x: Math.round(r.x),
y: Math.round(r.y),
w: Math.round(r.width),
h: Math.round(r.height)
}; # coordinates in screenshot space
✓ Why a real box matters
  • ✓ The cursor lands at the exact center of the control
  • ✓ Works even if the layout is responsive
  • ✓ The highlight/zoom frames the right element
  • ✓ It's what gives the video a professional look
✗ Estimating "by eye" fails
  • ✗ Guessed coordinates miss the control
  • ✗ Break with any layout change
  • ✗ Don't follow scrolling or element size
  • ✗ Give it that amateurish "almost there" look
6

🖱️ The cursor targets the boxes — shot↔coordinate consistency

With the screenshots and bounding boxes ready, the cursor animates on the main timeline, targeting the center of each bounding box. Consistency is the key to getting it right: same viewport, same scroll position.

Main Concept

The cursor is global (animated on the main timeline, not per scene), with the hotspot at its tip. The tip slides to the center of the target box and triggers a pulse + ripple on click. Since the box and shot come from the same viewport and scroll position, the tip lands exactly on what appears on screen.

Key concepts
🌐
Global cursor
Main timeline
📍
Hotspot at the end
Lands in the center
💥
Pulse + ripple
On click
🎯
Same scroll position
Box matches the shot
💡
Shot↔coordinate consistency is the golden rule

As long as the screenshot for that step and the target’s bounding box come from the same viewport and scroll position, the cursor will be aligned. Breaking this consistency is the #1 cause of a cursor being “out of place.”

📋 Module 1.2 summary

What you learned
  • ✓ HyperFrames rendering is deterministic — no network access
  • ✓ That’s why the live site is never loaded in the video
  • ✓ Real screenshots are captured BEFORE, 1 per state
  • ✓ The fixed viewport becomes the coordinate space
  • ✓ Bounding boxes come from getBoundingClientRect (viewport-relative)
  • ✓ The cursor targets the boxes; consistency between shot and coordinates is the rule
Next module
1.3
🧰 The API-key-free stack
The three local components that put the principle into practice: agent-browser for capture, HyperFrames for rendering, and Kokoro TTS for narration—all on the machine.
Go to module 1.3 →