⚙️ Deterministic rendering
HyperFrames renders the HTML into frames deterministically: given the same HTML and timeline, the video comes out identical. To do this, there is no network access during rendering.
Deterministic means the render doesn’t depend on anything that can vary—it doesn’t fetch, doesn’t load remote images or call APIs along the way. Each frame is a pure function of the HTML + time.
This is exactly what makes the video reproducible and ensures the animation matches the audio frame by frame.
If a frame depended on a request, the result would change based on latency, server state, or the data at that moment. The render would lose predictability—and synchronization with the narration would break.
🚫 Never the live site
As a direct consequence of deterministic rendering, the app is never loaded live inside the video. No iframe pointing to the real URL during rendering.
- ✓ Real screenshots saved in
assets/shots/ - ✓ Local images referenced in the video HTML
- ✓ Everything embedded before rendering starts
- ✓ Animation over the static image
- ✗
<iframe src="https://app…">in the video - ✗ Load images via remote URL during rendering
- ✗ Wait for the app to respond during rendering
- ✗ Rely on fonts served via CDN (use local ones)
The static screenshot is placed inside a browser frame (bar + URL) and gets a cursor + zoom overlay. To viewers, it looks like a live screen recording — but it’s an image plus animation.
📸 Capture real screenshots FIRST
The step that makes all the difference happens before rendering: the app is navigated for real with agent-browser, and a screenshot is taken for each state in the journey.
agent-browser set viewport 1280 800 and open the URL. The viewport is set once and kept fixed.
Fill in a field, click a button, wait for a result — the action demonstrated in that step.
Each state becomes a assets/shots/NN-id.png — the visual foundation for that step.
Along with the shot, the actual box of the element the cursor will target is measured. All of this becomes steps.json.
Automated (capture.mjs + actions.json) for predictable apps, or manually drive agent-browser step by step when there's a login or dynamic states. The details are in Track 2: Capture.
🖼️ The fixed viewport becomes the coordinate space
This is the most important concept in the module: the viewport used for capture defines the coordinate system everything is positioned in afterward. That’s why it’s sacred.
If the capture was made in a 1280×800 viewport, then the screenshot is 1280×800 and any coordinate (x, y) refers to that space. The cursor, zoom, and frame all use this same coordinate system.
If the screenshot comes from one viewport and the bounding box from another, the coordinates won’t match: the cursor lands in the wrong place. That’s why the viewport must be identical throughout the capture.
📦 Viewport-relative bounding boxes
The position of each target element isn’t estimated “by eye”: it comes from getBoundingClientRect, which returns actual coordinates relative to the viewport.
- ✓ The cursor lands at the exact center of the control
- ✓ Works even if the layout is responsive
- ✓ The highlight/zoom frames the right element
- ✓ It's what gives the video a professional look
- ✗ Guessed coordinates miss the control
- ✗ Break with any layout change
- ✗ Don't follow scrolling or element size
- ✗ Give it that amateurish "almost there" look
🖱️ The cursor targets the boxes — shot↔coordinate consistency
With the screenshots and bounding boxes ready, the cursor animates on the main timeline, targeting the center of each bounding box. Consistency is the key to getting it right: same viewport, same scroll position.
The cursor is global (animated on the main timeline, not per scene), with the hotspot at its tip. The tip slides to the center of the target box and triggers a pulse + ripple on click. Since the box and shot come from the same viewport and scroll position, the tip lands exactly on what appears on screen.
As long as the screenshot for that step and the target’s bounding box come from the same viewport and scroll position, the cursor will be aligned. Breaking this consistency is the #1 cause of a cursor being “out of place.”
📋 Module 1.2 summary
- ✓ HyperFrames rendering is deterministic — no network access
- ✓ That’s why the live site is never loaded in the video
- ✓ Real screenshots are captured BEFORE, 1 per state
- ✓ The fixed viewport becomes the coordinate space
- ✓ Bounding boxes come from getBoundingClientRect (viewport-relative)
- ✓ The cursor targets the boxes; consistency between shot and coordinates is the rule