PTENES
MODULE 2.4

⏳ Long pages, React inputs & multi-state

The tricky cases in real demos: scrolling long pages, handling React/Vue-controlled inputs, waiting for a condition instead of a fixed time, and capturing asynchronous flows. Honestly: what already works vs. what’s on the roadmap.

7
Topics
~30
Minutes
Advanced
Level
Edge
Type
asynchronous multi-state flow COMPLETION POLLING BETWEEN STEPS analyze → status waiting_ approval dub processing completed ✓ waitFor (roadmap) — wait for a CONDITION, not a fixed amount of time poll by text / selector / status change · named substeps validated in inemaVOX: 14 steps · 2:08
🗺️
Read first: what it is today vs. what’s on the roadmap

This module covers the tricky cases. Some already work in the capture.mjs; others still require manually driving agent-browser and are in the backlog. Each topic clearly marks its status.

1

📜 Scroll before the screenshot roadmap

Long pages require scrollIntoView in the section before capturing. Today, done manually.

The problem

O capture.mjs doesn't scroll the page. On inemaVOX screens, it was necessary to drive agent-browser manually, making scrollIntoView per section before each screenshot.

📊 The planned solution (backlog #1)
Scroll / scrollTo action
Add a scroll action for each step to capture.mjs.
scrollIntoView({block})
Scroll to the target before the screenshot. The bboxes (viewport-relative) match the shot at that scroll position.
💡
It's the item that saves the most manual work

That’s why it’s backlog priority #1. Until it arrives, scrolling manually before measuring/capturing works—it just takes more effort.

2

⚛️ React/Vue-controlled inputs partial

Set .value shows the text but doesn’t trigger the state—buttons stay disabled.

✗ Don't
  • ✗ el.value = "768" via plain eval
  • ✗ React doesn't notice → button remains disabled
  • ✗ The screenshot shows the control inactive
✓ Do
  • ✓ setValue triggers input+change
  • ✓ Ideal (roadmap): fill native to Playwright
  • ✓ Check that the button is enabled in the shot
setValue triggers the events (excerpt from capture.mjs)
el.value = "768";
el.dispatchEvent(new Event('input',{bubbles:true}));
el.dispatchEvent(new Event('change',{bubbles:true}));
💡
Backlog #2: use native fill

The ideal is the capture.mjs call agent-browser fill @ref (Playwright’s native fill) instead of setting .value — simulates real typing and reliably triggers the state.

3

⏱️ Wait for a CONDITION, not a fixed amount of time roadmap

Today, capture uses wait (ms). Ideally, a waitFor text/selector.

Approach wait (ms) — today waitFor — roadmap
Criterionfixed timingreal condition
Slow generationguess the marginwaits for the <img>
Risktoo early / slowhappens at the right time
💡
Slow generation: poll a real <img>

flux2-klein took ~2.5 min in the POC. Before the result screenshot, poll for a <img> actually loaded in the panel — don’t rely on the clock alone. Changing wait by waitFor is backlog item #3.

4

🔄 Multi-state asynchronous flows handmade

Analyze → approve → dub → complete: captured as named sub-steps, with completion polling.

The inemaVOX pipeline
1
Analyze

Start processing; the status changes. Capture the "analyzing" state.

2
Approve (waiting_approval)

The flow pauses for approval. It captures the waiting state and the approval action.

3
Dub

Long process; poll for completion between captures.

4
Completed

Final result, zoomed in. Ends the walkthrough before the CTA.

⚠️
The poll was a hand-written script

In the POC, the completion poll (poll-job*.sh) was built outside capture.mjs. Embedding it as named substeps with polling is backlog item #4.

5

⏸️ States like waiting_approval

Event-triggered states (not time-triggered) need a status poll — not a wait fixed.

Why fixed timing doesn’t work

One waiting_approval changes when someone approves—not when the clock strikes. Waiting X ms may capture too early (or too late).

Status polling (concept)
// repeat until the status is no longer "waiting_approval"
while (status === "waiting_approval") {
status = readStatus(); // eval in the DOM / API
sleep(2000);
} // → then capture the next state
💡
It's what connects this topic to waitFor

O waiting_approval is the use case that justifies the waitFor of topic 3: wait for a condition (status changed) instead of a duration.

6

⚠️ Capture pitfalls

The mistakes that cost the most time—address them before rendering. Details in gotchas.md.

✗ Capture pitfalls
  • ✗ Inconsistent viewport → cursor misses the target
  • ✗ @eN refs shift after DOM changes
  • ✗ eval --json nested: data URL in data.result
  • ✗ Screenshot exceeding the canvas boundaries
✓ Fixes
  • ✓ Same viewport for bbox and shot; recapture everything
  • ✓ CSS selectors / {tag,text} + re-snapshot
  • ✓ Read data.result; decode base64 if saving
  • ✓ Viewport 1280×800 → screenshot x320..1600, y148..948
💡
Save the generated result

If the app generates a file (e.g., an image), get the src (data: URL) via eval—which comes in data.result — and decode base64 to PNG, or click the download button.

7

🗺️ What already works vs. what's on the roadmap

The skill's honest boundary: what capture.mjs does today and what still requires manual work.

✓ Works today (v1)
  • ✓ actions.json: fill, click, clickText, setValue, wait
  • ✓ Bounding boxes via getBoundingClientRect → steps.json
  • ✓ 1 shot per state + frame + cursor + zoom + CTA
  • ✓ Long, multi-state pages, driven manually
⏳ Roadmap / backlog
  • 1. Scroll/scrollIntoView in capture.mjs
  • 2. Playwright's native fill for React inputs
  • 3. waitFor (condition) instead of a fixed wait
  • 4. Embedded multi-state polling · 5. 9:16 · 6. v3 recording
⚠️
Known limitations (be honest with the user)

Dynamic state becomes a static screenshot (for real motion, use v3 and record the screen). Apps with login require test credentials. The natural format is 16:9 — 9:16 would require reframing.

🎯 Module summary

  • ✓ Long pages: manual scrollIntoView for now (scroll = backlog item #1)
  • ✓ React inputs: use setValue (triggers input/change); native fill is #2
  • ✓ Wait for a condition (waitFor) > fixed time; poll for the actual <img> (no. 3)
  • ✓ Multi-state (analyze→approve→dub→done) with polling (no. 4)
  • ✓ Apply the gotchas before rendering; always mark what’s on the roadmap
Continues in Track 3
T3
🎬 Composition & Render
With steps.json ready, T3 builds the video: frame, cursor, zoom, Kokoro narration, and HyperFrames rendering to MP4.
Go to Track 3 →