PTENES
MODULE 1.1

🎯 Demo vs. Explainer

The golden rule of this module: show the app, don’t explain the concept. The skill video-demonstrativo navigates a real web application and turns it into a narrated walkthrough — unlike its sibling video-explicativo, which animates concepts.

7
Topics
~30
Minutes
Basic
Level
Theory
Type
TWO SISTER SKILLS 🖱️ DEMONSTRATION localhost:8000 real app screen shows the app being used 🎬 EXPLAINER motion graphics explain a concept 🎯 Has a screen? → demo Is it a concept without an interface? → explainer
1

🎬 What is a walkthrough

A walkthrough — or demo video — goes through a real web application, screen by screen, showing the app in use: opening, clicking, filling in fields, generating, seeing the result. Nothing is drawn; everything is captured from the actual screen.

Main Concept

A walkthrough answers "how do I use this?" showing the response in the interface itself. Instead of describing buttons and fields, it appears to interact with them—with a cursor that lands exactly on the control and narration explaining each step.

In practice: 5 to 8 steps plus the final CTA make a video about 35 to 50 seconds long. The typical arc is abrir o app → ação 1 → ação 2 → … → resultado → CTA.

Anatomy of a step in the script (STEPS.md)
# each step = one action + one narration sentence
Step 1 (intro) — app home screen
narration: "This is INEMA’s image generator."
Step 2 — write the prompt
narration: "We write what we want to generate."
Step 3 — click Generate
narration: "And we click Generate."
Final step (zoom) — the result
narration: "Done: the image appears in seconds."
💡
The first step introduces, the last celebrates

The opening step is usually the home screen (marked intro:true) to present the app. The last content step is usually the result (zoom:true), with a gentle zoom that highlights the result.

Key concepts
🖥️
Real app
Real screen
👣
Step by step
5–8 steps
🗣️
Narrated
1 sentence/step
⏱️
~35–50s
Typical duration
2

⚖️ Demonstrative vs. explanatory

These are two sibling skills with the same rendering engine (HyperFrames) and the same CTA, but opposite purposes. One showcases a real app; the other explains an abstract concept with animated scenes.

The two skills side by side
Aspect 🖱️ video-demonstrativo 🎬 video-explicativo
ObjectiveShow an app in useExplain a concept
VisualReal screenshotsAnimated motion graphics
InputApp link / localhostOne subject / topic
Captureagent-browser navigates the appNo capture
Format16:9 (landscape screens)16:9 e 9:16
✓ It's demonstrative when...
  • ✓ There’s an app/site with a navigable screen
  • ✓ Did the user provide a link or localhost
  • ✓ Want to show how a feature works in practice
  • ✓ The focus is the product, not the idea behind it
🎬 It’s an explainer when...
  • → The topic is abstract, with no interface to show
  • → You want to illustrate with diagrams and animations
  • → Need a vertical version (Shorts/Reels)
  • → The focus is educating viewers about a concept
💡
Same DNA, opposite purposes

Both end with the INEMA.CLUB CTA and use local TTS + HyperFrames. The difference is where the visuals come from: captured screen (demonstrative) versus drawn scene (explanatory).

3

🤔 When to use each one

There’s a one-question test that determines the skill in seconds — and clears up any doubt before you even start the script.

The one-question test

"Is there a screen to show?"

If the answer is yes—there’s an app, a localhost, or a navigable URL—it’s demonstrative. If there's no interface, just an idea to illustrate, it's explanatory.

Decision tree
1
Did the user provide a link or localhost?

If so, that’s the strongest sign it’s a demo. The skill’s own entry is the app link.

2
Does the request say "demo," "walkthrough," or "show how to use"?

Demo verbs (“show,” “record the screen,” “step by step using”) point to the demo flow.

3
Is the request "explain X," "video about Y," with no app?

That’s explanatory—the content will be motion graphics, not a screen recording.

✓
If in doubt, ask: "is there an app live?"

A single question to the user resolves ambiguous cases at no cost.

💡
App needs to be live

Demonstration requires the app to be accessible during capture (e.g., localhost:8000 running). If the app isn't live, there's nothing to navigate — confirm this before promising the video.

4

📋 Use cases

A walkthrough shines whenever the goal is to show a product in action. Four situations cover most requests.

Four classic use cases
🚪
Product onboarding

A new user sees the app's first steps: sign-up, main screen, first valuable action. It reduces friction for someone who's just arrived.

✨
Feature tutorial

Demonstrates a specific feature end to end — useful when a new feature needs a quick “how it works.”

🚀
Launch

An announcement video that shows the real product in action at launch, with the INEMA.CLUB CTA at the end.

🛟
Support

Answers “how do I do X?” with a short video showing the exact path on screen—ready to attach to a support ticket or help center.

📊 Validated in real apps
Simple app
Image generation — 7-step flow, ~42s. The skill’s first POC.
Complex app
AI dubbing suite—a 14-step tutorial, 2:08, with React-controlled input and an asynchronous multi-state workflow.
💡
The more concrete the use case, the better the script

Knowing whether it’s onboarding, a feature, a launch, or support helps you choose which steps go into the video—and which “result” deserves the final zoom.

5

🖱️ What the skill provides

Four elements turn static screenshots into a video that looks like a professional screen recording: a browser frame, an animated cursor, highlights/zoom, and narration.

Main Concept

Real screenshots go inside a browser frame (bar + URL) to look like a real window. On top, a global animated cursor aims at the real bounding box of each control, clicks with a pulse + ripple, and the result gets a gentle zoom. TTS narration ties it all together.

The four elements
🪟
Frame
Bar + URL
🖱️
Cursor
Aim at the actual box
🔍
Highlight/zoom
In the result
🗣️
Narration
PT-BR TTS
✓ What the skill does
  • ✓ Cursor lands exactly on the control (real box, not by eye)
  • ✓ Click becomes a visible pulse + ripple
  • ✓ A smooth zoom highlights the final result
  • ✓ Timing syncs with WAVs measured by ffprobe
✗ What the skill does NOT do (honest limitations)
  • ✗ Dynamic state (video, live data) becomes a static screenshot
  • ✗ Doesn't load the live site inside the video
  • ✗ App with login requires test credentials
  • ✗ Real movement only in v3 mode (roadmap, not implemented yet)
6

📐 Output format

The output is an MP4 in 16:9, narrated in PT-BR and ending with the INEMA.CLUB CTA. The landscape format isn’t aesthetic—it’s technical: app screens are wide.

The output in numbers
16:9
Aspect ratio
1920×1080
MP4
Container
via FFmpeg
30
FPS
render high
PT-BR
Narration
Local Kokoro
The final CTA (default)

Every output ends on the same scene: "CONTINUE AT" + INEMA.CLUB (INEMA in cream, .CLUB in amber with glow) + 🌐 inema.club. The narration ends with "This is content from INEMA dot CLUB. Visit: inema dot club." It’s already included in the composition template.

⚠️
9:16 isn't the natural format

Because app screens are landscape, the template generates 16:9. A vertical version would require cropping/reframing the active region — it's on the roadmap, not part of the standard workflow. Don't promise Shorts for a walkthrough without aligning on the reframing.

7

🔗 The input is the link

Unlike skills that start with a brief or script, a demo starts literally with the app link—and it needs to be accessible.

The starting point
# everything starts with an accessible link
Public URL → https://app.exemplo.com
Localhost → http://localhost:8000/

# agent-browser opens in a fixed viewport
agent-browser set viewport 1280 800
agent-browser open http://localhost:8000/
✓ Input requirements
  • ✓ The app is live and responding
  • ✓ The target screen is reachable from the URL
  • ✓ Test credentials are available if login is required
  • ✓ Width fits in 16:9 (≤ ~1280)
✗ When input gets stuck
  • ✗ App offline — there’s nothing to navigate
  • ✗ Login without credentials (capture public screens only)
  • ✗ Screen depends on live data that disappears
  • ✗ URL inaccessible from the capturing machine
💡
Confirm the link before any script

The first practical step is to make sure the link opens and the target journey is navigable. Without that, the rest of the pipeline (capture, narration, render) has nothing to work with.

📋 Module 1.1 summary

What you learned
  • ✓ Walkthrough = real app navigated step by step, with narration
  • ✓ Demonstration SHOWS the app; explanatory EXPLAINS a concept
  • ✓ 1-question test: "is there a screen to show?"
  • ✓ Use cases: onboarding, feature, launch, support
  • ✓ Output: frame + cursor + zoom + narration
  • ✓ 16:9 output, narrated, with CTA; input is the app link
Next module
1.2
🧠 Capture first, animate later
The principle behind the skill: why rendering is deterministic, why you capture real screens first, and how the fixed viewport becomes the cursor’s coordinate space.
Go to module 1.2 →