🎬 What is a walkthrough
A walkthrough — or demo video — goes through a real web application, screen by screen, showing the app in use: opening, clicking, filling in fields, generating, seeing the result. Nothing is drawn; everything is captured from the actual screen.
A walkthrough answers "how do I use this?" showing the response in the interface itself. Instead of describing buttons and fields, it appears to interact with them—with a cursor that lands exactly on the control and narration explaining each step.
In practice: 5 to 8 steps plus the final CTA make a video about 35 to 50 seconds long. The typical arc is abrir o app → ação 1 → ação 2 → … → resultado → CTA.
The opening step is usually the home screen (marked intro:true) to present the app. The last content step is usually the result (zoom:true), with a gentle zoom that highlights the result.
⚖️ Demonstrative vs. explanatory
These are two sibling skills with the same rendering engine (HyperFrames) and the same CTA, but opposite purposes. One showcases a real app; the other explains an abstract concept with animated scenes.
| Aspect | 🖱️ video-demonstrativo | 🎬 video-explicativo |
|---|---|---|
| Objective | Show an app in use | Explain a concept |
| Visual | Real screenshots | Animated motion graphics |
| Input | App link / localhost | One subject / topic |
| Capture | agent-browser navigates the app | No capture |
| Format | 16:9 (landscape screens) | 16:9 e 9:16 |
- ✓ There’s an app/site with a navigable screen
- ✓ Did the user provide a link or localhost
- ✓ Want to show how a feature works in practice
- ✓ The focus is the product, not the idea behind it
- → The topic is abstract, with no interface to show
- → You want to illustrate with diagrams and animations
- → Need a vertical version (Shorts/Reels)
- → The focus is educating viewers about a concept
Both end with the INEMA.CLUB CTA and use local TTS + HyperFrames. The difference is where the visuals come from: captured screen (demonstrative) versus drawn scene (explanatory).
🤔 When to use each one
There’s a one-question test that determines the skill in seconds — and clears up any doubt before you even start the script.
"Is there a screen to show?"
If the answer is yes—there’s an app, a localhost, or a navigable URL—it’s demonstrative. If there's no interface, just an idea to illustrate, it's explanatory.
If so, that’s the strongest sign it’s a demo. The skill’s own entry is the app link.
Demo verbs (“show,” “record the screen,” “step by step using”) point to the demo flow.
That’s explanatory—the content will be motion graphics, not a screen recording.
A single question to the user resolves ambiguous cases at no cost.
Demonstration requires the app to be accessible during capture (e.g., localhost:8000 running). If the app isn't live, there's nothing to navigate — confirm this before promising the video.
📋 Use cases
A walkthrough shines whenever the goal is to show a product in action. Four situations cover most requests.
A new user sees the app's first steps: sign-up, main screen, first valuable action. It reduces friction for someone who's just arrived.
Demonstrates a specific feature end to end — useful when a new feature needs a quick “how it works.”
An announcement video that shows the real product in action at launch, with the INEMA.CLUB CTA at the end.
Answers “how do I do X?” with a short video showing the exact path on screen—ready to attach to a support ticket or help center.
Knowing whether it’s onboarding, a feature, a launch, or support helps you choose which steps go into the video—and which “result” deserves the final zoom.
🖱️ What the skill provides
Four elements turn static screenshots into a video that looks like a professional screen recording: a browser frame, an animated cursor, highlights/zoom, and narration.
Real screenshots go inside a browser frame (bar + URL) to look like a real window. On top, a global animated cursor aims at the real bounding box of each control, clicks with a pulse + ripple, and the result gets a gentle zoom. TTS narration ties it all together.
- ✓ Cursor lands exactly on the control (real box, not by eye)
- ✓ Click becomes a visible pulse + ripple
- ✓ A smooth zoom highlights the final result
- ✓ Timing syncs with WAVs measured by ffprobe
- ✗ Dynamic state (video, live data) becomes a static screenshot
- ✗ Doesn't load the live site inside the video
- ✗ App with login requires test credentials
- ✗ Real movement only in v3 mode (roadmap, not implemented yet)
📐 Output format
The output is an MP4 in 16:9, narrated in PT-BR and ending with the INEMA.CLUB CTA. The landscape format isn’t aesthetic—it’s technical: app screens are wide.
Every output ends on the same scene: "CONTINUE AT" + INEMA.CLUB (INEMA in cream, .CLUB in amber with glow) + 🌐 inema.club. The narration ends with "This is content from INEMA dot CLUB. Visit: inema dot club." It’s already included in the composition template.
Because app screens are landscape, the template generates 16:9. A vertical version would require cropping/reframing the active region — it's on the roadmap, not part of the standard workflow. Don't promise Shorts for a walkthrough without aligning on the reframing.
🔗 The input is the link
Unlike skills that start with a brief or script, a demo starts literally with the app link—and it needs to be accessible.
- ✓ The app is live and responding
- ✓ The target screen is reachable from the URL
- ✓ Test credentials are available if login is required
- ✓ Width fits in 16:9 (≤ ~1280)
- ✗ App offline — there’s nothing to navigate
- ✗ Login without credentials (capture public screens only)
- ✗ Screen depends on live data that disappears
- ✗ URL inaccessible from the capturing machine
The first practical step is to make sure the link opens and the target journey is navigable. Without that, the rest of the pipeline (capture, narration, render) has nothing to work with.
📋 Module 1.1 summary
- ✓ Walkthrough = real app navigated step by step, with narration
- ✓ Demonstration SHOWS the app; explanatory EXPLAINS a concept
- ✓ 1-question test: "is there a screen to show?"
- ✓ Use cases: onboarding, feature, launch, support
- ✓ Output: frame + cursor + zoom + narration
- ✓ 16:9 output, narrated, with CTA; input is the app link