PTENES
MODULE 1.2

🔄 What Changes in the Way You Edit

You stop fiddling with tools and start say what you want. These are six principles that replace "dragging a clip onto the timeline" with "describing the result and letting the engine execute." Each defines, in plain language, a rule that recurs throughout the course.

6
Topics
~35min
Duration
Basic
Level
Theory
Type
Sections read in this module0% · 0 of 0
1

📝 Specification First, Execution Later

The number one mindset shift: you says what you want first — “agile ad reel with my cold open” — and the skill decides how get there (which cuts, which animations, which B-roll). In a traditional editor, you’re stuck with the "how" all the time: dragging clips, adjusting the frame. Here, the "how" is the engine’s problem.

New here? — “specification” and “the HOW”

A specification is the description of the result you want, in words—not clicks. The "how" (the execution) are the technical steps to get to this result: cut here, animate there, add this caption. The shift is: you take care of the what; the skill handles the how.

🎯 The golden rule

Describe the destination, not the path. When you specify it well ("I want a serious news-style pace, no music"), the engine has the freedom to put things together better than if you micromanaged every cut.

What it is

Replace “operate the tool” with “describe the result” — intent before execution.

Why learn

You think like a director, not an operator. Less time on technical details, more time on intent.

Key concepts

Specification, intention, execution, the what vs. the how.

2

⏱️ Measured pacing: no more than 4s without a visual hit

A hard rule for the engine: the screen never stays still for more than 4 seconds. Every window of up to 4s has to include a visual hit — something that refreshes the image and keeps the viewer’s eye. It’s not decoration: it’s what keeps the person from swiping away.

New here? — what counts as a “visual hit”

A visual hit is any change that refreshes the screen: a edit (shot change), a zoom (zoom in/out), a chip (an incoming tag/label), a B-roll (overlay image) or a reveal (something that appears with animation). Any of these resets the 4-second timer.

0s 4s 8s 12s 16s 20s edit zoom chip B-roll reveal Each segment ≤ 4s → a new hit. The screen never "rests."
Green = engine moves (cuts, zooms, reveals); cyan = content moves (chip, B-roll). No more than 4 seconds ever pass between two markers.

💡 Reading tip

When reviewing one of your reels, count “one thousand one, one thousand two…” If more than four seconds pass without anything changing on screen, that’s where the audience drops off. The engine treats that gap as error, not as a style.

Key concepts: visual hit, 4s window, gap (the time spent idle that becomes an error), retention.

3

🎥 Real B-roll over generic footage

When you mention a repository, website, or product, the engine doesn’t throw a generic stock image on top. It fetches the real capture of that — the page itself, the repo itself — and places it inside a browser frame. That changes how people perceive it: instead of “a nice little video,” it becomes “a video that shows the real thing.”

✓ Real B-roll (search)

  • ✓Capture of the page/repo you mentioned
  • ✓Inside a browser frame, for credibility
  • ✓Visual proof: "this is exactly what I’m talking about"

✗ Generic B-roll (avoid)

  • ✗Stock photo of “code on screen”
  • ✗Illustration that isn’t the real product
  • ✗Template look: fills space, proves nothing

🔍 Why it matters

Real B-roll is one of the things that most clearly separates a reel "edited by a person" from one "assembled from a template." Showing the actual source builds authority and is hard to fake—so the engine prioritizes a capture before any generic image.

Key concepts: Real B-roll, capture, browser frame, visual proof.

4

🔁 Always Use the Last Take

This frees you up while recording: got a line wrong? Repeat and continue. When the same line appears several times in a row, the engine understands that the earlier ones were attempts and keeps the last good one. You don't need to stop, cut, or re-record from scratch — record it "straight through" and the skill cleans it up.

How the engine reads three takes of the same line

take 1 "Forja Reel, it…" (stuck) discards ✗
take 2 "Forja Reel tur— oh" (wrong) discards ✗
take 3 "Forja Reel turns your raw footage into a reel." (good) keeps ✓

💡 How to make the most of it while recording

Don’t stop to “re-record it properly.” Start over and same phrase right after it, in sequence. Saying the correct version last is the signal the engine uses to choose — the cleaner the last one, the better the cut.

Key concepts: take, repetition, “last good one,” record straight through.

5

🎛️ Adapt to the video, don’t follow a fixed template

A fixed template treats everything the same—and that’s exactly why it “smells like AI.” The engine does the opposite: it watches the video and proposes a treatment for that content. A news story calls for seriousness and an urgent pace; an ad calls for energy and a sales-driven cold open. Same process, different results.

📰 News

  • •Serious tone, no flashy music
  • •Urgent pacing, straight to the point
  • •B-roll from the original news source

📣 Ad

  • •High energy, salesy cold open
  • •Punchy pacing, more prominent motion
  • •Focus on benefits and calls to action

🧭 The core idea

There is no “Forja Reel template.” There is an engine that analyzes before assembling and chooses the right treatment for the content at hand. That’s what makes each reel feel thoughtfully made, not stamped out.

Key concepts: treatment, fixed template, content analysis, adaptation.

6

📐 Keep the Resolution and Put Text at Mouth Level

Two finishing touches we tend to miss when we’re in a rush. First: the engine keeps the resolution that came in — if you recorded in 2K, the output is 2K; if you recorded in 4K, the output is 4K, without squeezing the quality. Second: the caption sits at mic height, not tucked underneath where the app bar covers it.

New here? — “resolution,” “2K/4K,” and “mic height”

Resolution is the number of pixels in the image—the higher it is, the sharper the image. 2K/4K are two common tiers (4K has about four times as many pixels as Full HD). "Mic height" is the vertical band where the caption should sit: above the Instagram/TikTok button bar, roughly at the height of the person’s mouth as they speak—so no one has to crane their neck or miss text hidden under the interface.

✓ What the engine guarantees

  • ✓Input 2K, output 2K · input 4K, output 4K
  • ✓Caption in the safe zone, above the app bar
  • ✓Highlighted caption keyword

✗ The errors it avoids

  • ✗Export at a smaller size and lose sharpness
  • ✗Caption stuck at the bottom, covered by the interface
  • ✗Text that's too small to read on a phone

✅ Closes Track 1

Finishing touches are what separate “almost pro” from “pro.” Keeping the resolution and putting the caption in the right place are easy to get right and costly to get wrong—the engine handles both by default so you don’t miss them when you’re in a hurry.

Key concepts: resolution, 2K/4K, safe zone, mic height.

✅ Module summary

✓
Specifies, doesn’t operate — you say what; the skill decides how.
✓
No more than 4s without movement — cuts, zooms, chips, B-roll, or reveals refresh the screen.
✓
Real B-roll and final shot — real capture instead of generic footage; record straight through, and the skill keeps the good takes.
✓
Adapts and ends well — content-aware treatment, resolution preserved, and captions at mic height.

Next module:

This is the last module in Track 1. You already master the why e o what — now move on to the Track 2 · Inside the skill, where we open up the engine and see the 5 phases in action.