PTENES
MODULE 2.1

⚙️ The engine’s 5 phases

Every child skill edits by following the same five-phase process: cuts, reads and plans, builds motion and B-roll, applies SFX and renders with QC, and finally goes through a reviewer in a loop. Each phase has its quality gate—that's what makes the result reliably good, not a matter of luck.

7
Topics
~45min
Duration
Intermediate
Level
Theory
Type
Sections read in this module0% · 0 of 0
1

⚙️ Phase 1 — Deterministic cutting

The first phase is the most mechanical — and it’s the foundation for everything. The engine doesn’t “guess” where to cut: it measures the audio energy and finds the sections where you actually speak (the voice islands). Then it removes silences, long pauses, “let me repeat that,” and false starts, always keeping the last shot of each sentence — the good one.

New here? — “deterministic” and “voice island”

Deterministic means the same input produces exactly the same output every time—without randomness. Voice island is each continuous segment containing speech, separated from neighboring segments by silence. The engine detects these islands by energy (silencedetect), not by transcription.

PHASE 1 Cut silencedetect PHASE 1.5 Read + plan /watch · beat sheet PHASES 2/3 Motion + B-roll cold open · ≤4s PHASE 4 SFX + render Automatic QC PHASE 5 Reviewer PASS / FAIL FAIL → fix and repeat until PASS
The map of the whole workflow: phases 1 through 4 move forward in a line (blue); phase 5 closes the loop with the cyan arrow back. Keep this diagram handy—each topic in this module is one of these boxes.
What it is

Automatic cut based on audio energy, keeping the last take of each sentence.

Why learn

A predictable, clean cut is the foundation for every subsequent phase.

Key concepts

silencedetect, island, last take, determinism.

2

👁️ Phase 1.5 — Watch the Video and Propose a Treatment

Between the cut and the edit, there’s a “middle” phase that makes all the difference: the engine watches its own cut. Using /watch, it looks at the frames, understands what the video is about, and suggests a treatment: which opening to use, what kind of B-roll fits, where to speed things up and where to give it room. The result is a beat sheet — the editing plan.

New here? — “/watch” and “beat sheet”

/watch is a visual QC capability: the agent sees the video frames, not just the text. Beat sheet is the map of the reel’s “beats” (rhythmic hits): where the cold open comes in, where B-roll goes, where things speed up. It’s the edit plan before you start editing.

🎯 Why adapt instead of applying a template

A news story calls for urgency and data; a tutorial calls for clarity and steps; an ad calls for a hook and proof. Analyzing the video before editing lets the engine choose the right treatment for THAT content—instead of stamping the same template on everything, which would cause visual fatigue and an “AI look.”

Key concepts: /watch, treatment, beat sheet, opening, adaptation.

3

🎞️ Phases 2/3 — Motion graphics + real B-roll

With the clean cut and plan in hand, the engine assembles: the Hyperframes generates the animations (motion graphics) and the Firecrawl brings real screenshots of websites and repositories for the B-roll. Two hard rules apply here: mandatory cold open (impact in the first few seconds) and never more than 4s without a visual punch.

✓ What these phases deliver

  • ✓A hook-setting cold open in the first 1–3s
  • ✓Motion graphics with a deterministic timeline
  • ✓Real B-roll (capture of the cited site/repo)

✗ What the gate prevents

  • ✗More than 4s of a static screen with only speech
  • ✗Open without a hook (no cold open)
  • ✗Generic stock image instead of the real thing

💡 Reading tip

Hyperframes and Firecrawl come up again, with real commands, in module 2.2. Here, just remember each one’s role: Hyperframes = animation; Firecrawl = bringing in the real world from outside.

Key concepts: Hyperframes, Firecrawl, cold open, 4-second rule.

4

🔊 Phase 4 — SFX, Rendering, and QC

Now the sound goes in and the file comes out. The engine overlays Subtle SFX (whoosh on transitions, ching on numbers) under the voice — and no music by default, so it doesn’t compete with the speech. Then render and run a Automatic QC that looks for gaps and overlaps; when it finds a problem, it regenerates that section.

♪

SFX under the voice

Short, subtle effects that “produce” the cut without becoming noise pollution. Music is excluded by default.

▶

Render

The layers (video, motion, captions, SFX) are “flattened” into a single MP4—preserving the input resolution.

✓

Automatic QC

A scan looks for pacing gaps and overlapping sounds; anything that fails is regenerated, not delivered.

Key concepts: SFX, render, QC (quality control), regeneration.

5

🔁 Phase 5 — Independent Reviewer in a Loop

The final phase is a safety lock: a independent reviewer — a subagent with clean context — watches the render (via /watch) and issues a verdict: PASS or FAIL. If it's FAIL, it says what's wrong, the engine fixes it, and the cycle repeats. Nothing is delivered before a PASS.

Render Clean reviewer /watch → PASS/FAIL Delivery only after PASS PASS FAIL → fix and go back
The blue path (PASS) opens only when the reviewer approves. While it fails, the cyan arrow (FAIL) sends the render back for fixes. The loop repeats until it passes.

Key concepts: reviewer, subagent, PASS/FAIL, quality loop.

6

🧪 Why a Separate Reviewer Works

Why not let the agent that assembled the video judge the result? Because the builder gets contaminated because it made the edit: it knows the intent, mentally "fills in" what’s missing, and doesn’t see its own mistakes. The reviewer is a subagent who has never seen the edit — it judges only what’s on screen, as a viewer would.

🧠 The core idea

A fresh perspective catches what an invested perspective misses. It’s the same reason we’re better at reviewing someone else’s writing than our own: distance creates objectivity. The engine makes that distance part of the process by giving the reviewer clean context.

Agent that assembled it

Knows the intent behind every cut. Tends to approve because they “know what they meant”—even if the screen doesn’t show it.

Independent reviewer

It doesn’t have that kind of memory. If something isn’t clear just by looking, it fails—which is exactly what a viewer would experience.

Key concepts: clean context, author bias, independent verification.

7

🗝️ Key Engine Concepts

Three ideas tie together everything you've seen and return in module 2.2 (where they become scripts). Lock them in now: they explain why the engine is reliable, not just clever.

Hard gate
A blocking check: if it doesn’t pass, the process stops and nothing moves forward. Examples: verify-cut.py (clean cut) and lint-timeline.py (pacing ≤4s).
Beat sheet
The reel’s pacing plan: where the cold open comes in, where B-roll goes, where the pace picks up. It comes out of Phase 1.5 and guides the edit.
Quality loop
Repeat review→fix until a PASS. Turns “does it look good?” into an objective stopping condition.

Self-recovery (optional): why the Phase 5 reviewer is a subagent separate?

✅ Module summary

✓
Phase 1 cuts by energy — silencedetect + last take, deterministically.
✓
Phase 1.5 reads and plans — /watch proposes the treatment and builds the beat sheet.
✓
Phases 2/3 and 4 assemble and finalize — motion + real B-roll, cold open, ≤4s, SFX, render, and QC.
✓
Phase 5 closes the loop — an independent reviewer approves (PASS) or returns it (FAIL) until it’s good.

Next module:

2.2 — Gates, scripts, and determinism (the actual commands behind each phase)