PTENES
MODULE 2.4

🎞️ Scene composition

Adapt the composition-template.mjs for each new video: fill in AUDIO[], write sceneN(), code the GSAP tweens in anim() and align everything to a single source of truth for timing.

6
Topics
~35
Minutes
Medium
Level
Practical
Type
AUDIO[] single source · ffprobe LEAD=0.5 · TAIL=0.9 · FADE=0.45 data-start data-duration .scene .clip anim(i,t) GSAP tweens scene-inner sceneN() + BODIES HTML for Each Scene CAPTIONS[] caption per scene · cap-N index.html node build-index.mjs MP4 ✓ always synchronized AUDIO[] → single timing source → audio and animation always in sync
1

🧩 build-index.mjs — the project generator

Each video starts by copying scripts/composition-template.mjs such as build-index.mjs at the project root. This file is the generator: it runs once and writes index.html, and you render.

Main Concept

The template already provides premium dark CSS, a persistent background (glow/grid/grain), caption layer, progress bar, and overrides for 9:16. You only need to fill in 4 things: AUDIO[], CAPTIONS[], sceneN() e anim().

Run node build-index.mjs for 16:9 and node build-index.mjs --vertical for 9:16. Both write index.html — render immediately after each generation.

composition-template.mjs — header and imports
// scripts/composition-template.mjs (copy as build-index.mjs)
import { writeFileSync, readFileSync } from "node:fs";

// Local @font-face — paths relative to the project root
const FONT_CSS = readFileSync(
new URL("./assets/fonts/fonts.css", import.meta.url), "utf8")
.replace(/\.\/fonts\//g, "assets/fonts/");

// Format: 16:9 by default; pass --vertical for 9:16 (Shorts)
const VERT = process.argv.includes("--vertical");
const W = VERT ? 1080 : 1920;
const H = VERT ? 1920 : 1080;
const OUT = "index.html"; // always index.html
✓ Best practices for creating build-index.mjs
  • ✓ Copy the entire template — don’t create it from scratch
  • ✓ Keep scene 9 (CTA INEMA.CLUB) unchanged
  • ✓ Render immediately after each node build-index.mjs
  • ✓ Use writeFileSync for index.html — the template already does this
✗ Common pitfalls
  • ✗ Don't edit index.html directly — it will be overwritten
  • ✗ Don't use the Google Fonts CDN—it disappears in headless rendering
  • ✗ Don't remove CTA scene 9—it's the standard sign-off
  • ✗ Don't mix --vertical and 16:9 without rerunning the generator
📦 What the template includes for free
Premium dark CSS
Palette --bg:#0D1321, color variables, local fonts, overrides for 9:16.
Background layer
Radial glow, 64px grid, SVG grain—all with data-layout-ignore.
Progress + captions
Animated progress bar and caption layer on alternating track (2/4) — ready.
Key concepts
📄
build-index.mjs
Single generator
🔄
node → index.html
Each round
📐
--vertical flag
1080×1920 Shorts
🔒
Scene 9 untouched
INEMA.CLUB CTA
2

🔢 AUDIO[] — actual durations measured by ffprobe

The array AUDIO[] is the only timing data input for the entire project. Each element is the duration in seconds of the narration WAV for the corresponding scene, including CTA scene 9.

Why use ffprobe instead of estimating?

The duration of a WAV generated by Kokoro varies depending on the text, the speed (--speed 0.98) and the phonetics. If your estimate is wrong, the audio and visuals will be out of sync forever. Always measure.

The command ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/sN.wav prints the exact number in seconds — paste it directly into the array.

Real AUDIO[] + series S calculation (composition-template.mjs)
// ACTUAL durations measured with ffprobe (9 scenes; s9 = CTA)
const AUDIO = [
10.944, // s1 — intro
14.464, // s2 — folder + file
13.760, // s3 — SKILL.md anatomy
17.237333, // s4 — progressive disclosure
10.858667, // s5 — where they live
13.525333, // s6 — advanced level
11.114667, // real example
8.170667, // closing
3.840 // s9 — CTA inema.club
];

// Timing constants — DO NOT change
const LEAD = 0.5; // visual establishes before the voice
const TAIL = 0.9; // hold after the voice
const FADE = 0.45; // scene-inner fade in/out

// S series: accumulates absolute timing by scene
let t = 0;
const S = AUDIO.map((a, i) => {
const dur = LEAD + a + TAIL;
const o = { i: i + 1, start: round(t), dur: round(dur),
audioStart: round(t + LEAD), audioDur: round(a), end: round(t + dur) };
t += dur;
return o;
});
Example output: cumulative start times for the 9 scenes
1
s1: start=0 dur=12.344 audio@0.5 (10.944s)
2
s2: start=12.344 dur=15.864 audio@12.844 (14.464s)
...
s3–s8: accumulate sequentially — LEAD + audioDur + TAIL each
9
s9: start=~96.x dur=5.240 audio@~97.x (3.840s) — CTA
💡
Measure in batches with a shell script

Use for i in 1 2 3 4 5 6 7 8 9; do echo "s$i: $(ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/s$i.wav)"; done to generate the entire list at once and paste it into AUDIO[].

Key concepts
📏
ffprobe
Exact measurement
⏱️
LEAD=0.5
Visuals before voice
🔚
TAIL=0.9
Holds after voice
🎯
s9 = CTA
Always included
3

🎨 sceneN() — HTML for each scene

Each scene is a JavaScript function that returns an HTML string. The content is inserted into <div class="scene-inner">, inside a <section class="scene clip"> with timing attributes calculated from AUDIO[].

Structure generated by the template for each scene

The assembly transforms each item from S (array derived from AUDIO[]) in a <section> with three critical attributes: data-start, data-duration e data-track-index. Alternating tracks (1/3) avoid overlap at scene edges.

Assembly: .scene .clip with data-start/duration/track-index
// ---------- ASSEMBLY (composition-template.mjs) ----------
const scenesHTML = S.map((s, idx) => `
<section id="s${`s.i`}" class="scene clip"
data-start="${`s.start`}"
data-duration="${`s.dur`}"
data-track-index="${`s.i % 2 === 1 ? 1 : 3`}">
<div class="scene-inner" id="scene-inner-${`s.i`}">
${`BODIES[idx]()`}
</div>
</section>`
).join("");
Example: scene1() — opening scene of the Skills video
function scene1() {
return `
<div class="eyebrow" id="s1-eyebrow">
<span class="dot"></span>CLAUDE CODE · SKILLS
</div>
<h1 class="title">
<span class="word" id="s1-w1">Skills</span>
<span class="word accent" id="s1-w2">in Claude Code</span>
</h1>
<div class="rule" id="s1-rule"></div>
<p class="subhead" id="s1-sub">from the first principle to advanced</p>
<div class="reg tl" id="s1-r1"></div>
<div class="reg br" id="s1-r2"></div>
`;
}

// BODIES = ordered list of functions (scene9 = CTA, do not remove)
const BODIES = [scene1, scene2, scene3, scene4, scene5, scene6, scene7, scene8, scene9];
✓ sceneN() best practices
  • ✓ Reuse the template's CSS classes (.kicker, .h2, .grid2)
  • ✓ Give id unique for each animatable element (s1-word, etc.)
  • ✓ Off-canvas decorative elements: use data-layout-ignore
  • ✓ Animate the .scene-inner, never the .clip wrapper
✗ Errors that break the render
  • ✗ Duplicate IDs across scenes — GSAP will animate the wrong one
  • ✗ Animations in the .clip — the framework forces opacity:1 in the active clip
  • ✗ Reorder BODIES without adjusting the corresponding IDs
  • ✗ Remove or reposition scene9 (CTA fixed at the end)
Key concepts
🏷️
Unique IDs
Per scene + element
🔁
Tracks 1/3
Scene toggling
🎭
scene-inner
Animatable wrapper
📋
BODIES[]
Scene order
4

✨ anim(i,t) — GSAP tweens per scene

The function anim(i, t) receives the scene index and absolute start time. It pushes GSAP code strings into a local array, which is later concatenated into <script> in the generated HTML.

Why generate code as a string and not execute it directly?

The generator runs in Node.js, but GSAP executes in the browser (headless Chrome). anim() builds code strings to be embedded in the HTML — when the browser loads it, these strings already have the absolute times calculated from AUDIO[].

The helper at(d) (shortcut to round(t + d)) calculates the offset relative to the start of the scene. Use it in all tweens to keep the code readable.

anim() — scene-inner entrance/exit pattern + case 1
// composition-template.mjs — anim() function
function anim(i, t) {
const L = [];
const P = (s) => L.push(s);
const at = (d) => round(t + d);

// inner entry/exit (common to all scenes)
P(`tl.fromTo("#scene-inner-${`i`}",
{opacity:0},{opacity:1,duration:${`FADE`},ease:"power2.out"},${`t`});`);
P(`tl.to("#scene-inner-${`i`}",
{opacity:0,duration:${`FADE`},ease:"power2.in"},${`round(S[i-1].end - FADE)`});`);
P(`tl.set("#scene-inner-${`i`}",{opacity:0},${`round(S[i-1].end)`});`);

// case 1: opening scene
switch (i) {
case 1:
P(`tl.from("#s1-eyebrow",{y:-24,opacity:0,duration:.55,ease:"power3.out"},${`at(0.15)`});`);
P(`tl.from("#s1-w1",{y:70,opacity:0,duration:.7,ease:"power4.out"},${`at(0.35)`});`);
P(`tl.from("#s1-w2",{y:70,opacity:0,duration:.7,ease:"power4.out"},${`at(0.55)`});`);
P(`tl.fromTo("#s1-rule",{scaleX:0},{scaleX:1,duration:.7,ease:"expo.out",
transformOrigin:"left center"},${`at(0.95)`});`);
P(`tl.from("#s1-sub",{y:20,opacity:0,duration:.6,ease:"power2.out"},${`at(1.15)`});`);
P(`tl.fromTo("#s1-cur",{opacity:1},{opacity:0,duration:.5,repeat:18,
yoyo:true,ease:"none"},${`at(1.6)`});`);
break;
// ... cases 2-9 ...
}
// caption fade in/out
P(`tl.fromTo("#cap-${`i`}",{opacity:0,y:14},{opacity:1,y:0,duration:.5,ease:"power2.out"},${`at(0.35)`});`);
P(`tl.to("#cap-${`i`}",{opacity:0,duration:.4,ease:"power2.in"},${`round(S[i-1].end - 0.55)`});`);
return L.join("\n ");
}
💡
Animation golden rules

1. Always animate the .scene-inner, never the .clip wrapper. 2. Use absolute times calculated by at(d) — never hardcode numbers. 3. The FADE=0.45 is the same for input and output across all scenes, ensuring a uniform transition.

⚠️
Scenes on alternating tracks — not optional

Odd-numbered scenes stay on data-track-index="1" and pairs in data-track-index="3". Captions follow the same pattern (2/4). This prevents the “edge overlap” where two frames from adjacent scenes appear simultaneously—a classic error described in references/gotchas.md.

Key concepts
✨
fromTo / from
GSAP tweens
🕐
at(d) helper
Relative offset
🔀
FADE=0.45
Uniform transition
🧱
switch(i)
Tweens per scene
5

💬 CAPTIONS[] — one short caption per scene

The array CAPTIONS[] has exactly the same length as AUDIO[]. Each string is displayed at the bottom of its corresponding scene as an accessible caption, synchronized via data-start/duration inherited from the same S[idx].

Captions as cognitive reinforcement, not as a transcript

A good caption captures the central idea for the scene in 5–10 words — it does not transcribe the audio. Viewers watching without sound should understand the point of each scene from the caption + visual alone.

The caption runs in data-track-index="2" (odd-numbered scenes) or "4" (pairs) — alternating tracks to avoid overlap with the previous scene at the transition boundary.

Real CAPTIONS[] + caption HTML assembly
// composition-template.mjs — CAPTIONS and assembly
const CAPTIONS = [
"Skills in Claude Code — from beginner to advanced",
"One Skill = one folder + one SKILL.md file",
"name + description — the description is the trigger",
"Progressive disclosure: loads only when needed",
"Where they live: .claude/skills (project or global)",
"Advanced: scripts, references, and templates",
"This video was made with the HyperFrames Skill",
"Start with a SKILL.md. Now it’s up to you.",
"More content at inema.club", // s9 = CTA
];

// automatic assembly — same timing as S
const captionsHTML = S.map((s, idx) => `
<div class="caption clip" id="cap-${`s.i`}"
data-start="${`s.start`}"
data-duration="${`s.dur`}"
data-track-index="${`s.i % 2 === 1 ? 2 : 4`}">
${`CAPTIONS[idx]`}
</div>`
).join("");
✓ Good captions
  • ✓ Short sentence that summarizes the central concept
  • ✓ Includes exact technical terms (file names, commands)
  • ✓ Readable for viewers watching without audio
  • ✓ The s9 caption mentions the site: inema.club
✗ Poor captions
  • ✗ Verbatim audio transcription (too long)
  • ✗ Vague: "an important idea about the topic"
  • ✗ Full captions — makes quick reading harder
  • ✗ Different AUDIO[] length (will cause an error)
💡
Caption as a video indexer

On YouTube, captions are indexed by the search engine. Accurate captions with correct technical terms (ffprobe, GSAP, build-index.mjs) improve the video's ranking in technical searches.

Key concepts
💬
1 caption/scene
Same length as AUDIO[]
🔁
Tracks 2/4
Caption toggling
📐
Automatic sync
Same data-start as the scene
🔍
SEO-friendly
Exact technical terms
6

⏱️ Single-source timing—audio and animation always in sync

The most important principle of composition-template: AUDIO[] generates everything — data-start, data-duration, the GSAP tween timings e the attributes of the <audio>. No number is duplicated or hardcoded elsewhere.

Single source of truth — formal definition

The three constants LEAD=0.5, TAIL=0.9 e FADE=0.45 combined with the values in AUDIO[] determine each scene’s start time, duration, and fade — and therefore the exact timing of all GSAP tweens. Changing a value in AUDIO[] automatically propagates to everything.

Formulas: dur = LEAD + audioDur + TAIL · audioStart = start + LEAD · Entrance tween = t · Exit tween = end - FADE.

Complete workflow: AUDIO[] → S → HTML attributes + GSAP + <audio>
// EVERYTHING derives from AUDIO[] + LEAD/TAIL/FADE — never duplicate
// 1. Series S: each object carries all the required timings
const S = AUDIO.map((a, i) => {
const dur = LEAD + a + TAIL; // total scene duration
return {
i: i + 1,
start: round(t), // → data-start
dur: round(dur), // → data-duration
audioStart: round(t + LEAD), // → audio data-start
audioDur: round(a), // → audio data-duration
end: round(t + dur), // → exit tween
};
});

// 2. Audio — data-start = audioStart (starts AFTER the visual LEAD)
const audioHTML = S.map((s) => `
<audio id="a${`s.i`}" data-start="${`s.audioStart`}"
data-duration="${`s.audioDur`}" data-track-index="20"
src="assets/audio/s${`s.i`}.wav"></audio>`
).join("");

// 3. Tweens generated via anim() — use S[i-1].start and S[i-1].end
const animJS = S.map((s) => anim(s.i, s.start)).join("\n ");
📐 Timing formulas — quick reference card
Duration of each scene
dur = LEAD + audioDur + TAIL
= 0.5 + medido + 0.9
Audio start (delays the LEAD)
audioStart = start + LEAD
The visuals come before the voice
Exit tween for scene-inner
exitAt = S[i-1].end - FADE
= end - 0.45
Total composition duration
TOTAL = sum(all dur)
The progress bar uses this value
✓ What the single source guarantees
  • ✓ Change a WAV → run the generator → everything syncs
  • ✓ Add a scene → add an entry in AUDIO[], CAPTIONS[], sceneN(), anim() case N
  • ✓ No hardcoded numbers outside AUDIO[]/LEAD/TAIL/FADE
  • ✓ npx hyperframes lint detects track desynchronization
✗ Common violations
  • ✗ Hardcode time in the tween: gsap.to(..., 14.5)
  • ✗ Adjust data-start in the generated HTML — it will be overwritten
  • ✗ AUDIO[] and CAPTIONS[] arrays of different lengths
  • ✗ Use round() inconsistent → frame offset
Key concepts
🎯
Single AUDIO[]
Everything derives from it
🔗
Series S
Timing by object
⚡
round(n)
ms precision
🔄
Propagates everything
Change AUDIO[], run

📋 Module 2.4 Summary

What you learned
  • ✓ Copy the template as build-index.mjs and run with node
  • ✓ Fill in AUDIO[] with real durations from ffprobe (includes CTA s9)
  • ✓ Writing sceneN() reusing the template’s CSS classes
  • ✓ Code GSAP tweens in anim(i,t) always animating the .scene-inner
  • ✓ Fill in CAPTIONS[] with short, technical phrases
  • ✓ Understand the timing formulas: LEAD=0.5, TAIL=0.9, FADE=0.45
Next module
2.5
✅ Validate & render
Run npx hyperframes lint for 0 errors, inspect --samples 16 for 0 layout issues, and generate the final MP4 at --quality high for 16:9 and 9:16.
Go to module 2.5 →