PTENES
MODULE 3.3

⚙️ The build-index.mjs generator

Complete anatomy of the composition-template.mjs: how AUDIO[] controls everything, the calculation of S[], GSAP scene and tween construction, complete HTML assembly, 9:16 overrides, and the invariant CTA.

6
Topics
~40
Minutes
Advanced
Level
Technical
Type
AUDIO[] single source of timing S[] start · dur · audioStart · end LEAD=0.5 · TAIL=0.9 · FADE=0.45 scenesHTML .scene .clip · sceneN() captionsHTML tracks 2 / 4 audioHTML track 20 · s.audioStart animJS anim(i, t) · GSAP tweens index.html writeFileSync · always synchronous AUDIO[] → S[] → 4 streams → index.html
1

🔢 AUDIO[] as the single source of timing

The array AUDIO[] contains the REAL durations (measured with ffprobe) for each WAV narration. It's the only variable you need to fill in so that all video timing—HTML, GSAP, and audio—stays automatically synchronized.

Main Concept

The duration array controls everything. No time values appear elsewhere in the code. Change a WAV file, update the corresponding entry in AUDIO[], run the generator — the entire video recalibrates automatically.

This is possible because the generator runs in Node.js and produces a static HTML file with all timings already calculated. The browser receives ready-to-use numbers — it doesn’t need to calculate anything at runtime.

composition-template.mjs — AUDIO[] with actual durations measured by ffprobe
// ACTUAL durations measured with ffprobe (9 scenes; s9 = CTA — do not remove)
const AUDIO = [
10.944, // s1 — intro
14.464, // s2 — folder + file
13.760, // s3 — SKILL.md anatomy
17.237333, // s4 — progressive disclosure
10.858667, // s5 — where they live
13.525333, // s6 — advanced level
11.114667, // real example
8.170667, // closing
3.840 // s9 — CTA inema.club (invariant)
];

// Measure all at once with a shell script:
// for i in 1 2 3 4 5 6 7 8 9; do
// echo "s$i: $(ffprobe -v error -show_entries format=duration \
// -of default=noprint_wrappers=1:nokey=1 assets/audio/s$i.wav)"
// done
✓ What AUDIO[] controls
  • ✓ data-start e data-duration of each .scene.clip
  • ✓ Absolute times for all GSAP tweens via anim(i, s.start)
  • ✓ data-start of the elements <audio> in track 20
  • ✓ TOTAL of the composition and progress bar
✗ Why you should never estimate
  • ✗ Kokoro generates different durations depending on the text and speed (--speed 0.98)
  • ✗ Wrong estimate = permanently out-of-sync audio and visuals
  • ✗ Headless rendering doesn’t alert you—the video simply turns out wrong
  • ✗ Never guess durations — always use ffprobe
Key concepts
🔢
AUDIO[]
Only timing input
📏
ffprobe
Required measurement
🔄
Propagates everything
HTML + GSAP + audio
🔒
s9 = CTA
Always included
2

⏱️ The calculation of S[]—the timing series

The series S is generated by a single AUDIO.map(). Each object carries all the timing needed for scenes, captions, audio, and tweens—calculated once and propagated everywhere.

The three constants and the accumulator

Three constants control the rhythm: LEAD=0.5 (visual comes in before the voice), TAIL=0.9 (visual holds after the voice) and FADE=0.45 (fade-in/out duration of the .scene-inner). An accumulator t grows with each scene.

Fundamental formulas: dur = LEAD + a + TAIL · audioStart = t + LEAD · end = t + dur.

composition-template.mjs — full S[] calculation with AUDIO.map()
// Timing constants — never change
const LEAD = 0.5; // visual before the voice
const TAIL = 0.9; // hold after the voice
const FADE = 0.45; // scene-inner fade in/out

// rounding helper (3 decimal places) — avoids frame offsets
const round = (n) => Math.round(n * 1000) / 1000;

// absolute time accumulator
let t = 0;
const S = AUDIO.map((a, i) => {
const dur = LEAD + a + TAIL; // total scene duration
const o = {
i: i + 1, // 1-based index
start: round(t), // → data-start of the .clip
dur: round(dur), // → data-duration of the .clip
audioStart: round(t + LEAD), // → data-start of the <audio>
audioDur: round(a), // → data-duration of the <audio>
end: round(t + dur), // → exit tween (end - FADE)
};
t += dur; // advances the accumulator
return o;
});

// Total composition duration — feeds the progress bar
const TOTAL = round(t);
Actual output: S[] objects from the first scenes
1
s1: start=0 dur=12.344 audioStart=0.5 audioDur=10.944 end=12.344
2
s2: start=12.344 dur=15.864 audioStart=12.844 audioDur=14.464 end=28.208
…
s3–s8: accumulate sequentially. Each start = the previous scene’s end.
9
s9: start=~96.x dur=5.240 audioStart=~97.x audioDur=3.840 end=~101.x — CTA
💡
The generator prints the calculated values

After generating the HTML, the template displays this in the console: OUT gerado · W×H · TOTAL = Xs · N cenas, followed by one line per scene with start, dur, audio@, and audioDur. Check these numbers before rendering.

Key concepts
⏱️
LEAD=0.5
Visuals before voice
🔚
TAIL=0.9
Holds after voice
🌊
FADE=0.45
scene-inner transition
➕
round(n)
3 decimal places
3

🎬 sceneN() and anim(i,t) per scene

Each scene has two functions: sceneN() returns the inner HTML of .scene-inner, e anim(i,t) generates GSAP code strings that will be embedded in the <script> in the final HTML. The fade on the .scene-inner is always generated by the shared part of anim().

Why generate tweens as code strings?

The generator runs in Node.js, but GSAP executes in the browser (headless Chrome). anim() builds JavaScript strings with the already calculated absolute times — when the browser loads the HTML, these strings run with exact numbers, without recalculation. The fade of the .scene-inner is shared by all scenes; the internal switch adds each scene's specific tweens.

sceneN() — actual example from composition-template.mjs (scene 1)
// Each sceneN() returns plain HTML — no timing, no GSAP
function scene1() {
return `
<div class="eyebrow" id="s1-eyebrow">
<span class="dot"></span>CLAUDE CODE · SKILLS
</div>
<h1 class="title">
<span class="word" id="s1-w1">Skills</span>
<span class="word accent" id="s1-w2">in Claude Code</span>
</h1>
<div class="rule" id="s1-rule"></div>
<p class="subhead" id="s1-sub">from the first principle to advanced</p>
`;
}

// BODIES = ordered list of functions (scene9 = CTA, do not remove)
const BODIES = [scene1, scene2, scene3, scene4, scene5, scene6, scene7, scene8, scene9];
anim(i,t) — scene-inner fade + case-specific tweens
// anim() generates GSAP strings embedded in the generated HTML's <script>
function anim(i, t) {
const L = []; // string buffer
const P = (s) => L.push(s); // push helper
const at = (d) => round(t + d); // relative offset

// --- .scene-inner fade in/out (common to ALL scenes) ---
P(`tl.fromTo("#scene-inner-${i}",
{opacity:0},{opacity:1,duration:${FADE},ease:"power2.out"},${t});`);
P(`tl.to("#scene-inner-${i}",
{opacity:0,duration:${FADE},ease:"power2.in"},${round(S[i-1].end - FADE)});`);
P(`tl.set("#scene-inner-${i}",{opacity:0},${round(S[i-1].end)});`);

// --- scene-specific tweens ---
switch (i) {
case 1:
P(`tl.from("#s1-eyebrow",{y:-24,opacity:0,duration:.55,ease:"power3.out"},${at(0.15)});`);
P(`tl.from("#s1-w1",{y:70,opacity:0,duration:.7,ease:"power4.out"},${at(0.35)});`);
P(`tl.from("#s1-w2",{y:70,opacity:0,duration:.7,ease:"power4.out"},${at(0.55)});`);
P(`tl.fromTo("#s1-rule",{scaleX:0},{scaleX:1,duration:.7,ease:"expo.out",
transformOrigin:"left center"},${at(0.95)});`);
P(`tl.from("#s1-sub",{y:20,opacity:0,duration:.6,ease:"power2.out"},${at(1.15)});`);
break;
// ... cases 2–8 with scene-specific tweens ...
// case 9 = CTA — do not modify
}
// caption fade in/out
P(`tl.fromTo("#cap-${i}",{opacity:0,y:14},{opacity:1,y:0,duration:.5,ease:"power2.out"},${at(0.35)});`);
P(`tl.to("#cap-${i}",{opacity:0,duration:.4,ease:"power2.in"},${round(S[i-1].end - 0.55)});`);
return L.join("\n ");
}
⚠️
Always animate .scene-inner — never .clip

HyperFrames forces opacity:1 in the active clip. If you try to animate the .clip or the .scene, the fade animation doesn't work. The animatable wrapper is always the child .scene-inner. The template already generates the fades for it—don't remove these lines.

Key concepts
🎭
scene-inner
Only animatable wrapper
🕐
at(d)
Offset relative to the scene
📋
BODIES[]
Scene order
🔀
switch(i)
Tweens per scene
4

🧱 Assembling the complete HTML

The four streams (scenesHTML, captionsHTML, audioHTML in track 20, animJS) are assembled into a final template string with the paused GSAP timeline registered in window.__timelines["main"]. The background (glow/grid) and the sentinel tl.set({},{},TOTAL) are also part of the assembly.

Four streams, one HTML file

The final assembly is one giant template string that stitches the four streams together. Captions are on alternating tracks (2/4), and audio is on track 20 (special). The GSAP timeline is created paused and registered in window.__timelines["main"] — the HyperFrames player controls it externally.

The sentinel tl.set({}, {}, TOTAL) extends the timeline to the end of the composition, ensuring the progress bar reaches the end even when there are no tweens after the last scene.

Generation of the four streams from S[]
// composition-template.mjs — the four streams
// 1. scenesHTML: .scene.clip with alternating tracks 1/3
const scenesHTML = S.map((s, idx) => `
<section id="s${s.i}" class="scene clip"
data-start="${s.start}" data-duration="${s.dur}"
data-track-index="${s.i % 2 === 1 ? 1 : 3}">
<div class="scene-inner" id="scene-inner-${s.i}">
${BODIES[idx]()}
</div>
</section>`
).join("");

// 2. captionsHTML: same timing as S, tracks 2/4
const captionsHTML = S.map((s, idx) => `
<div class="caption clip" id="cap-${s.i}"
data-start="${s.start}" data-duration="${s.dur}"
data-track-index="${s.i % 2 === 1 ? 2 : 4}">
${CAPTIONS[idx]}
</div>`
).join("");

// 3. audioHTML: data-start = audioStart (starts after the visual LEAD), track 20
const audioHTML = S.map((s) => `
<audio id="a${s.i}" data-start="${s.audioStart}"
data-duration="${s.audioDur}" data-track-index="20"
src="assets/audio/s${s.i}.wav"></audio>`
).join("");

// 4. animJS: GSAP code strings — embedded in the <script>
const animJS = S.map((s) => anim(s.i, s.start)).join("\n ");
Final <script> block — paused timeline, environment, sentinel
// Generated HTML fragment — the final <script>
<script>
window.__timelines = window.__timelines || {};
const tl = gsap.timeline({ paused: true });
const TOTAL = 101.234; // value calculated by the generator

// ambient animation — repeat covers the TOTAL
tl.to("#glow",{scale:1.22,opacity:.55,duration:4.5,yoyo:true,
repeat:Math.ceil(TOTAL/4.5)+1,ease:"sine.inOut"},0);
tl.to("#glow2",{scale:1.18,duration:6,yoyo:true,
repeat:Math.ceil(TOTAL/6)+1,ease:"sine.inOut"},0);
tl.to("#grid",{backgroundPositionY:"+=128",duration:18,
repeat:Math.ceil(TOTAL/18)+1,ease:"none"},0);
tl.fromTo("#progress",{scaleX:0},{scaleX:1,duration:TOTAL,ease:"none"},0);

// scene tweens (generated by animJS)
// ${animJS} — strings generated by anim()

// sentinel: extends the timeline to the end of the composition
tl.set({}, {}, TOTAL);
window.__timelines["main"] = tl;
</script>
🏗️ .bg-layer structure (persistent environment)
#glow + #glow2
Two pulsing radial gradients with yoyo:true covering the entire TOTAL. data-layout-ignore — does not interfere with the HyperFrames layout.
#grid
Grid pattern that scrolls infinitely with backgroundPositionY+=128 throughout the video. Creates a sense of movement.
#grain + #ghost
Subtle SVG grain and ghost text SKILL.md decorative. Both with data-layout-ignore.
Key concepts
⏸️
paused:true
Player controls externally
🔊
Track 20
Reserved for audio
🏁
tl.set({},TOTAL)
End sentinel
📌
__timelines
Public player API
5

📱 9:16 overrides via body.v and the --vertical flag

The generator supports two formats: 16:9 (1920×1080) standard and 9:16 (1080×1920) for Shorts. The flag --vertical switches W and H, adds the class body.v in the generated HTML and applies automatic CSS overrides. The output is always index.html.

One generator, two formats

The same build-index.mjs generates 16:9 and 9:16. The timing logic is identical — only the W/H dimensions change, along with the CSS for the body.v adjusts fonts, padding, and layout. You render twice with the same font, producing two synchronized formats.

Typical workflow: node build-index.mjs → renders 16:9 → node build-index.mjs --vertical → renders 9:16. Both overwrite index.html — never edit the HTML directly.

composition-template.mjs — --vertical detection and W/H swap
// Format detection at the top of the file
const VERT = process.argv.includes("--vertical");
const W = VERT ? 1080 : 1920; // composition width
const H = VERT ? 1920 : 1080; // composition height
const OUT = "index.html"; // always index.html — both formats

// body receives class "v" when --vertical
// In the generated HTML: <body${VERT ? ' class="v"' : ''}>

// Root composition attributes
<div id="composition"
data-start="0" data-duration="${TOTAL}"
data-width="${W}" data-height="${H}">

// Confirmation log in the terminal
consolelog(`${OUT} generated · ${W}×${H} · TOTAL = ${TOTAL}s · ${S.length} scenes`);
S.forEach(s => console.log(
` s${s.i}: start=${s.start} dur=${s.dur} audio@${s.audioStart} (${s.audioDur}s)`));
CSS overrides for body.v — 9:16 layout adjustments
/* Overrides applied automatically when body has class="v" */
body.v .title { font-size: 88px; }
body.v .subhead { font-size: 38px; }
body.v .kicker { font-size: 28px; }
body.v .h2 { font-size: 62px; }
body.v .reg { font-size: 32px; }
body.v .caption { font-size: 36px; bottom: 140px; }
body.v .grid2 { grid-template-columns: 1fr; }
body.v #ghost { font-size: 340px; opacity: .018; }
📐 16:9 — Main YouTube format
node build-index.mjs
Generates 1920×1080. No class v in the body. Default CSS with large fonts for 1080p screens.
📱 9:16 — Shorts / Reels
node build-index.mjs --vertical
Generates 1080×1920. Adds class="v" to the body. CSS overrides adjust fonts and layout for mobile.
Key concepts
📐
1920×1080
16:9 default
📱
1080×1920
9:16 Shorts
🏷️
body.v
Mode CSS class
💾
index.html
Always overwritten
6

🏁 CTA scene9() / case 9 — invariant signature

Scene 9 is the standard signature for all INEMA.CLUB videos. It shows “CONTINUES AT” + INEMA.CLUB with a glow and the URL 🌐 inema.club. It must never be removed, repositioned, or changed—it's part of the channel identity.

Why the CTA is invariant

Scene 9 works as a brand signature: anyone watching an INEMA.CLUB video always sees the same ending. This builds recognition and directs viewers to the site. The template already includes the CTA narration (s9.wav) and the HTML—you only provide the actual duration in the AUDIO[].

Rule: AUDIO[8] (index 8, scene 9) is always the duration of assets/audio/s9.wav. Default narration: "This is INEMA dot CLUB content. Go to: inema dot club."

composition-template.mjs — scene9() and case 9 of anim()
// scene9() — CTA INEMA.CLUB — do not modify
function scene9() {
return `
<div class="cta-wrap" id="s9-wrap">
<div class="cta-eyebrow" id="s9-eye">CONTINUE AT</div>
<div class="cta-logo" id="s9-logo">
<span class="cta-inema" id="s9-inema">INEMA</span>
<span class="cta-club" id="s9-club">.CLUB</span>
</div>
<div class="cta-url" id="s9-url">🌐 inema.club</div>
</div>
`;
}

// case 9 in anim() — CTA animations
case 9:
P(`tl.from("#s9-eye",{y:-20,opacity:0,duration:.5,ease:"power3.out"},${at(0.12)});`);
P(`tl.from("#s9-inema",{y:60,opacity:0,duration:.7,ease:"power4.out"},${at(0.28)});`);
P(`tl.from("#s9-club",{y:60,opacity:0,duration:.7,ease:"power4.out"},${at(0.42)});`);
P(`tl.from("#s9-url",{opacity:0,duration:.5,ease:"power2.out"},${at(0.85)});`);
P(`tl.to("#s9-logo",{textShadow:"0 0 40px #facc15",duration:1,
yoyo:true,repeat:4,ease:"sine.inOut"},${at(1.0)});`);
break;
⚠️
Golden rule: scene9() never goes away

The CTA is in position 9 of BODIES[] and at index 8 of AUDIO[]. Removing or repositioning it breaks the timing of the entire video and removes the channel signature. When adapting the template for a new video, change only scenes 1–8 and update the durations in AUDIO[0..7]. AUDIO[8] is always the CTA WAV.

CTA CSS — INEMA.CLUB visual identity
/* CTA palette — cream + amber with glow */
.cta-eyebrow { font-size: 28px; letter-spacing: .25em; color: #9ca3af; }
.cta-inema { font-size: 128px; font-weight: 800; color: #fef3c7; /* cream */ }
.cta-club { font-size: 128px; font-weight: 800; color: #f59e0b; /* amber */ }
.cta-url { font-size: 36px; color: #60a5fa; letter-spacing: .05em; }
✓ Checklist for adapting the template
  • ✓ Copy the entire template as build-index.mjs
  • ✓ Fill in AUDIO[0..7] with ffprobe — keep AUDIO[8]
  • ✓ Write scene1()–scene8() — keep scene9()
  • ✓ Code case 1–case 8 in anim() — keep case 9
  • ✓ Run the generator and check the timing log before rendering
✗ Never do this with the CTA
  • ✗ Remove scene9() of BODIES[]
  • ✗ Remove AUDIO[8] or leave the array with fewer than 9 entries
  • ✗ Replace scene 9 with content from another video
  • ✗ Change the colors .cta-inema / .cta-club
Key concepts
🏁
scene9() fixed
Default signature
🎨
Cream + amber
INEMA.CLUB identity
🔊
AUDIO[8]
CTA WAV
🔒
case 9 untouched
CTA animations

📋 Module 3.3 Summary

What you learned
  • ✓ AUDIO[] controls 100% of the timing — HTML, GSAP, and audio
  • ✓ S[] = AUDIO.map() with LEAD/TAIL/FADE, calculate start, dur, audioStart, end
  • ✓ sceneN() returns HTML; anim(i,t) generates GSAP strings with a fade on the .scene-inner
  • ✓ Four streams assembled in HTML: scenesHTML, captionsHTML, audioHTML (track 20), animJS
  • ✓ GSAP timeline paused at window.__timelines["main"], sentinel tl.set({},{},TOTAL)
  • ✓ Flag --vertical switches W/H and adds body.v — output always index.html
  • ✓ scene9() / case 9 is the INEMA.CLUB CTA — never remove
Next module
3.4
🧯 Gotchas & fixes
The most common errors when working with the generator — incorrect tracks, duplicate IDs, timing mismatches — and how to diagnose and fix each one before rendering.
Go to module 3.4 →