🧩 build-index.mjs — the project generator
Each video starts by copying scripts/composition-template.mjs such as build-index.mjs at the project root. This file is the generator: it runs once and writes index.html, and you render.
The template already provides premium dark CSS, a persistent background (glow/grid/grain), caption layer, progress bar, and overrides for 9:16. You only need to fill in 4 things: AUDIO[], CAPTIONS[], sceneN() e anim().
Run node build-index.mjs for 16:9 and node build-index.mjs --vertical for 9:16. Both write index.html — render immediately after each generation.
- ✓ Copy the entire template — don’t create it from scratch
- ✓ Keep scene 9 (CTA INEMA.CLUB) unchanged
- ✓ Render immediately after each
node build-index.mjs - ✓ Use
writeFileSyncforindex.html— the template already does this
- ✗ Don't edit
index.htmldirectly — it will be overwritten - ✗ Don't use the Google Fonts CDN—it disappears in headless rendering
- ✗ Don't remove CTA scene 9—it's the standard sign-off
- ✗ Don't mix
--verticaland 16:9 without rerunning the generator
--bg:#0D1321, color variables, local fonts, overrides for 9:16.data-layout-ignore.🔢 AUDIO[] — actual durations measured by ffprobe
The array AUDIO[] is the only timing data input for the entire project. Each element is the duration in seconds of the narration WAV for the corresponding scene, including CTA scene 9.
The duration of a WAV generated by Kokoro varies depending on the text, the speed (--speed 0.98) and the phonetics. If your estimate is wrong, the audio and visuals will be out of sync forever. Always measure.
The command ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/sN.wav prints the exact number in seconds — paste it directly into the array.
Use for i in 1 2 3 4 5 6 7 8 9; do echo "s$i: $(ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 assets/audio/s$i.wav)"; done to generate the entire list at once and paste it into AUDIO[].
🎨 sceneN() — HTML for each scene
Each scene is a JavaScript function that returns an HTML string. The content is inserted into <div class="scene-inner">, inside a <section class="scene clip"> with timing attributes calculated from AUDIO[].
The assembly transforms each item from S (array derived from AUDIO[]) in a <section> with three critical attributes: data-start, data-duration e data-track-index. Alternating tracks (1/3) avoid overlap at scene edges.
- ✓ Reuse the template's CSS classes (
.kicker,.h2,.grid2) - ✓ Give
idunique for each animatable element (s1-word, etc.) - ✓ Off-canvas decorative elements: use
data-layout-ignore - ✓ Animate the
.scene-inner, never the.clipwrapper
- ✗ Duplicate IDs across scenes — GSAP will animate the wrong one
- ✗ Animations in the
.clip— the framework forcesopacity:1in the active clip - ✗ Reorder BODIES without adjusting the corresponding IDs
- ✗ Remove or reposition
scene9(CTA fixed at the end)
✨ anim(i,t) — GSAP tweens per scene
The function anim(i, t) receives the scene index and absolute start time. It pushes GSAP code strings into a local array, which is later concatenated into <script> in the generated HTML.
The generator runs in Node.js, but GSAP executes in the browser (headless Chrome). anim() builds code strings to be embedded in the HTML — when the browser loads it, these strings already have the absolute times calculated from AUDIO[].
The helper at(d) (shortcut to round(t + d)) calculates the offset relative to the start of the scene. Use it in all tweens to keep the code readable.
1. Always animate the .scene-inner, never the .clip wrapper. 2. Use absolute times calculated by at(d) — never hardcode numbers. 3. The FADE=0.45 is the same for input and output across all scenes, ensuring a uniform transition.
Odd-numbered scenes stay on data-track-index="1" and pairs in data-track-index="3". Captions follow the same pattern (2/4). This prevents the “edge overlap” where two frames from adjacent scenes appear simultaneously—a classic error described in references/gotchas.md.
💬 CAPTIONS[] — one short caption per scene
The array CAPTIONS[] has exactly the same length as AUDIO[]. Each string is displayed at the bottom of its corresponding scene as an accessible caption, synchronized via data-start/duration inherited from the same S[idx].
A good caption captures the central idea for the scene in 5–10 words — it does not transcribe the audio. Viewers watching without sound should understand the point of each scene from the caption + visual alone.
The caption runs in data-track-index="2" (odd-numbered scenes) or "4" (pairs) — alternating tracks to avoid overlap with the previous scene at the transition boundary.
- ✓ Short sentence that summarizes the central concept
- ✓ Includes exact technical terms (file names, commands)
- ✓ Readable for viewers watching without audio
- ✓ The s9 caption mentions the site:
inema.club
- ✗ Verbatim audio transcription (too long)
- ✗ Vague: "an important idea about the topic"
- ✗ Full captions — makes quick reading harder
- ✗ Different AUDIO[] length (will cause an error)
On YouTube, captions are indexed by the search engine. Accurate captions with correct technical terms (ffprobe, GSAP, build-index.mjs) improve the video's ranking in technical searches.
⏱️ Single-source timing—audio and animation always in sync
The most important principle of composition-template: AUDIO[] generates everything — data-start, data-duration, the GSAP tween timings e the attributes of the <audio>. No number is duplicated or hardcoded elsewhere.
The three constants LEAD=0.5, TAIL=0.9 e FADE=0.45 combined with the values in AUDIO[] determine each scene’s start time, duration, and fade — and therefore the exact timing of all GSAP tweens. Changing a value in AUDIO[] automatically propagates to everything.
Formulas: dur = LEAD + audioDur + TAIL · audioStart = start + LEAD · Entrance tween = t · Exit tween = end - FADE.
- ✓ Change a WAV → run the generator → everything syncs
- ✓ Add a scene → add an entry in AUDIO[], CAPTIONS[], sceneN(), anim() case N
- ✓ No hardcoded numbers outside AUDIO[]/LEAD/TAIL/FADE
- ✓
npx hyperframes lintdetects track desynchronization
- ✗ Hardcode time in the tween:
gsap.to(..., 14.5) - ✗ Adjust
data-startin the generated HTML — it will be overwritten - ✗ AUDIO[] and CAPTIONS[] arrays of different lengths
- ✗ Use
round()inconsistent → frame offset
📋 Module 2.4 Summary
- ✓ Copy the template as
build-index.mjsand run withnode - ✓ Fill in
AUDIO[]with real durations from ffprobe (includes CTA s9) - ✓ Writing
sceneN()reusing the template’s CSS classes - ✓ Code GSAP tweens in
anim(i,t)always animating the.scene-inner - ✓ Fill in
CAPTIONS[]with short, technical phrases - ✓ Understand the timing formulas:
LEAD=0.5,TAIL=0.9,FADE=0.45
npx hyperframes lint for 0 errors, inspect --samples 16 for 0 layout issues, and generate the final MP4 at --quality high for 16:9 and 9:16.