🔢 AUDIO[] as the single source of timing
The array AUDIO[] contains the REAL durations (measured with ffprobe) for each WAV narration. It's the only variable you need to fill in so that all video timing—HTML, GSAP, and audio—stays automatically synchronized.
The duration array controls everything. No time values appear elsewhere in the code. Change a WAV file, update the corresponding entry in AUDIO[], run the generator — the entire video recalibrates automatically.
This is possible because the generator runs in Node.js and produces a static HTML file with all timings already calculated. The browser receives ready-to-use numbers — it doesn’t need to calculate anything at runtime.
- ✓
data-startedata-durationof each.scene.clip - ✓ Absolute times for all GSAP tweens via
anim(i, s.start) - ✓
data-startof the elements<audio>in track 20 - ✓
TOTALof the composition and progress bar
- ✗ Kokoro generates different durations depending on the text and speed (
--speed 0.98) - ✗ Wrong estimate = permanently out-of-sync audio and visuals
- ✗ Headless rendering doesn’t alert you—the video simply turns out wrong
- ✗ Never guess durations — always use
ffprobe
⏱️ The calculation of S[]—the timing series
The series S is generated by a single AUDIO.map(). Each object carries all the timing needed for scenes, captions, audio, and tweens—calculated once and propagated everywhere.
Three constants control the rhythm: LEAD=0.5 (visual comes in before the voice), TAIL=0.9 (visual holds after the voice) and FADE=0.45 (fade-in/out duration of the .scene-inner). An accumulator t grows with each scene.
Fundamental formulas: dur = LEAD + a + TAIL · audioStart = t + LEAD · end = t + dur.
After generating the HTML, the template displays this in the console: OUT gerado · W×H · TOTAL = Xs · N cenas, followed by one line per scene with start, dur, audio@, and audioDur. Check these numbers before rendering.
🎬 sceneN() and anim(i,t) per scene
Each scene has two functions: sceneN() returns the inner HTML of .scene-inner, e anim(i,t) generates GSAP code strings that will be embedded in the <script> in the final HTML. The fade on the .scene-inner is always generated by the shared part of anim().
The generator runs in Node.js, but GSAP executes in the browser (headless Chrome). anim() builds JavaScript strings with the already calculated absolute times — when the browser loads the HTML, these strings run with exact numbers, without recalculation. The fade of the .scene-inner is shared by all scenes; the internal switch adds each scene's specific tweens.
HyperFrames forces opacity:1 in the active clip. If you try to animate the .clip or the .scene, the fade animation doesn't work. The animatable wrapper is always the child .scene-inner. The template already generates the fades for it—don't remove these lines.
🧱 Assembling the complete HTML
The four streams (scenesHTML, captionsHTML, audioHTML in track 20, animJS) are assembled into a final template string with the paused GSAP timeline registered in window.__timelines["main"]. The background (glow/grid) and the sentinel tl.set({},{},TOTAL) are also part of the assembly.
The final assembly is one giant template string that stitches the four streams together. Captions are on alternating tracks (2/4), and audio is on track 20 (special). The GSAP timeline is created paused and registered in window.__timelines["main"] — the HyperFrames player controls it externally.
The sentinel tl.set({}, {}, TOTAL) extends the timeline to the end of the composition, ensuring the progress bar reaches the end even when there are no tweens after the last scene.
yoyo:true covering the entire TOTAL. data-layout-ignore — does not interfere with the HyperFrames layout.backgroundPositionY+=128 throughout the video. Creates a sense of movement.SKILL.md decorative. Both with data-layout-ignore.📱 9:16 overrides via body.v and the --vertical flag
The generator supports two formats: 16:9 (1920×1080) standard and 9:16 (1080×1920) for Shorts. The flag --vertical switches W and H, adds the class body.v in the generated HTML and applies automatic CSS overrides. The output is always index.html.
The same build-index.mjs generates 16:9 and 9:16. The timing logic is identical — only the W/H dimensions change, along with the CSS for the body.v adjusts fonts, padding, and layout. You render twice with the same font, producing two synchronized formats.
Typical workflow: node build-index.mjs → renders 16:9 → node build-index.mjs --vertical → renders 9:16. Both overwrite index.html — never edit the HTML directly.
v in the body. Default CSS with large fonts for 1080p screens.class="v" to the body. CSS overrides adjust fonts and layout for mobile.🏁 CTA scene9() / case 9 — invariant signature
Scene 9 is the standard signature for all INEMA.CLUB videos. It shows “CONTINUES AT” + INEMA.CLUB with a glow and the URL 🌐 inema.club. It must never be removed, repositioned, or changed—it's part of the channel identity.
Scene 9 works as a brand signature: anyone watching an INEMA.CLUB video always sees the same ending. This builds recognition and directs viewers to the site. The template already includes the CTA narration (s9.wav) and the HTML—you only provide the actual duration in the AUDIO[].
Rule: AUDIO[8] (index 8, scene 9) is always the duration of assets/audio/s9.wav. Default narration: "This is INEMA dot CLUB content. Go to: inema dot club."
The CTA is in position 9 of BODIES[] and at index 8 of AUDIO[]. Removing or repositioning it breaks the timing of the entire video and removes the channel signature. When adapting the template for a new video, change only scenes 1–8 and update the durations in AUDIO[0..7]. AUDIO[8] is always the CTA WAV.
- ✓ Copy the entire template as
build-index.mjs - ✓ Fill in
AUDIO[0..7]with ffprobe — keepAUDIO[8] - ✓ Write
scene1()–scene8()— keepscene9() - ✓ Code
case 1–case 8inanim()— keepcase 9 - ✓ Run the generator and check the timing log before rendering
- ✗ Remove
scene9()ofBODIES[] - ✗ Remove
AUDIO[8]or leave the array with fewer than 9 entries - ✗ Replace scene 9 with content from another video
- ✗ Change the colors
.cta-inema/.cta-club
📋 Module 3.3 Summary
- ✓
AUDIO[]controls 100% of the timing — HTML, GSAP, and audio - ✓
S[] = AUDIO.map()with LEAD/TAIL/FADE, calculate start, dur, audioStart, end - ✓
sceneN()returns HTML;anim(i,t)generates GSAP strings with a fade on the.scene-inner - ✓ Four streams assembled in HTML: scenesHTML, captionsHTML, audioHTML (track 20), animJS
- ✓ GSAP timeline paused at
window.__timelines["main"], sentineltl.set({},{},TOTAL) - ✓ Flag
--verticalswitches W/H and addsbody.v— output alwaysindex.html - ✓
scene9()/case 9is the INEMA.CLUB CTA — never remove