PTENES
TRACK 3

🔧 Inside the skill

Complete reverse engineering of the skill video-explicativo. You’ll open the hood: folder, references, house style, the generator build-index.mjs and the gotchas that save hours of debugging.

4
Modules
24
Topics
~2h
Duration
Advanced
Level
AUDIO[] single source of truth S[] timing LEAD/TAIL/FADE audio tracks scenes 1/3 alternating GSAP timeline slower, audio seek index.html synchronized & renderable

AUDIO[] — the single source that generates synchronized audio + GSAP timeline

Track map

Detailed content

3.1~30 min

🗂️ Skill structure

Complete folder anatomy video-explicativo/: the SKILL.md at the center, the three specialized references, and the three scripts that automate the heavy lifting.

What it is:

The skill root contains SKILL.md (the orchestrator), the subfolder scripts/ with 3 executable files and the subfolder references/ with 3 specification files. Nothing else.

Why learn:

Knowing the folder map is the first step to understanding what the skill does, modifying it without breaking anything, and creating new skills using the same pattern.

Key concepts:

Root = SKILL.md; specialization in subfolders; separation of orchestrator/reference/scripts.

What it is:

pipeline.md describes the 7 steps in the production workflow; house-style.md defines the color palette, typography, and visual rules; gotchas.md list the most common errors and their specific fixes.

Why learn:

Knowing which reference to consult for each question avoids unnecessary reading — each file has a precise scope and doesn’t overlap with the others.

Key concepts:

Domain specialization; loaded on demand; no redundancy between files.

What it is:

fetch-fonts.mjs downloads the local .woff2 fonts; narration-template.sh generates the TTS script with voice pf_dora --speed 0.98; composition-template.mjs is the central generator that produces the index.html final with all scenes.

Why learn:

Each script solves a specific automation problem—together, they eliminate repetitive manual work and ensure deterministic output.

Key concepts:

Script-based automation; local sources; TTS; deterministic HTML generation.

What it is:

The SKILL.md defines 7 numbered steps: (1) confirm the topic, (2) write the script, (3) generate TTS audio, (4) download sources, (5) run the generator, (6) validate the output, (7) deliver the files. Each step points to the right reference.

Why learn:

The numbered flow is what Claude executes in order—understanding it lets you debug which step an error occurred in.

Key concepts:

Numbered workflow; responsibility for each step; references linked by step.

What it is:

The SKILL.md instructs Claude to read house-style.md only before building the visual scenes, and gotchas.md only before validating the output—never all at once.

Why learn:

Preserves context for the actual task instead of spending tokens loading specifications that may not be necessary.

Key concepts:

On demand; context savings; conditional reading.

What it is:

The SKILL.md includes the author's fixed preferences: narration in PT-BR, a premium dark amber palette, dual output (16:9 for YouTube and 9:16 for Shorts), and a final CTA always pointing to INEMA.CLUB.

Why learn:

These defaults save users from having to repeat preferences with every request—they're built into the skill for brand consistency.

Key concepts:

Brand defaults; PT-BR; dual output; consistent CTA.

View Full
3.2~30 min

🎨 Premium dark house style

The complete style guide for the explanatory video: background #0D1321, amber accent #FFC300, Sora/Inter/JetBrains Mono typography and the rules that distinguish video scale from web scale.

What it is:

Six fixed variables: background #0D1321, panel #1D2D44, border #3E5C76, main text #F0EBD8, amber accent #FFC300 and code highlighting #2EC4B6. No other accent color is allowed.

Why learn:

The palette is the channel’s visual identity. Straying from it breaks brand consistency—the amber should be the only warm spot on screen.

Key concepts:

Closed color system; amber = sole accent; code in cyan.

What it is:

Sora for titles and headlines (strong visual impact), Inter for body text (high readability), JetBrains Mono for code and data. The three families cover every video use case.

Why learn:

Mixing fonts outside the system creates visual noise. The rule is simple: title → Sora, text → Inter, code → JetBrains.

Key concepts:

Sora for display; Inter for body; JetBrains for code; never use Google Fonts CDN at runtime.

What it is:

Every scene has a background layer made up of 4 elements: (1) amber radial glow (glow), (2) low-opacity ghost text (ghost text), (3) dot grid (grid) and (4) SVG grain texture (grain). This layer is never animated — it stays static.

Why learn:

The layer creates depth without distraction. Understanding that it’s static prevents the mistake of trying to animate it (which causes a conflict with the framework).

Key concepts:

Static layer; 4 elements; depth without distraction.

What it is:

The sources are downloaded by the script fetch-fonts.mjs in format .woff2 and referenced via @font-face local. The Google Fonts CDN is prohibited because the renderer may not have internet access.

Why learn:

A video rendered without access to the CDN will display fallback fonts — breaking the visual design. Local fonts ensure offline consistency.

Key concepts:

local woff2; @font-face; fetch-fonts.mjs; offline-safe.

What it is:

Video isn’t the web: each scene should have no more than 8-10 visual elements to avoid overwhelming the viewer. Headlines range from 64px (subtitle) to 172px (maximum emphasis) — sizes impossible on a website, but necessary for viewing on TV or mobile.

Why learn:

Designers coming from the web tend to put too much information in each scene. The 8–10 element rule is the antidote.

Key concepts:

Maximum 10 elements; large fonts; visual breathing room; viewer focus.

What it is:

Every video ends with the scene scene9() (CTA), which displays the amber logo and the address INEMA.CLUB in maximum emphasis. This scene is generated automatically by the composition-template and must not be omitted.

Why learn:

The CTA is the channel’s conversion point. Included by default, it ensures no video is delivered without the branded call to action.

Key concepts:

scene9() automatic; INEMA.CLUB; channel conversion; never omit.

View Full
3.3~30 min

⚙️ The build-index.mjs generator

The technical heart of the skill: the array AUDIO[] as the single source of truth that calculates timing, assembles scenes, generates captions and tracks, and builds the synchronized GSAP timeline.

What it is:

AUDIO[] is the only array the author edits: each entry has the MP3 file and duration in seconds. From there, the generator derives everything—scene timing, transition duration, audio track positions in the HTML, and GSAP timeline timestamps.

Why learn:

A single source eliminates sync issues: audio and the timeline can't disagree because they both read the same data.

Key concepts:

Single source of truth; AUDIO[] array; automatic derivation.

What it is:

The generator calculates the array S[] adding up LEAD=0.5s (silence before audio), TAIL=0.9s (silence afterward) and FADE=0.45s (overlapping entrance/exit). For each scene: start = soma das durações anteriores, dur = audio.dur + LEAD + TAIL, audioStart = LEAD.

Why learn:

Understanding the calculation lets you adjust the parameters for videos with a different pace without breaking synchronization.

Key concepts:

LEAD 0.5s; TAIL 0.9s; FADE 0.45s; S[] derived from AUDIO[].

What it is:

Each scene is a function sceneN(s) that returns the scene's HTML string using the timing object s. The animations are declared with anim(element, props, delay), which accumulates a list of GSAP tweens — never direct CSS.

Why learn:

Separating the visual structure (HTML) from the animation (GSAP via anim()) makes it easier to edit one without touching the other.

Key concepts:

sceneN() returns HTML; anim() declares tweens; separation of structure and motion.

What it is:

The generator builds the HTML in 4 blocks: (1) .scene divs with the visual content, (2) .caption divs with synchronized captions, (3) <audio> alternating tracks (scenes 1/3 on track A, 2/4 on track B), and (4) the block <script> with the GSAP timeline initialized in a paused state.

Why learn:

Knowing the structure of the 4 blocks lets you edit the output manually when needed without losing the synchronization logic.

Key concepts:

4 blocks; paused timeline; playback controlled by audio events.

What it is:

Running the generator with the flag --vertical, it adds the class .v to the <body> and applies CSS overrides that reposition elements for the 9:16 format (1080×1920px). The generated file is always index.html, overwriting the 16:9 version.

Why learn:

Understanding the override mechanism prevents confusion between the two formats and explains why both share the same base file.

Key concepts:

body.v; CSS overrides; --vertical flag; one file, two formats.

What it is:

The composition-template defines scene9() as the hard-coded final CTA scene. It is always injected after the content scenes, regardless of how many scenes the video has — if the video has 7 scenes, the last one is always the CTA scene9.

Why learn:

Knowing it's automatic prevents duplicating the CTA manually, and knowing it can't be omitted ensures brand consistency across all output.

Key concepts:

scene9() hard-coded; always last; INEMA.CLUB; never duplicate.

View Full
3.4~30 min

🧯 Gotchas & fixes

The errors that come up every time and the exact fix for each one: from which element to animate to font issues, overlapping clips, and generator determinism.

What it is:

The HyperFrames framework sets opacity:1 in all the .clip during frame capture — opacity animations on them are ignored. The correct target for fade-in/out is always .scene-inner, the scene's internal container.

Why learn:

This is the most common gotcha #1. The video looks correct in the browser, but the final render shows flickering scenes or no transitions.

Key concepts:

Correct target = .scene-inner; .clip = forbidden for opacity; framework override.

What it is:

Odd-numbered scenes (1, 3, 5…) use the <audio> track A; even-numbered scenes (2, 4, 6…) use track B. This allows one track to finish naturally while the other already loads the next audio — eliminating the problem of overlapping clips where two audio tracks play at the same time.

Why learn:

If you put all the audio on the same track, there will inevitably be overlap during transitions with FADE. The alternating pattern solves this structurally.

Key concepts:

Track A = odd; Track B = even; overlap eliminated; preload on the inactive track.

What it is:

Purely decorative elements positioned outside the visible area (e.g., overflowing glows, negatively positioned ghost texts) should have the attribute data-layout-ignore. Without it, Remotion’s layout engine includes them in the bounding box and may crop or shift the scene.

Why learn:

Off-canvas decorative elements without the attribute are the #1 cause of compositions with unexpected cropping in the render.

Key concepts:

data-layout-ignore; off-canvas; bounding box; invisible decorative elements.

What it is:

The gotchas.md prohibits any @import url(fonts.googleapis.com) in the generated HTML. The sentinel variable google_fonts_import is used in the templates to ensure the generator throws an error if it detects the string in the output—always forcing the use of @font-face local.

Why learn:

The render environment has no internet access. An external import causes missing fonts, text in a sans-serif fallback, and a completely broken look.

Key concepts:

Google Fonts CDN is prohibited; use local @font-face; google_fonts_import sentinel.

What it is:

On Windows and in git-bash, ffmpeg tries to read interactive stdin even in non-interactive calls, causing the process to hang indefinitely. The flag -nostdin disables this behavior. All ffmpeg commands in the skill scripts include it by default.

Why learn:

If you remove the flag when adapting a script, the process will silently hang on Windows without an error message — making it difficult to diagnose.

Key concepts:

-nostdin; Windows/git-bash; interactive stdin; stuck process.

What it is:

The generator and templates prohibit Date.now(), Math.random() and any fetch() inside the renderable HTML. Remotion renders each frame independently — non-deterministic code produces inconsistent frames and video flickering.

Why learn:

This is the hardest cause to diagnose: the video looks correct in preview but jitters in the final render. The simple rule: if a value changes between runs, it doesn't belong in the template.

Key concepts:

Fully deterministic; no Date.now; no Math.random; no runtime fetch; fixed values.

View Full
← All tracks Track 4: Applications →