Track map
Detailed content
🗂️ Skill structure
Complete folder anatomy video-explicativo/: the SKILL.md at the center, the three specialized references, and the three scripts that automate the heavy lifting.
The skill root contains SKILL.md (the orchestrator), the subfolder scripts/ with 3 executable files and the subfolder references/ with 3 specification files. Nothing else.
Knowing the folder map is the first step to understanding what the skill does, modifying it without breaking anything, and creating new skills using the same pattern.
Root = SKILL.md; specialization in subfolders; separation of orchestrator/reference/scripts.
pipeline.md describes the 7 steps in the production workflow; house-style.md defines the color palette, typography, and visual rules; gotchas.md list the most common errors and their specific fixes.
Knowing which reference to consult for each question avoids unnecessary reading — each file has a precise scope and doesn’t overlap with the others.
Domain specialization; loaded on demand; no redundancy between files.
fetch-fonts.mjs downloads the local .woff2 fonts; narration-template.sh generates the TTS script with voice pf_dora --speed 0.98; composition-template.mjs is the central generator that produces the index.html final with all scenes.
Each script solves a specific automation problem—together, they eliminate repetitive manual work and ensure deterministic output.
Script-based automation; local sources; TTS; deterministic HTML generation.
The SKILL.md defines 7 numbered steps: (1) confirm the topic, (2) write the script, (3) generate TTS audio, (4) download sources, (5) run the generator, (6) validate the output, (7) deliver the files. Each step points to the right reference.
The numbered flow is what Claude executes in order—understanding it lets you debug which step an error occurred in.
Numbered workflow; responsibility for each step; references linked by step.
The SKILL.md instructs Claude to read house-style.md only before building the visual scenes, and gotchas.md only before validating the output—never all at once.
Preserves context for the actual task instead of spending tokens loading specifications that may not be necessary.
On demand; context savings; conditional reading.
The SKILL.md includes the author's fixed preferences: narration in PT-BR, a premium dark amber palette, dual output (16:9 for YouTube and 9:16 for Shorts), and a final CTA always pointing to INEMA.CLUB.
These defaults save users from having to repeat preferences with every request—they're built into the skill for brand consistency.
Brand defaults; PT-BR; dual output; consistent CTA.
🎨 Premium dark house style
The complete style guide for the explanatory video: background #0D1321, amber accent #FFC300, Sora/Inter/JetBrains Mono typography and the rules that distinguish video scale from web scale.
Six fixed variables: background #0D1321, panel #1D2D44, border #3E5C76, main text #F0EBD8, amber accent #FFC300 and code highlighting #2EC4B6. No other accent color is allowed.
The palette is the channel’s visual identity. Straying from it breaks brand consistency—the amber should be the only warm spot on screen.
Closed color system; amber = sole accent; code in cyan.
Sora for titles and headlines (strong visual impact), Inter for body text (high readability), JetBrains Mono for code and data. The three families cover every video use case.
Mixing fonts outside the system creates visual noise. The rule is simple: title → Sora, text → Inter, code → JetBrains.
Sora for display; Inter for body; JetBrains for code; never use Google Fonts CDN at runtime.
Every scene has a background layer made up of 4 elements: (1) amber radial glow (glow), (2) low-opacity ghost text (ghost text), (3) dot grid (grid) and (4) SVG grain texture (grain). This layer is never animated — it stays static.
The layer creates depth without distraction. Understanding that it’s static prevents the mistake of trying to animate it (which causes a conflict with the framework).
Static layer; 4 elements; depth without distraction.
The sources are downloaded by the script fetch-fonts.mjs in format .woff2 and referenced via @font-face local. The Google Fonts CDN is prohibited because the renderer may not have internet access.
A video rendered without access to the CDN will display fallback fonts — breaking the visual design. Local fonts ensure offline consistency.
local woff2; @font-face; fetch-fonts.mjs; offline-safe.
Video isn’t the web: each scene should have no more than 8-10 visual elements to avoid overwhelming the viewer. Headlines range from 64px (subtitle) to 172px (maximum emphasis) — sizes impossible on a website, but necessary for viewing on TV or mobile.
Designers coming from the web tend to put too much information in each scene. The 8–10 element rule is the antidote.
Maximum 10 elements; large fonts; visual breathing room; viewer focus.
Every video ends with the scene scene9() (CTA), which displays the amber logo and the address INEMA.CLUB in maximum emphasis. This scene is generated automatically by the composition-template and must not be omitted.
The CTA is the channel’s conversion point. Included by default, it ensures no video is delivered without the branded call to action.
scene9() automatic; INEMA.CLUB; channel conversion; never omit.
⚙️ The build-index.mjs generator
The technical heart of the skill: the array AUDIO[] as the single source of truth that calculates timing, assembles scenes, generates captions and tracks, and builds the synchronized GSAP timeline.
AUDIO[] is the only array the author edits: each entry has the MP3 file and duration in seconds. From there, the generator derives everything—scene timing, transition duration, audio track positions in the HTML, and GSAP timeline timestamps.
A single source eliminates sync issues: audio and the timeline can't disagree because they both read the same data.
Single source of truth; AUDIO[] array; automatic derivation.
The generator calculates the array S[] adding up LEAD=0.5s (silence before audio), TAIL=0.9s (silence afterward) and FADE=0.45s (overlapping entrance/exit). For each scene: start = soma das durações anteriores, dur = audio.dur + LEAD + TAIL, audioStart = LEAD.
Understanding the calculation lets you adjust the parameters for videos with a different pace without breaking synchronization.
LEAD 0.5s; TAIL 0.9s; FADE 0.45s; S[] derived from AUDIO[].
Each scene is a function sceneN(s) that returns the scene's HTML string using the timing object s. The animations are declared with anim(element, props, delay), which accumulates a list of GSAP tweens — never direct CSS.
Separating the visual structure (HTML) from the animation (GSAP via anim()) makes it easier to edit one without touching the other.
sceneN() returns HTML; anim() declares tweens; separation of structure and motion.
The generator builds the HTML in 4 blocks: (1) .scene divs with the visual content, (2) .caption divs with synchronized captions, (3) <audio> alternating tracks (scenes 1/3 on track A, 2/4 on track B), and (4) the block <script> with the GSAP timeline initialized in a paused state.
Knowing the structure of the 4 blocks lets you edit the output manually when needed without losing the synchronization logic.
4 blocks; paused timeline; playback controlled by audio events.
Running the generator with the flag --vertical, it adds the class .v to the <body> and applies CSS overrides that reposition elements for the 9:16 format (1080×1920px). The generated file is always index.html, overwriting the 16:9 version.
Understanding the override mechanism prevents confusion between the two formats and explains why both share the same base file.
body.v; CSS overrides; --vertical flag; one file, two formats.
The composition-template defines scene9() as the hard-coded final CTA scene. It is always injected after the content scenes, regardless of how many scenes the video has — if the video has 7 scenes, the last one is always the CTA scene9.
Knowing it's automatic prevents duplicating the CTA manually, and knowing it can't be omitted ensures brand consistency across all output.
scene9() hard-coded; always last; INEMA.CLUB; never duplicate.
🧯 Gotchas & fixes
The errors that come up every time and the exact fix for each one: from which element to animate to font issues, overlapping clips, and generator determinism.
The HyperFrames framework sets opacity:1 in all the .clip during frame capture — opacity animations on them are ignored. The correct target for fade-in/out is always .scene-inner, the scene's internal container.
This is the most common gotcha #1. The video looks correct in the browser, but the final render shows flickering scenes or no transitions.
Correct target = .scene-inner; .clip = forbidden for opacity; framework override.
Odd-numbered scenes (1, 3, 5…) use the <audio> track A; even-numbered scenes (2, 4, 6…) use track B. This allows one track to finish naturally while the other already loads the next audio — eliminating the problem of overlapping clips where two audio tracks play at the same time.
If you put all the audio on the same track, there will inevitably be overlap during transitions with FADE. The alternating pattern solves this structurally.
Track A = odd; Track B = even; overlap eliminated; preload on the inactive track.
Purely decorative elements positioned outside the visible area (e.g., overflowing glows, negatively positioned ghost texts) should have the attribute data-layout-ignore. Without it, Remotion’s layout engine includes them in the bounding box and may crop or shift the scene.
Off-canvas decorative elements without the attribute are the #1 cause of compositions with unexpected cropping in the render.
data-layout-ignore; off-canvas; bounding box; invisible decorative elements.
The gotchas.md prohibits any @import url(fonts.googleapis.com) in the generated HTML. The sentinel variable google_fonts_import is used in the templates to ensure the generator throws an error if it detects the string in the output—always forcing the use of @font-face local.
The render environment has no internet access. An external import causes missing fonts, text in a sans-serif fallback, and a completely broken look.
Google Fonts CDN is prohibited; use local @font-face; google_fonts_import sentinel.
On Windows and in git-bash, ffmpeg tries to read interactive stdin even in non-interactive calls, causing the process to hang indefinitely. The flag -nostdin disables this behavior. All ffmpeg commands in the skill scripts include it by default.
If you remove the flag when adapting a script, the process will silently hang on Windows without an error message — making it difficult to diagnose.
-nostdin; Windows/git-bash; interactive stdin; stuck process.
The generator and templates prohibit Date.now(), Math.random() and any fetch() inside the renderable HTML. Remotion renders each frame independently — non-deterministic code produces inconsistent frames and video flickering.
This is the hardest cause to diagnose: the video looks correct in preview but jitters in the final render. The simple rule: if a value changes between runs, it doesn't belong in the template.
Fully deterministic; no Date.now; no Math.random; no runtime fetch; fixed values.