A map of ~45 projects that produce or process video and images: deterministic rendering, local and cloud AI generation, direction, orchestrators, and image, voice, and music support — each with what it does, how to run it, and a link to its guide.
This page brings together tools that evolved in parallel. The idea is to stop rediscovering “which project does X”: find the right path here in seconds and go straight to the guide.
Deterministic, cloud AI, local AI, image, direction, orchestration, and post-production. Every video request falls into one of these.
Most run on your machine/DGX: flux2-klein (image), inemavox (voice), HyperFrames/Remotion/pixflow (rendering). Local video AI is optional.
The orchestrators reuse the same building blocks — direction from MDD, images from inemaimg, voice from inemavox, rendering from HyperFrames.
Before writing a new pipeline, see what has already been measured. These five solved the problem and left notes — including about the flaws.
Gold standard for an animated story: derived character anchors (not generated in parallel), character description repeated in every scene, narration by paragraph, no on-screen text. Outputs in ~/projetos/output/videos-agnes/.
NOTAS-API.md includes the rules that affect your bottom line, the model's flaws, and the rate limits. Includes the filter trap: a child character deterministically returns HTTP 400.
Access via MCP and CLI, with account, defaults, and dubbing with a cloned voice. API notes measured in ~/projetos/wifi/KLING-NOTAS-API-2026-08-02.md.
2.5D parallax with real depth, grain, LUT, and vignette. Deterministic: same input, same result — the opposite of a video generator.
The same character carries through a folder, comic, motion comic, and entire series, maintaining visual consistency.
Start with the type of video you want. The flow below points you in the right direction; the catalog has details on each component.
Camera/parallax over images, motion graphics, animated comics. → pixflow, inemafilme, inemaref, HyperFrames, Remotion.
Story, episode, consistent character. → videos-agnes (simple, US$ 0) or videoanima (with shot breakdown and swappable provider).
Cloud/subscription: klingaimcp, seedance2. Local on GPU: VideosDGX2 (Wan 2.2), HuMo, skyreelsv3.
Scripted talking avatar. → heygenmcp (subscription credits), videosavatar, or HuMo for local lip-sync.
Explainer, app demo, video lesson, marketing reels. → video producer, video-explicativo, video-demonstrativo, videos-cursos-inema, timesmkt3.
Dub, cut fillers, add captions, replace the soundtrack. → dublar pro, inemavox, musicaclone, inemadlp.
Each card: what it does, input→output/how to run it, and the link. Where a published page exists, the link goes to the guide; otherwise, to the repo. The colored dot indicates the status; (local) = no remote on GitHub.
Cinematic film from still images: parallax with real depth + grain/LUT/vignette/bloom, camera, and transitions.
YAML (movie spec) + images → MP4 16:9. Depth-Anything-V2 + WebGL/GLSL + Remotion + FFmpeg.
repo ↗Structured script (historia.json) becomes a narrated parallax film about the pixflow engine — camera chosen for emotion, music, and SFX.
historia.json → MP4. It doesn't write the story: that's the skill roteiro.
Narrative factory: character folder → comic → motion comic and an entire series from a topic.
JSON (bible + script) + references → MP4 in Forms A (slideshow), B (camera journey), and C (directed). Python + headless Chrome + TTS + FFmpeg.
guide ↗Core engine behind almost all skills: animates HTML+GSAP and records via headless Chrome. cchyperframes = editor.
HTML + TTS narration → MP4 16:9/9:16. Node + Playwright + Kokoro + FFmpeg.
cchyperframes ↗ · skill ↗React video framework (kinetic text/data layer). remotion-templates = 81 ready-to-use components.
TSX + media → MP4/WebM/GIF via FFmpeg.
remotion ↗ · templates ↗The entire page is a single cinematic shot that progresses as the visitor scrolls. Pure code (GSAP) or footage from a generator.
topic/footage → scroll-film site. It’s not a slide deck or corporate website.
guide ↗MCP server (119 tools, guardrailed) for agents to edit video: cuts, merges, crops, overlays, captions, effects.
JSON specs → edited video + analysis. Python 3.11 + FFmpeg + HyperFrames.
guide ↗Story → cinematic breakdown → A/B keyframes → clips → narrated vertical film. Measures each clip’s motion and redoes any that come out static. Being adjusted: see the post-mortem in the repo.
YAML shot breakdown → MP4 9:16. Image and video provider can be switched via parameter (agnes/inemaimg/kie · agnes/kling/klingai/kie).
guide ↗Story/short story → narrated animated film via Agnes AI, delivered on Telegram. The simple path: no breakdown, no provider selection.
story → model sheet + 2 images/scene + clips + local narration → MP4.
guide ↗Kling AI via MCP and CLI with your own subscription (Pro/SVIP) — includes dubbing with a cloned voice. Measured API notes.
image/prompt → clip. Accepts a local path via CLI; via fal.ai, it requires a keyframe at a public URL.
guide ↗Talking avatar in HeyGen using the subscription credits (via MCP), not the API wallet.
script → MP4 9:16. The sister heygen-cli does the same through the API key (paid separately).
Script → talking avatar in HeyGen, using the configured looks/characters in the account.
text → MP4 9:16, PT-BR voice, avatar_iii engine.
guide ↗A photo becomes a miniature set inside a practical-effects studio rig, and the simulation destroys the miniature in a single behind-the-scenes shot.
photo → behind-the-scenes video. Recipes tested for Kling and Agnes.
guide ↗Finished interior photo → renovation time-lapse video: generates the “before construction” with the same camera and architecture, and the before→after video.
1 photo → before-and-after MP4. GPT Image + Higgsfield Seedance.
guide ↗Generates the prompt cinematic, structured for Seedance 2.0; generation runs on FAL.ai. seedance2en is the PT/EN variant.
description → 300–450-word prompt + generation link.
repo ↗Local T2V + I2V on DGX Spark with Wan 2.2 (14B MoE and fast 5B), via ComfyUI + web UI.
text/image → MP4. UI on port :7862. Weights ~25–90GB VRAM.
guide ↗Human-centric, audio-driven video (fine lip sync) with HuMo 17B/1.7B. It’s the local alternative to a cloud avatar.
text + audio (+ image) → 480/720p. 32G+ GPU. ComfyUI support.
(local — no remote)Complete SkyReels V3 web UI: reference-to-video, video-to-video, and talking avatar, with a job queue.
text/image/audio → video. PyTorch + diffusers + FastAPI.
guide ↗Desktop app (Electron) with 4 studios — image, video, lip sync, and cinema — aggregating 200+ models.
various → video/image. macOS/Windows/Linux.
repo ↗Image server with model hot-swapping: flux2-klein (ecosystem default), Qwen-Edit, FLUX.2-dev, ERNIE. No content filter, accepts native refs, has a seed.
prompt (+1–4 refs) → PNG. HTTP API at localhost:8000. ~31s/image on the DGX.
Standalone image via Agnes (agnes-image-2.1-flash), at no cost. CLI gerar.py. Watch the filter: child returns HTTP 400.
prompt (+refs) → PNG. Companion to videos-agnes.
The Agnes API dissected across ~70 calls: text, image, and video. The rules that make money, model flaws, and rate limits.
measured documentation — read before writing a new integration.
guide ↗AI image upscaling (Real-ESRGAN, 2x/4x) — useful for refining frames before rendering.
low-res image → 2x/4x image. Electron app.
repo ↗KairoBoost prompt engine (art direction) + end-to-end case (stills → 9:16 video with narration).
intent + preset → refined prompt → inemaimg + HyperFrames.
repo ↗Dynamic Direction Master: turns a topic into a direction package (scene, camera, continuity) for AI generators.
topic → panel storyboard + final and negative prompts for Seedance/Kling/Veo/Runway.
guide ↗Camera direction (18 moves, validated A/B rules) over images + narration.
images + narration → JSON shot breakdown + pixflow YAML → MP4.
guide ↗Generates a motor-agnostic editing plan (Viral5 / Hero / Save-the-Cat), with a Python package vpe.
topic + preset → plano-edicao.json + RESUMO → render via HyperFrames.
guide ↗Local-first filmmaking engine driven by knowledge packs (concept proven in the Hormozi-12 case).
topic → images (flux2) + narration (inemavox) + parallax (pixflow).
guide ↗Complete production pipeline: plan → direction + image → voice → render in 3 layers (cinema + text + illustration).
link/source → professional 16:9 + 9:16 MP4. Reuses mdd, flux2, inemavox, HyperFrames, Remotion.
guide ↗Narrated educational video in PT-BR, from script to INEMA.CLUB CTA. Motion graphics — doesn’t show a real app.
topic → scenes → local TTS → HTML/GSAP → HyperFrames → MP4 16:9 + 9:16.
guide ↗Automatic web app walkthrough: navigates the app for real, captures the actual screens and narrates them, with an animated cursor and zoom.
link → UI capture → narration → 16:9 + 9:16 MP4.
guide ↗Generates videos for an INEMA course at 3 selectable levels: landing page, learning paths, and in-depth lesson by module/day.
course site → script → inemavox → HyperFrames → MP4.
repo ↗Promotional Reels for 12 audiences, with human gate in the middle: the bot writes the text and stops until you approve.
audience catalog → copy → avatar → 9:16 reel.
guide ↗Same idea, three videos for different audiences instead of one: reach, authority, and promotional. Separate system, with its own engine and templates.
/aprovar C#N unlocks avatar, download, and reel.
Queues (FIFO) and runs renders for the skills above in the background, with notifications and a dashboard. Serializes the GPU.
command (CLI/Telegram) → enqueue → worker → MP4 + notification.
repo ↗Marketing agent team: generates Reels/ads (native Remotion, quick) and delegates long-form original work to mkvideos. Canonical for the timesmkt/2/3, imkt4/5 family.
/campaign → 5 stages → image (inemaimg) + voice (chatterbox) + video → publishes.
guide ↗GPU-first voice suite: zero-shot cloning (Chatterbox) + transcription (Whisper) + dubbing + music and SFX search (Pixabay/Freesound).
text/URL → dubbed wav/video. FastAPI daemon :8010. Default voice in the ecosystem: rachel.
Same suite without a GPU: Groq Whisper + Edge TTS. For VPS/laptop.
text/URL → audio/subtitles/clips.
repo ↗Open-source zero-shot voice cloning library (23 languages). Foundation for inemavox.
text + 10s ref → wav. pip install chatterbox-tts.
Soundtrack for videos — music cloning/generation, so you don’t have to rely only on a library.
reference/prompt → audio track.
guide ↗Ecosystem downloader (yt-dlp): where music, ambience, and SFX come from when a video needs audio.
URL → local audio/video.
guide ↗Dubbing + editing + transcription with a web interface. Multiple ASR/TTS/LLM options. Canonical — dublar/2/4 are superseded.
video → dubbed video. FastAPI + Next. Repo: dublarv5.
Local PT-BR TTS built into HyperFrames skills (voices pf_dora / bella / rachel). A lightweight alternative to inemavox.
text → wav, within the skills pipeline.
(within the skills)Smart editing: removes fillers (umm/uh), color grading, captions, and overlays via LLM→EDL.
raw footage → assembled video. FFmpeg + LLM + render.
repo ↗The local-first paths depend on a few services and binaries. Start the ones your path uses.
flux2-klein server and more.
# starts the image server cd ~/projetos/inemaimg docker compose up # :8000
Narration / cloning / SFX.
# voice daemon cd ~/projetos/inemavox uvicorn app:api # :8010
Node + FFmpeg + Chrome (HyperFrames).
# quick check node -v && ffmpeg -version echo $CHROMIUM_BIN
Example of the most common journey: a topic becomes a professional video by reusing the building blocks. Switch the skill based on the type (explainer, demo, series, marketing).
The orchestrators call these services; keep them running.
cd ~/projetos/inemaimg && docker compose up -d # image :8000 cd ~/projetos/inemavox && uvicorn app:api & # voice :8010
Explainer → video-explicativo. App demo → video-demonstrativo. Pro/landing page → videoprodutor. Story → videos-agnes or videoanima. Series/comic → inemaref.
# each skill responds to help /video-explicativo "o que é engenharia de contexto"
For paths using a paid provider, check the pacing and cost before the first call.
python3 rodar.py exemplo --so-decupagem # videoanima: validates without generating
For batches, queue them in mkivideos and move on — it serializes GPU work and notifies you when it’s done.
/mkivideos explicativo "tema" # queues and renders in the background
The output goes into the default output folder; move the final file to its destination.
ls ~/projetos/output/<projeto>/ # MP4 16:9 and/or 9:16
Several projects have variants. Use the canonical ones; archive or remove the rest.
ivox2 is for research only — disposable.skyreels, humo2, HuMo_clean, HuMo_velho. VideosDGX is the previous generation of DGX2.videosanimados e mkvideos are older engines from the same family.dublar, dublar2, dublar4 are older versions.timesmkt/timesmkt2 superseded; imkt4 duplicate; imkt5 = new architecture (still no video generator).*-video these aren't toolsagentic-video, aiosagi-video, consultoria2k-video, hermes21c-video, mapacliente-video, my-video, opus48-video, polyskills-video, segrobot-video, uany-video, vla-video and others are HyperFrames compositions by a single piece — don't look for a rendering engine there.