Narrated vertical pipeline, with image and video providers selectable by parameter β and a YAML contract you can review before spending anything.

The difference between a film and a slideshow is in the shot breakdown: each shot knows its framing, movement, and where the gesture ends. That's what videoanima writes before generating any images.
Framing, angle, movement, speed, look, and pacing become prompt fragments β vocabulary distilled from a style analysis, not improvised adjectives.
Image in agnes, inemaimg or kie; video at agnes, kling, klingai or kie. Each measured pitfall lives in its adapter.
Each clip is measured after generation; anything that comes out still is redone. There's a negative prompt against still frame and A/B keyframes so the generator has something to animate.
The pipeline runs end to end, but it is in an open correction cycle based on the films it produced. The first cut came out poorly (static scenes, inconsistent character, choppy narration) β what was measured and what changed as a result is in the post-mortem doc/pos-mortem-montanha-leao.md. Still open: the narration stretches shots beyond the planned pacing, SFX are added without prior listening, and linked references between shots (shot N using Nβ1) haven't been implemented.
Everything lives in output: a standalone ID resolves to ~/projetos/output/videoanima/<id>/decupagem.yaml. O historias/<id>/ from the repo is just a seed β on the first run, the shot breakdown is copied to the output, and the copy there becomes the active one.
The pacing check runs first, and the cost preview appears before the first paid API call. --so-decupagem validates without generating anything.
One anchor goes through text2img, and the others derive from it through img2img β generating in parallel would produce different individuals because none of these generators has an identity seed.
Slow motion and speed ramps, handheld, zoom punches, a ducked soundtrack, SFX, and loudnorm are added during editing with ffmpeg.
Python 3 and ffmpeg are required. Local services and keys depend on the providers you choose.
The entire edit is done with ffmpeg. Without it, the film can't be completed.
# check ffmpeg -version python3 --version
Local TTS in localhost:8010, chatterbox engine. The voice comes from the shot breakdown (voz:).
# needs to respond curl localhost:8010/health
Local image server in localhost:8000. Required for --img inemaimg β the path for a child character, which gets caught by the Agnes filter.
# only if using --img inemaimg curl localhost:8000/health
Local LLM in localhost:11434, used by the automatic shot breakdown (decupagem_llm.py).
# only for LLM-based shot breakdown curl localhost:11434/api/tags
Read at runtime from .env from the ecosystem β nothing is copied into the repo.
# AGNES_API_KEY, KIE_API_KEY, FAL_KEY # in ~/projetos/openpcbotv2/.env # or ~/projetos/wifi/.env
klingai e kie need the keyframe at a public URL β pair it with --img agnes or --img kie.
# valid combination --img kie --video klingai
Run all commands from the repo root. The order matters: validate pacing and cost before making any paid calls.
Checks pacing and shows the cost preview. Does not generate images, clips, or audio.
python3 rodar.py exemplo --so-decupagem # validates pacing + cost, doesn't spend anything
The first run copies the seed from the repo to the output. From then on, edit the one in the output β that's the one that counts.
# seed in the repo (first time only) historias/exemplo/ # the live one, which you edit ~/projetos/output/videoanima/exemplo/decupagem.yaml
Generates the character anchors and A/B keyframes, assembles the contact sheet, and stops. This is where you catch inconsistent characters before paying for the clips.
python3 rodar.py exemplo --so-imagens
Choose the image and video providers. The default for both is agnes (US$ 0 cost, batch submission).
python3 rodar.py exemplo --img agnes --video agnes # Kling via fal.ai, with a keyframe at a public URL python3 rodar.py exemplo --img kie --video klingai
Measured on 2026-08-01: Agnes consistently returns HTTP 400 for "boy"/"child" β the same prompt works with an adult.
python3 rodar.py minha-historia --img inemaimg --video agnes
--sim doesn't ask before spending. --prompt-minimo sends only the motion instructions to the video generator β useful when the long prompt is getting in the way of the animation.
python3 rodar.py exemplo --sim --prompt-minimo
Each step skips anything already on disk. Delete the artifact you want to redo and run it again β only that artifact is regenerated.
# redo only the clip for shot 07 rm ~/projetos/output/videoanima/exemplo/clipes/p07*.mp4 python3 rodar.py exemplo
Frames from the vertical films produced during the correction cycle β the same ones that informed the post-mortem.


tipo_plano and the environment bible; the state plan is not sent to any generator.The roadmap is driven by defects measured in the films themselves β each item came from a cut that went wrong.
still frame and a gate that measures the clipβs motion and redoes anything that comes out static.tipo_plano: tableau is not sent to a video generator β the camera moves over the still image. blocos applies the right look to the shot range, instead of all 50 shots inheriting a global "daylight" setting.