O otv transcribes, splits the speech into numbered units, and asks the model only for a score from 0 to 10 for each ID. The code converts the ID to a timestamp and makes the cut β so it never lands in the middle of a word, and running it twice produces the same result.

A 20- to 30-minute lesson, podcast, or talk goes in; a output.mp4 about two minutes long with the best content. Each phase writes a JSON file that you can read, edit, and re-render at no cost.
It sees [042] 4.2s talking_head "texto" and returns a score. Timing comes from the word-level timestamped transcription, so the cut boundary is a real word boundary β never an invented timestamp.
The selection (backpack by score, quota per topic, hook and closing anchors, coherence) is code. Even notas.json, even plan.json. Editing the plan by hand and rendering again costs zero model calls.
Transcription with Groq or local Whisper; scoring with GLM, Gemini, Ollama, or Claude Code itself; TTS with inemavox or ElevenLabs; image generation with flux-2-klein. One line in the config.yaml or a flag.
Each phase reads the previous phaseβs artifact and writes its own. Rerunning one phase does not redo the others β and existing files are reused unless you ask --forcar.
The default. Keeps the presenter and cuts only what works. Variant A+ (--substituir gerado) replaces presenter segments with generated illustrations and preserves the original audio.
Keeps only slides, screen demos, and charts; the model writes a script in Brazilian Portuguese, narrated with TTS, while the original audio becomes a bed at β18 dB.
Same selection as B, without the narration layer β for when the visual material already explains itself.
No container or build. Python 3.12, ffmpeg, and the keys in .env as always.
Handle all the cutting, audio mixing, and headline. Rendering runs inside a scope with a memory limit.
# Debian/Ubuntu sudo apt install ffmpeg
yt-dlp to download, PySceneDetect to find the cuts, mediapipe and OpenCV to detect faces.
# in the project root pip install -r requirements.txt
Read at runtime from ~/projetos/openpcbotv2/.env e ~/projetos/wifi/.env. No keys are copied into the project.
# what is used GROQ_API_KEY OPENROUTER_API_KEY FAL_KEY # only in mode A+
All commands below are real and were run during end-to-end validation. The <id> is the name of the folder created in trabalho/.
One line does it all: downloads, transcribes, detects scenes, scores, selects, and renders. The result is copied to ~/projetos/output/otimizevideo/<id>/.
python3 otv.py run "https://www.youtube.com/watch?v=..." --modo A --alvo 120
O status lists the artifacts, headline, and each segment with its timestamp, duration, and visual classification.
python3 otv.py status <id> # plan: mode A Β· 136.6s across 18 segments
Each phase logs time and cost in trabalho/<id>/custos.json. A typical condensation costs a few cents.
python3 otv.py custo <id> # total US$0.0023
Open the plan.json, remove or adjust a segment and render again. No model calls are made β manual cutting is free.
$EDITOR trabalho/<id>/plan.json python3 otv.py render <id>
The phases are independent: rescoring with another model does not redo the transcription or download.
python3 otv.py pontuar <id> --provedor gemini --forcar python3 otv.py selecionar <id> && python3 otv.py render <id>
Keeps only slides, demos, and charts. Requires visual classification by a model, so pass --visual.
python3 otv.py run "<url>" --modo B --visual glm # roteiro.md + narracao/*.wav
Each segment with a face is replaced by a generated image with a slow Ken Burns effect. The audio remains original, so the speech does not change.
python3 otv.py run "<url>" --modo A --visual glm --substituir gerado
A 1206 s English source about longevity and AI, condensed to 136,6 s in 18 segments for US$0,0023 in model costs. Rendering takes 21 seconds.


The pipeline is implemented and validated end to end. What comes next is scale and format.