PTENES
CLI Β· Python Β· ffmpeg

30 minutes become 2 β€” and the model never chooses where to cut

O otv transcribes, splits the speech into numbered units, and asks the model only for a score from 0 to 10 for each ID. The code converts the ID to a timestamp and makes the cut β€” so it never lands in the middle of a word, and running it twice produces the same result.

Cover image for the otimizevideo project
What it is

A video cutter where every decision is auditable

A 20- to 30-minute lesson, podcast, or talk goes in; a output.mp4 about two minutes long with the best content. Each phase writes a JSON file that you can read, edit, and re-render at no cost.

🎯 The model scores, it doesn’t cut

It sees [042] 4.2s talking_head "texto" and returns a score. Timing comes from the word-level timestamped transcription, so the cut boundary is a real word boundary β€” never an invented timestamp.

πŸ” Deterministic and inexpensive to redo

The selection (backpack by score, quota per topic, hook and closing anchors, coherence) is code. Even notas.json, even plan.json. Editing the plan by hand and rendering again costs zero model calls.

πŸ”Œ Provider can be switched by phase

Transcription with Groq or local Whisper; scoring with GLM, Gemini, Ollama, or Claude Code itself; TTS with inemavox or ElevenLabs; image generation with flux-2-klein. One line in the config.yaml or a flag.

How it works

Nine phases, one JSON between each

Each phase reads the previous phase’s artifact and writes its own. Rerunning one phase does not redo the others β€” and existing files are reused unless you ask --forcar.

ingest→ transcribe→ scenes→ score→ units→ score→ select→ narrate→ render

Mode A β€” condensed

The default. Keeps the presenter and cuts only what works. Variant A+ (--substituir gerado) replaces presenter segments with generated illustrations and preserves the original audio.

Mode B β€” no presenter

Keeps only slides, screen demos, and charts; the model writes a script in Brazilian Portuguese, narrated with TTS, while the original audio becomes a bed at βˆ’18 dB.

Mode C β€” visuals only

Same selection as B, without the narration layer β€” for when the visual material already explains itself.

Prerequisites

What you need on your machine

No container or build. Python 3.12, ffmpeg, and the keys in .env as always.

ffmpeg and ffprobe

Handle all the cutting, audio mixing, and headline. Rendering runs inside a scope with a memory limit.

# Debian/Ubuntu
sudo apt install ffmpeg

Python dependencies

yt-dlp to download, PySceneDetect to find the cuts, mediapipe and OpenCV to detect faces.

# in the project root
pip install -r requirements.txt

API keys

Read at runtime from ~/projetos/openpcbotv2/.env e ~/projetos/wifi/.env. No keys are copied into the project.

# what is used
GROQ_API_KEY
OPENROUTER_API_KEY
FAL_KEY  # only in mode A+
User guide Β· step by step

From link to a two-minute cut

All commands below are real and were run during end-to-end validation. The <id> is the name of the folder created in trabalho/.

1

Run the entire pipeline

One line does it all: downloads, transcribes, detects scenes, scores, selects, and renders. The result is copied to ~/projetos/output/otimizevideo/<id>/.

python3 otv.py run "https://www.youtube.com/watch?v=..." --modo A --alvo 120
2

See what was selected

O status lists the artifacts, headline, and each segment with its timestamp, duration, and visual classification.

python3 otv.py status <id>  # plan: mode A Β· 136.6s across 18 segments
3

Check the cost

Each phase logs time and cost in trabalho/<id>/custos.json. A typical condensation costs a few cents.

python3 otv.py custo <id>  # total US$0.0023
4

Don’t like a segment? Edit the plan

Open the plan.json, remove or adjust a segment and render again. No model calls are made β€” manual cutting is free.

$EDITOR trabalho/<id>/plan.json
python3 otv.py render <id>
5

Swap the provider for a single phase

The phases are independent: rescoring with another model does not redo the transcription or download.

python3 otv.py pontuar <id> --provedor gemini --forcar
python3 otv.py selecionar <id> && python3 otv.py render <id>
6

Mode B β€” video with narration, no presenter

Keeps only slides, demos, and charts. Requires visual classification by a model, so pass --visual.

python3 otv.py run "<url>" --modo B --visual glm  # roteiro.md + narracao/*.wav
7

Mode A+ β€” presenter becomes an illustration

Each segment with a face is replaced by a generated image with a slow Ken Burns effect. The audio remains original, so the speech does not change.

python3 otv.py run "<url>" --modo A --visual glm --substituir gerado
Examples

Run on a real 20-minute video

A 1206 s English source about longevity and AI, condensed to 136,6 s in 18 segments for US$0,0023 in model costs. Rendering takes 21 seconds.

Frame from a segment replaced with an illustration in mode A+
Mode A+: the presenter segment became an AI-generated conceptual still life β€” the audio remains original.
Frame from a segment preserved from the original video
A segment preserved from the source: frame-accurate cuts, starting and ending at word boundaries.
Roadmap

Where the project is

The pipeline is implemented and validated end to end. What comes next is scale and format.

Ready
Complete pipeline, modes A, A+, B, and CNine phases, CLI with standalone phases, 110 tests, end-to-end validation documented in the spec.
Ready
Render robustnessOne ffmpeg input per segment and a memory cap β€” previously, a single filtergraph used up to 60,9 GB and froze the machine.
Next
9:16 output and batch processingVertical cuts for Reels and Shorts, and an entire playlist with one command.
Idea
Mode A+ with real b-rollReplace the presenter with stock footage instead of a static image.