You provide a topic or a link (page/video). The skill detects the type, decides on the best editing action, chooses a preset, and outputs plano-edicao.json + RESUMO.md — structured data, renderer-agnostic. Optional rendering already works: motion graphics + flux2-klein b-roll, local, no API key.
# from input to plan input → detect → analyze → best action → preset → PLANO (json + resumo) → (opcional) render → MP4 # quick start vpe scaffold "como fazer pão caseiro" --preset viral \ > plano-edicao.json vpe validate plano-edicao.json # guardrails vpe resumo plano-edicao.json # to approve
Programmatic editing is fragmented (FFmpeg, Auto-Editor, MoviePy, Remotion, HyperFrames). What’s missing is a layer that decides what to do from natural language and outputs a clean, portable plan. That’s what this is.
Classifies the input as topic, page or video, analyzes the best editing action and records why in best_action_rationale.
EditPlan in JSON (pydantic) + RESUMO.md readable. It can be consumed by FFmpeg, HyperFrames, Remotion, or other systems.
Motion graphics via HyperFrames + b-roll generated with flux2-klein. Local, deterministic, no API key.
5 style presets (action · smooth · promo · sales · viral) are parameter profiles for same schema — changing the preset changes pacing, transitions, soundtrack, captions, and aspect ratio without rewriting the base plan.
The core (plan generation) is just Python. Rendering is optional and requires the local motion graphics stack.
Install in editable mode; exposes the console script vpe.
python3 -m venv .venv && . .venv/bin/activate pip install -e ".[dev]"
The skill orchestrates vpe and fills the creative timeline using natural language.
# https://claude.com/claude-codeFor rendering only: HyperFrames (Chrome+FFmpeg), Kokoro (voiceover), and flux2-klein (b-roll).
# see knowledge/render.mdReal commands for vpe. The skill (via SKILL.md) fills the timeline between the scaffold and validation.
Creates the environment and exposes the command vpe.
pip install -e ".[dev]" # exposes `vpe`
vpe detects the input (topic/page/video) and applies the selected preset.
vpe scaffold "como fazer pão caseiro" --preset viral > plano-edicao.json
Script/scenes for a topic or page; source_in/out (real cuts) for video. See examples/.
# timeline: hook → point → ... → cta (each beat with headline/narration/caption)
Reset the [error] before any rendering (empty timeline, no hook, etc.).
vpe validate plano-edicao.json # exit 0 = ok
Readable document with source, objective, preset, and a beat-by-beat timeline. Rendering is optional, on request.
vpe resumo plano-edicao.json > RESUMO.md
Complete sample plan in examples/exemplo-viral-pao.json.
// vpe scaffold "https://youtu.be/..." --preset viral "source": { "kind": "video", ... }, "style": { "preset": "viral", "pacing": { "avg_cut_seconds": 1.1, "energy": 0.95 }, "captions": { "style": "karaoke" } }, "output": { "aspect": "9:16", "duration_target_seconds": 30 }
YouTube input classified as video; preset applied ~1.1s cuts, karaoke captions, and 9:16. The timeline starts empty.
# Homemade bread in 20s - Objetivo: viral - Preset: viral (corte ~1.1s, energia 0.95) - Saída: 9:16, alvo 20s 1. hook Headline: Pão em 20 segundos?! 2. point Headline: 3 ingredientes 3. cta Headline: INEMA.CLUB
The same plan becomes readable Markdown with a beat-by-beat timeline.
Plan schema: meta · source · intent · style · output · timeline[] · render. Integrates with MDD for the cinematic direction of generative beats.
vpe. 35 tests..otio for NLEs.