PTENES
Skill /watch · Claude Code · Codex · claude.ai

Give Claude eyes to watch video

Paste a URL or a local file. Claude downloads it, extracts frames at scene cuts, transcribes it, and answers based on actually seen e heard the video.

/watch https://youtu.be/… "what's the hook?"
[watch] downloading via yt-dlp…
[watch] frames at scene cuts (1 per shot)…
[watch] examining the hook under a microscope for the first 10s…
[watch] transcribed via Groq whisper-large-v3

# watch: relatório do vídeo
- Frames:  47 @ scene-change
- Hook 0-10s:  pergunta direta + corte rápido
- Transcript:  via captions
- report.md pronto para ingest ✓
What it is

Claude has no video input. This skill gives it one.

A Python script downloads the video, turns it into JPEG frames (one per detected shot), gets a timestamped transcript (captions first, Whisper as fallback), and delivers it all to Claude Read read as an image. It responds by looking at what’s on screen and listening to what was said.

🎞️ Frames per scene cut

One frame per detected shot via ffmpeg select=gt(scene,…), not a fixed tick every N seconds. Token cost stays flat for long videos because the number of frames is limited by cuts, not duration.

🔬 Hook microscope 0-10s

Dense pass at 2 fps + word-level Whisper transcript for the first 10 seconds — where every video wins or loses attention. The report says what was on screen as each word landed.

📝 report.md + Obsidian

Fixed-schema report (TL;DR, key moments, hook, editorial profile with motion and camera movement, quotes, entities, concepts, transcript) with markers for Claude to fill in. Optional auto-save to your Obsidian vault.

How it works

From URL to answer, in one pipeline

Pure ffmpeg + yt-dlp + stdlib. Only the extracted audio goes over the network (Groq/OpenAI), and only when captions are missing and Whisper isn't disabled.

URL or file→ yt-dlp downloads→ ffmpeg: frames per scene→ pacing + motion + camera→ captions / Whisper→ Claude reads every frame→ answer + report.md→ Obsidian (optional)
End goal

Video becomes a connected note in your Obsidian

O /watch doesn’t stop at the answer. The goal is to turn the video into a structured knowledge node in your Second Brain — framed by why you watched (the --intent). It records in two layers.

📄 report.md (always)

Fixed-schema artifact in raw/watched/<slug>/ + the hero frames: TL;DR through the lens of intent, key moments, a 0–10s hook microscope, editorial profile, quotes, entities as [[wikilinks]], concepts, and the full transcript.

🧠 Wiki ingestion (with consent)

If the vault has one CLAUDE.md with an Ingest op, it runs and writes/updates wiki/entities/ (people, companies, tools), wiki/concepts/ (frameworks), wiki/sources/ (the video page) and a line in the log.md.

The unifying goal: video becomes a first-class node in your knowledge graph — people, tools, and concepts linked via [[wikilink]] to the notes you already have. Honest detail: the report.md the script generates it automatically; the wiki is only populated if your vault defines the Ingest op — the skill delegates this step to your Second Brain’s contract; it doesn’t come with one ready-made.

Prerequisites

Almost zero config to get started

First /watch the preflight checks everything. On macOS, it installs automatically via brew; on Linux/Windows, it prints the exact commands. Captions cover most public videos for free — the Whisper key is only needed when the video has no captions.

ffmpeg + yt-dlp

Download, frame extraction, and audio extraction. Installed on the first run.

# macOS (automatic)
brew install ffmpeg yt-dlp
# Linux
sudo apt install ffmpeg
pipx install yt-dlp

Whisper key (optional)

Only for videos without captions. Groq preferred (cheaper/faster) or OpenAI.

# ~/.config/watch/.env (chmod 600)
GROQ_API_KEY=…   # console.groq.com/keys
OPENAI_API_KEY=… # platform.openai.com
User guide · step by step

Install and watch

Real commands. In Claude Code, it's a plugin; in Codex/claude.ai, it's a skill. Once installed, the workflow is just /watch <url-ou-arquivo> [pergunta].

1

Install in Claude Code

Add the marketplace and install the plugin watch. (Two separate commands.)

/plugin marketplace add inematds/claude-watch
/plugin install watch@claude-watch
2

Preflight on the first run

Installs missing dependencies and creates ~/.config/watch/.env. Idempotent — safe to rerun.

python3 "$CLAUDE_PLUGIN_ROOT/scripts/setup.py"
3

Watch a video

URL (YouTube, Vimeo, TikTok, X… anything yt-dlp supports) or local path. The question becomes the report's intent.

/watch https://youtu.be/dQw4w9WgXcQ # what happens at 30s?
4

Focus on a section (denser, fewer tokens)

--start / --end in SS, MM:SS or HH:MM:SS. The transcript is filtered to the same window.

/watch "$URL" --start 2:15 --end 2:45
5

Auto-save to Obsidian (optional)

Point it to the vault, and report.md becomes a linked entry. Without this, the ingest step is silently skipped.

export WATCH_VAULT_DIR=/caminho/do/seu/vault
Examples

What people use it for

The same video, different lenses — the --intent shapes the TL;DR and report sections.

🪝 Analyze someone else’s content

"what hook did it open with?" — Claude looks at the first frames, reads the opening transcript, and breaks down the structure. Works for ad creative, podcast intros, and competitor launches.

🐞 Diagnose a bug from a video

They send over a broken screen recording. /watch bug.mov o que está errado? — it finds the frame where the problem appears and describes the cause.

⏩ Summarize a long video

/watch <url> resume isso — structure, key moments, what was said and shown. Faster than watching at 2x.

Roadmap

Versions and origin

Fork of claude-video by Bradley Bonanno (MIT) — the yt-dlp + ffmpeg + Whisper pipeline is his and runs unchanged.

v0.4.1
CurrentClassification of camera movement per shot (pan / tilt / zoom / static / handheld) via ffmpeg vidstabdetect, no opencv. Before: motion by shot (signalstats), frames at scene cuts, a 0–10s hook microscope, report.md, and auto-save to Obsidian.
Base
claude-video (Bradley Bonanno)yt-dlp download + captions, ffmpeg frames with auto-scaled fps, Groq/OpenAI Whisper backends, focused mode --start/--end and the SessionStart hook.
Next
More metrics and platformsTalking-head ratio (optional face detection), broader source coverage, and refinements to the ingest operation in Second Brain.