Paste a URL or a local file. Claude downloads it, extracts frames at scene cuts, transcribes it, and answers based on actually seen e heard the video.
/watch https://youtu.be/… "what's the hook?" [watch] downloading via yt-dlp… [watch] frames at scene cuts (1 per shot)… [watch] examining the hook under a microscope for the first 10s… [watch] transcribed via Groq whisper-large-v3 # watch: relatório do vídeo - Frames: 47 @ scene-change - Hook 0-10s: pergunta direta + corte rápido - Transcript: via captions - report.md pronto para ingest ✓
A Python script downloads the video, turns it into JPEG frames (one per detected shot), gets a timestamped transcript (captions first, Whisper as fallback), and delivers it all to Claude Read read as an image. It responds by looking at what’s on screen and listening to what was said.
One frame per detected shot via ffmpeg select=gt(scene,…), not a fixed tick every N seconds. Token cost stays flat for long videos because the number of frames is limited by cuts, not duration.
Dense pass at 2 fps + word-level Whisper transcript for the first 10 seconds — where every video wins or loses attention. The report says what was on screen as each word landed.
Fixed-schema report (TL;DR, key moments, hook, editorial profile with motion and camera movement, quotes, entities, concepts, transcript) with markers for Claude to fill in. Optional auto-save to your Obsidian vault.
Pure ffmpeg + yt-dlp + stdlib. Only the extracted audio goes over the network (Groq/OpenAI), and only when captions are missing and Whisper isn't disabled.
O /watch doesn’t stop at the answer. The goal is to turn the video into a structured knowledge node in your Second Brain — framed by why you watched (the --intent). It records in two layers.
Fixed-schema artifact in raw/watched/<slug>/ + the hero frames: TL;DR through the lens of intent, key moments, a 0–10s hook microscope, editorial profile, quotes, entities as [[wikilinks]], concepts, and the full transcript.
If the vault has one CLAUDE.md with an Ingest op, it runs and writes/updates wiki/entities/ (people, companies, tools), wiki/concepts/ (frameworks), wiki/sources/ (the video page) and a line in the log.md.
The unifying goal: video becomes a first-class node in your knowledge graph — people, tools, and concepts linked via [[wikilink]] to the notes you already have. Honest detail: the report.md the script generates it automatically; the wiki is only populated if your vault defines the Ingest op — the skill delegates this step to your Second Brain’s contract; it doesn’t come with one ready-made.
First /watch the preflight checks everything. On macOS, it installs automatically via brew; on Linux/Windows, it prints the exact commands. Captions cover most public videos for free — the Whisper key is only needed when the video has no captions.
Download, frame extraction, and audio extraction. Installed on the first run.
# macOS (automatic) brew install ffmpeg yt-dlp # Linux sudo apt install ffmpeg pipx install yt-dlp
Only for videos without captions. Groq preferred (cheaper/faster) or OpenAI.
# ~/.config/watch/.env (chmod 600) GROQ_API_KEY=… # console.groq.com/keys OPENAI_API_KEY=… # platform.openai.com
Real commands. In Claude Code, it's a plugin; in Codex/claude.ai, it's a skill. Once installed, the workflow is just /watch <url-ou-arquivo> [pergunta].
Add the marketplace and install the plugin watch. (Two separate commands.)
/plugin marketplace add inematds/claude-watch /plugin install watch@claude-watch
Installs missing dependencies and creates ~/.config/watch/.env. Idempotent — safe to rerun.
python3 "$CLAUDE_PLUGIN_ROOT/scripts/setup.py"
URL (YouTube, Vimeo, TikTok, X… anything yt-dlp supports) or local path. The question becomes the report's intent.
/watch https://youtu.be/dQw4w9WgXcQ # what happens at 30s?
--start / --end in SS, MM:SS or HH:MM:SS. The transcript is filtered to the same window.
/watch "$URL" --start 2:15 --end 2:45
Point it to the vault, and report.md becomes a linked entry. Without this, the ingest step is silently skipped.
export WATCH_VAULT_DIR=/caminho/do/seu/vault
The same video, different lenses — the --intent shapes the TL;DR and report sections.
"what hook did it open with?" — Claude looks at the first frames, reads the opening transcript, and breaks down the structure. Works for ad creative, podcast intros, and competitor launches.
They send over a broken screen recording. /watch bug.mov o que está errado? — it finds the frame where the problem appears and describes the cause.
/watch <url> resume isso — structure, key moments, what was said and shown. Faster than watching at 2x.
Fork of claude-video by Bradley Bonanno (MIT) — the yt-dlp + ffmpeg + Whisper pipeline is his and runs unchanged.
vidstabdetect, no opencv. Before: motion by shot (signalstats), frames at scene cuts, a 0–10s hook microscope, report.md, and auto-save to Obsidian.--start/--end and the SessionStart hook.