Paste a URL or a file and ask a question. The agent pulls captions, extracts frames, transcribes the audio, and reads each frame as an image. It responds as if it watched the video.
# paste the URL + the question /watch https://youtu.be/… o que acontece em 0:30? ▸ yt-dlp subtitles found (auto) ▸ ffmpeg 42 frames · scene-aware ▸ transcription captions · 0:00 → 3:12 ▸ Claude read 42 images + transcript → Em 0:30 aparece o menu de config, e a narração explica o passo 2…
The agent reads the page, runs the script, navigates the repo. What it not does out of the box is watch a video. This skill gives it that ability—pure-stdlib Python orchestrating yt-dlp + ffmpeg + an optional Whisper API.
Frames extracted as JPEGs + transcript with timestamps. The agent provides Read in each frame—the image goes straight into context and aligns with what was said.
YouTube, TikTok, Loom, X, Vimeo, Twitch URLs, and a few hundred other sites via yt-dlp — or a local file .mp4 .mov .mkv .webm.
A self-contained skill folder installs in Claude Code, Codex, Cursor, Copilot, Gemini CLI, and +50 Agent Skills. Zero config to get started.
The priority is to spend as little as possible: in transcript, a video with captions returns without downloading any video. When frames are needed, it downloads and extracts only what the run requests.
yt-dlp checks for captions first. If they exist (and in transcript mode), no video is downloaded.
ffmpeg extracts quick keyframes (efficient) or frames per scene cut (balanced), 512px, max 2 fps. In efficient, only pulls the keyframes that the video itself already stores every ~1s (-skip_frame nokey) — without re-decoding and comparing frame by frame, so it’s much faster than a full scene scan.
Each frame becomes a thumbnail and is compared with the last one kept; a pixel difference below the threshold discards the frame (static slide, static screen) before counting toward the budget — you only pay for frames that actually change.
Native captions (free) first; if unavailable, extracts mono 16 kHz audio and sends it to Whisper (Groq or OpenAI).
Captions cover most public videos for free. The Whisper key is only used when the video genuinely has no captions (local files, some TikToks/Vimeos).
Installed on the first run. On macOS, automatic via brew; Linux/Windows print the exact command. Preflight is a <100 ms lookup the next time.
# preflight + installer (idempotent) python3 scripts/setup.py
Install it in your tool, paste a URL, and ask a question. The rest — captions, download, frames, transcript — is handled by the script.
Adds the local marketplace and installs the plugin. Update later with /plugin update watch@claude-video.
/plugin marketplace add inematds/claude-video /plugin install watch@claude-video
The CLI for Agent Skills detects the hosts and copies the entire skill. -g installs globally; remove it for project scope.
npx skills add inematds/claude-video -g # global for your user
Source (URL or path) + your question. Without a question, it summarizes. If you don't ask anything, you get a structured summary with key moments.
/watch https://youtu.be/dQw4w9WgXcQ o que acontece no minuto 0:30? /watch ~/Movies/screen-recording.mp4 quando a UI quebra? /watch https://www.tiktok.com/@user/video/123 resume isso
When the question is about a moment, pass --start/--end (accepts SS, MM:SS, HH:MM:SS). The frame budget gets denser, and the transcript is filtered to the same interval.
/watch $URL --start 2:15 --end 2:45 # zoom at 30 s at 2 fps /watch video.mp4 --start 50 --end 60 # last 10 s /watch $URL --start 1:12:00 # from 1h12m to the end
Trades fidelity for speed and token cost. Default balanced. Set the default with WATCH_DETAIL= in the ~/.config/watch/.env.
| Mode | Frames | Usage |
|---|---|---|
transcript | 0 | Transcript only; skips the download when captions are available. |
efficient | up to 50 | Quick keyframes (~0.5 s to extract). |
balanced | up to 100 | Frames per scene cut. Default. |
token-burner | no limit | Scene-based, no cap — maximum coverage for long videos. |
/watch $URL --detail efficient # quick pass of 50 keyframes /watch $URL --resolution 1024 # read text on screen (slides, terminal)
Scene/keyframe selection can miss moments when someone points at the screen without cutting to a new scene—"look here," "as you can see," "notice this." The agent reads the transcript first, identifies these moments, and runs it again with --timestamps to force a frame right there. It’s the agent’s judgment, not a regex—and the clue frames are added to what the --detail had already chosen, without being discarded by the budget cutoff.
/watch $URL --detail transcript # 1. transcript with timestamps first /watch $URL --timestamps 4:32,7:10,9:55 # 2. rerun at moments when the speech points to the screen
The benefit shows up when the how matters as much as the what — video hooks, screen recording bugs, summaries of long content.
Look at the first frames + the transcript opening and break down the structure of a viral video, ad creative, or podcast intro.
/watch $URL qual foi o hook de abertura?
Got a screen recording of something broken? It finds the frame where the problem appears and describes it—often catching the cause.
/watch bug-repro.mov o que está dando errado?Pulls out the structure, key moments, and what was actually said and shown. Faster than watching at 2×.
/watch $URL resume isso
Boils a "game-changer" down to the few things that matter—substance without ten minutes of intro and overselling.
/watch $URL o que é NOVO de verdade — pula o hype
Self-contained, pure-stdlib skill. Every line below describes behavior that exists today—not a future roadmap.
transcript, a video with captions doesn’t even download the video.transcript, efficient (keyframes), balanced (scene) and token-burner (no cap). Frame budget based on duration to avoid exceeding the context.--no-dedup turns off.--start/--endDenser budget in the requested segment; transcript filtered to the same interval; timestamps always use the video's absolute time.--timestampsThe agent reads the transcript, identifies where the speaker points to the screen ("look here," "as you can see"), and forces a frame there — in addition to what scene selection already captured, so it isn't discarded by the budget cutoff.npx skills, and claude.ai per bundle .skill.