PTENES
Agent Skill · /watch

Your agent watches the video—not guessing from the title.

Paste a URL or a file and ask a question. The agent pulls captions, extracts frames, transcribes the audio, and reads each frame as an image. It responds as if it watched the video.

# paste the URL + the question
/watch https://youtu.be/… o que acontece em 0:30?

▸ yt-dlp   subtitles found (auto)
▸ ffmpeg   42 frames · scene-aware
▸ transcription   captions · 0:00 → 3:12
▸ Claude read 42 images + transcript

→ Em 0:30 aparece o menu de config,
   e a narração explica o passo 2…
What it is

A video input for your agent

The agent reads the page, runs the script, navigates the repo. What it not does out of the box is watch a video. This skill gives it that ability—pure-stdlib Python orchestrating yt-dlp + ffmpeg + an optional Whisper API.

👁️ See and hear

Frames extracted as JPEGs + transcript with timestamps. The agent provides Read in each frame—the image goes straight into context and aligns with what was said.

🌐 Any source

YouTube, TikTok, Loom, X, Vimeo, Twitch URLs, and a few hundred other sites via yt-dlp — or a local file .mp4 .mov .mkv .webm.

🧩 Multi-host

A self-contained skill folder installs in Claude Code, Codex, Cursor, Copilot, Gemini CLI, and +50 Agent Skills. Zero config to get started.

How it works

Captions first, download only what's needed

The priority is to spend as little as possible: in transcript, a video with captions returns without downloading any video. When frames are needed, it downloads and extracts only what the run requests.

URL / file + question→ yt-dlp · subtitles→ ffmpeg · frames→ dedup→ transcript→ Claude Read→ answer + timestamps→ cleanup

1 · Download

yt-dlp checks for captions first. If they exist (and in transcript mode), no video is downloaded.

2 · Frames

ffmpeg extracts quick keyframes (efficient) or frames per scene cut (balanced), 512px, max 2 fps. In efficient, only pulls the keyframes that the video itself already stores every ~1s (-skip_frame nokey) — without re-decoding and comparing frame by frame, so it’s much faster than a full scene scan.

3 · Dedup

Each frame becomes a thumbnail and is compared with the last one kept; a pixel difference below the threshold discards the frame (static slide, static screen) before counting toward the budget — you only pay for frames that actually change.

4 · Transcript

Native captions (free) first; if unavailable, extracts mono 16 kHz audio and sends it to Whisper (Groq or OpenAI).

Prerequisites

Two dependencies — and one optional key

Captions cover most public videos for free. The Whisper key is only used when the video genuinely has no captions (local files, some TikToks/Vimeos).

⚙️ yt-dlp + ffmpeg

Installed on the first run. On macOS, automatic via brew; Linux/Windows print the exact command. Preflight is a <100 ms lookup the next time.

# preflight + installer (idempotent)
python3 scripts/setup.py

🔑 Whisper key (optional)

Groq (preferred — cheaper and faster, whisper-large-v3) or OpenAI (whisper-1). Without a key, videos without captions return frames only.

# ~/.config/watch/.env  (mode 0600)
GROQ_API_KEY=...
# or
OPENAI_API_KEY=...
User guide · step by step

From install to your first answer

Install it in your tool, paste a URL, and ask a question. The rest — captions, download, frames, transcript — is handled by the script.

1

Install in Claude Code

Adds the local marketplace and installs the plugin. Update later with /plugin update watch@claude-video.

/plugin marketplace add inematds/claude-video
/plugin install watch@claude-video
2

…or in Codex, Cursor, Copilot, Gemini CLI (+50)

The CLI for Agent Skills detects the hosts and copies the entire skill. -g installs globally; remove it for project scope.

npx skills add inematds/claude-video -g  # global for your user
3

Watch

Source (URL or path) + your question. Without a question, it summarizes. If you don't ask anything, you get a structured summary with key moments.

/watch https://youtu.be/dQw4w9WgXcQ o que acontece no minuto 0:30?
/watch ~/Movies/screen-recording.mp4 quando a UI quebra?
/watch https://www.tiktok.com/@user/video/123 resume isso
4

Focus on a segment (more frames, fewer tokens)

When the question is about a moment, pass --start/--end (accepts SS, MM:SS, HH:MM:SS). The frame budget gets denser, and the transcript is filtered to the same interval.

/watch $URL --start 2:15 --end 2:45   # zoom at 30 s at 2 fps
/watch video.mp4 --start 50 --end 60   # last 10 s
/watch $URL --start 1:12:00           # from 1h12m to the end
5

Adjust the detail (one dial)

Trades fidelity for speed and token cost. Default balanced. Set the default with WATCH_DETAIL= in the ~/.config/watch/.env.

ModeFramesUsage
transcript0Transcript only; skips the download when captions are available.
efficientup to 50Quick keyframes (~0.5 s to extract).
balancedup to 100Frames per scene cut. Default.
token-burnerno limitScene-based, no cap — maximum coverage for long videos.
/watch $URL --detail efficient          # quick pass of 50 keyframes
/watch $URL --resolution 1024           # read text on screen (slides, terminal)
6

Capture the frame the speaker points to ("look here")

Scene/keyframe selection can miss moments when someone points at the screen without cutting to a new scene—"look here," "as you can see," "notice this." The agent reads the transcript first, identifies these moments, and runs it again with --timestamps to force a frame right there. It’s the agent’s judgment, not a regex—and the clue frames are added to what the --detail had already chosen, without being discarded by the budget cutoff.

/watch $URL --detail transcript         # 1. transcript with timestamps first
/watch $URL --timestamps 4:32,7:10,9:55 # 2. rerun at moments when the speech points to the screen
Examples

What people actually use it for

The benefit shows up when the how matters as much as the what — video hooks, screen recording bugs, summaries of long content.

🎯 Analyze someone else’s content

Look at the first frames + the transcript opening and break down the structure of a viral video, ad creative, or podcast intro.

/watch $URL qual foi o hook de abertura?

🐞 Diagnose a bug

Got a screen recording of something broken? It finds the frame where the problem appears and describes it—often catching the cause.

/watch bug-repro.mov o que está dando errado?

📝 Summarize a video

Pulls out the structure, key moments, and what was actually said and shown. Faster than watching at 2×.

/watch $URL resume isso

✂️ Cut through the hype of an update

Boils a "game-changer" down to the few things that matter—substance without ten minutes of intro and overselling.

/watch $URL o que é NOVO de verdade — pula o hype
Features · v0.2.0

What's already live

Self-contained, pure-stdlib skill. Every line below describes behavior that exists today—not a future roadmap.

Captions
Free transcript firstyt-dlp pulls manual or automatic subtitles from the source itself. In mode transcript, a video with captions doesn’t even download the video.
Whisper
Fallback with auto-chunkingIf there are no captions, extracts mono 16 kHz audio and transcribes it with Groq (preferred) or OpenAI. Audio over 25 MB is automatically split and reassembled.
Frames
4-mode dial + 2 fps maxtranscript, efficient (keyframes), balanced (scene) and token-burner (no cap). Frame budget based on duration to avoid exceeding the context.
Dedup
Collapses near-identical framesBrightness-difference pass against the last retained frame; a paused slide or static screen doesn't use up the budget. --no-dedup turns off.
Focus
Window with --start/--endDenser budget in the requested segment; transcript filtered to the same interval; timestamps always use the video's absolute time.
Cues
Frames per deduplication cue with --timestampsThe agent reads the transcript, identifies where the speaker points to the screen ("look here," "as you can see"), and forces a frame there — in addition to what scene selection already captured, so it isn't discarded by the budget cutoff.
Multi-host
One skill, +50 hostsClaude Code (plugin/marketplace), Codex, Cursor, Copilot, Gemini CLI via npx skills, and claude.ai per bundle .skill.