Sandbox
@jordanrendric/claude-video-vision

Claude Code plugin for video frames and audio

This plugin lets Claude Code watch videos by extracting frames with ffmpeg and sending audio through Gemini, Whisper, or OpenAI. It also supports YouTube URLs, setup guidance, and adaptive frame extraction based on what you ask.

1,295 stars153 forksTypeScriptUpdated 1mo ago
Who it's for

Builders who use Claude Code and want video files or YouTube links analyzed inside the agent.

What it delivers

You can ask Claude to inspect a video and get both visual frames and timestamped audio context.

What it does

Frame extraction

Uses ffmpeg to pull video frames so Claude can inspect the visuals directly.

Audio backends

Supports Gemini API, local Whisper, or OpenAI Whisper for transcription and audio events.

YouTube support

Accepts YouTube URLs and downloads them with yt-dlp while keeping source metadata and captions.

Adaptive extraction

Adjusts frame rate, time range, and resolution based on the question you ask.

Interactive setup

Includes a `/setup-video-vision` command that walks through backend and dependency setup.

MCP tools

Exposes tools like `video_watch`, `video_detail`, `video_info`, and `video_setup` for Claude Code.

How to get it

  1. 1Inside Claude Code, run these commands one at a time
    /plugin marketplace add https://github.com/jordanrendric/claude-video-vision
  2. 2Then
    /plugin install claude-video-vision
  3. 3Alternative: local development
    git clone https://github.com/jordanrendric/claude-video-vision.git
    claude --plugin-dir /path/to/claude-video-vision
  4. 4Inside Claude Code, run the interactive wizard
    /claude-video-vision:setup-video-vision
  5. 5Run
    /watch-video path/to/video.mp4
    /watch-video tutorial.mp4 "what language is used in this tutorial?"
    /watch-video https://www.youtube.com/watch?v=... "summarize this video"

README

claude-video-vision

Claude Code Video Vision

Give Claude the ability to watch and understand videos.

A Claude Code plugin that extracts frames via ffmpeg and processes audio via multiple backends (Gemini API, local Whisper, or OpenAI API). Claude receives frames as images and audio transcription with timestamps — the plugin is a perception layer, not an interpretation layer.

Features

  • Multimodal perception — Claude sees video frames directly and reads audio transcriptions with timestamps
  • YouTube URL support — Pass a YouTube URL directly; the MCP server downloads it with yt-dlp, preserving source metadata and captions for context
  • Flexible backends — Choose between cloud APIs or fully local processing
  • Adaptive extraction — Claude adjusts fps, time range, and resolution based on your question
  • Auto-installation — Whisper models download automatically on first use
  • Interactive setup wizard/setup-video-vision walks you through configuration

Quick Start

1. Install the plugin

Inside Claude Code, run these commands one at a time:

/plugin marketplace add https://github.com/jordanrendric/claude-video-vision

Then:

/plugin install claude-video-vision

The MCP server will auto-install via npx from npm on first use — no build step required.

Alternative: local development

git clone https://github.com/jordanrendric/claude-video-vision.git
claude --plugin-dir /path/to/claude-video-vision

2. Configure

Inside Claude Code, run the interactive wizard:

/claude-video-vision:setup-video-vision

It will walk you through backend selection, whisper configuration (if local), frame options, and dependency verification.

Usage

Slash command

/watch-video path/to/video.mp4
/watch-video tutorial.mp4 "what language is used in this tutorial?"
/watch-video https://www.youtube.com/watch?v=... "summarize this video"

Conversational

Just mention a video file or YouTube URL — Claude will detect it:

"analyze this video for me: ~/Downloads/demo.mp4"

"take a look at the first second of ~/videos/bug-report.mov"

"summarize this YouTube Short: https://www.youtube.com/shorts/..."

Claude adapts parameters automatically:

  • "the first second" → extracts at original fps from 00:00:00 to 00:00:01
  • "summarize this 1h lecture" → low fps, full duration
  • "what text is on screen at 1:30?" → high resolution, narrow time window

Backends

BackendAudio processingCostSetup
Gemini APINative (speech + non-speech events)Free tier: 1500 req/dayGEMINI_API_KEY env var
Local (Whisper)whisper.cpp or Python openai-whisperFree, fully offlinebrew install whisper-cpp + auto model download
OpenAI APIOpenAI Whisper APIPaid per usageOPENAI_API_KEY env var

All backends extract video frames via ffmpeg — Claude always has direct visual access.

Architecture

┌───────────────────────────────────────────────────────┐
│ Claude Code (your session)                            │
│                                                       │
│  /watch-video  ──→  Skill: video-perception          │
│                        │                              │
│                        ▼                              │
│                  MCP tool: video_watch                │
│                        │                              │
└────────────────────────┼──────────────────────────────┘
                         │
                         ▼
      ┌────────────────────────────────────┐
      │ MCP Server (Node.js)               │
      │                                    │
      │  ┌──────────┐    ┌──────────────┐  │
      │  │ ffmpeg   │    │ Audio backend│  │
      │  │ frames   │ ║  │ (parallel)   │  │
      │  └──────────┘    └──────────────┘  │
      │       │                 │          │
      └───────┼─────────────────┼──────────┘
              ▼                 ▼
        base64 images     transcription
        + timestamps      + audio events
              │                 │
              └────────┬────────┘
                       ▼
              Claude receives both

Requirements

  • Node.js 20+ (for the MCP server)
  • ffmpeg (auto-detected, install instructions provided by setup wizard)
  • yt-dlp (optional, required only for YouTube URLs; brew install yt-dlp on macOS)
  • Backend-specific:
    • Gemini API: free API key from ai.google.dev
    • Local: brew install whisper-cpp (macOS) or equivalent
    • OpenAI: API key from OpenAI

MCP Tools

The plugin exposes 6 MCP tools:

  • video_watch — Extract frames + process audio (main tool)
  • video_analyze — Analyze video structure with ffmpeg filters before extraction
  • video_detail — Drill into specific cached or newly extracted moments
  • video_info — Get video metadata without processing
  • video_configure — Change settings
  • video_setup — Check and guide dependency installation

Slash Commands

  • /watch-video <path> [question] — Analyze a video
  • /setup-video-vision — Interactive configuration wizard

Configuration

Settings are stored in ~/.claude-video-vision/config.json:

{
  "backend": "local",
  "whisper_engine": "cpp",
  "whisper_model": "auto",
  "whisper_at": false,
  "frame_mode": "images",
  "frame_format": "jpeg",
  "frame_resolution": 512,
  "default_fps": "auto",
  "max_frames": 100,
  "frame_describer_model": "sonnet",
  "enable_index": false,
  "session_max_age_days": 7,
  "downloads_max_age_days": 7
}

frame_format can be jpeg, png, or webp. jpeg remains the default for backwards compatibility; png is useful for screen recordings where text and sharp UI edges should stay lossless.

Whisper models auto-download to ~/.claude-video-vision/models/ on first use. Available: tiny, base, small, medium, large-v3-turbo, large-v3, auto (picks best for your RAM).

YouTube Transcripts

For YouTube URLs, the server uses this transcript order:

  1. Manual YouTube subtitles when an English track is available.
  2. YouTube automatic captions when manual subtitles are not available.
  3. The configured audio backend when captions are missing, empty, or cover too little of a longer video.

Audio results label provenance with transcription_source, for example youtube_subtitles or youtube_auto_captions, so Claude can treat manual subtitles as stronger evidence than auto-captions.

Status

v1.0.0 — Initial release. Tested on macOS (Apple Silicon) with Local backend (whisper.cpp).

License

MIT — see LICENSE.

Author

Jordan Vasconcelos

Star History

Star History ChartStar History Chart

Files in the repo

Repository payload17 top-level entries
  • .claude-plugin
  • .github
  • agents
  • assets
  • commands
  • docs
  • mcp-server
  • skills
  • .gitignore
  • .mcp.json
  • CHANGELOG.md
  • CODE_OF_CONDUCT.md
  • CONTRIBUTING.md
  • LICENSE
  • PRIVACY.md
  • README.md
  • SECURITY.md

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More plugins

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

138k
1 add

Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.

82k
code-yeongyu/
oh-my-openagent

OmO: Just type "mass ulw" keyword with your prompt. Now you are the master of graph engineering.

69k

Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

94k

Opinionated Oxlint rules for rejecting low-evidence TypeScript and JavaScript patterns

4.3k