Sandbox
7 repos for multimodal · DesignClear

MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

86
wanshuiyin/
ARIS-Movie-Director

Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generation (intelligence lives in the agent; the diffusion model just renders). Image-based today, video next.

60

Wan 3.0 API Python SDK and MCP server for AI video generation: text-to-video, image-to-video, multimodal references, uploads, and async job polling.

78
hufeng173/
kunpeng-skill

Kunpeng-Skill is a powerful multimodal distillation toolkit that transforms high-value insights from repositories, websites, UIs, videos, images, and documents into reusable methodologies, model-agnostic regeneration specifications, and more.

54
jomeswang/
agnes-ai-skill

Agnes AI skill for text, image, and video APIs with persistent auth and OpenAI-style workflows

59

100+ curated Seedance 2.5 prompts with real video previews, plus an installable Agent Skill that optimizes prompts, creates storyboards, and generates videos via Seedance models.

193

Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.

4.3k