Sandbox
29 repos for multimodalClear

MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.

86

SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal

3.8k
yuezhiai/
jonex
yuezhiai/jonexConnectors

All-in-One Multimodal Parsing Engine + Ontology-Powered, LLM Wiki-Driven AI-Ready Knowledge Engine

1k

Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis

1.3k

The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra

39k
waybarrios/
vllm-mlx

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

1.6k
QinghongLin/
data2story-skill

Data Journalist Agent: Transforming Data into Verifiable Multimodal Story

155

Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.

328
ggozad/
haiku.rag

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

606
NPC-Worldwide/npcpyFrameworks & SDKs

The python library for research and development in NLP, multimodal LLMs, Agents, ML, Knowledge Graphs, and more.

1.5k
wanshuiyin/
ARIS-Movie-Director

Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generation (intelligence lives in the agent; the diffusion model just renders). Image-based today, video next.

60

Wan 3.0 API Python SDK and MCP server for AI video generation: text-to-video, image-to-video, multimodal references, uploads, and async job polling.

78
huangjunsen0406/
py-xiaozhi

Open-source AI assistant ecosystem with MCP integrations, multimodal workflows, IoT support, and cross-platform voice interaction.

3.5k
hufeng173/
kunpeng-skill

Kunpeng-Skill is a powerful multimodal distillation toolkit that transforms high-value insights from repositories, websites, UIs, videos, images, and documents into reusable methodologies, model-agnostic regeneration specifications, and more.

54

127 portable ecommerce skills for marketplace research, product discovery, keyword analysis, trends, sourcing, patent screening, multimodal tasks, and GEO workflows.

67
deepset-ai/haystackFrameworks & SDKs

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.

26k

Unreal Engine plugin for LLM/GenAI models & MCP UE5 server. OpenAI GPT-5, Deepseek R1, Claude Opus/Sonnet, Gemini 3, Grok 4, Alibaba Qwen, Kimi, ElevenLabs TTS, Inworld, OpenRouter, Groq, GLM, Ollama, Local, Meshy, Tripo, Hunyuan3D, Rodin, fal, Dashscope, Seedream. NPC AI, agentic, chat, 3D gen, TTS, multimodal, image gen. UnrealMCP/UnrealClaude

644
OpenBMB/UltraRAGFrameworks & SDKs

A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines

5.7k
datachain-ai/datachainFrameworks & SDKs

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

2.8k
Sumanth077/
Hands-On-AI-Engineering

A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.

3.4k
jomeswang/
agnes-ai-skill

Agnes AI skill for text, image, and video APIs with persistent auth and OpenAI-style workflows

59

YC (S26) | Open Computer History | Record your screen continuously locally and provide context to your agents (Claude, Codex, Openclaw, Hermes, Runner...)

22k

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.

49