Sandbox
8 repos for multimodal · DocsClear

SimpleMem: Efficient Lifelong Memory for LLM Agents — Text & Multimodal

3.8k
yuezhiai/
jonex
yuezhiai/jonexConnectors

All-in-One Multimodal Parsing Engine + Ontology-Powered, LLM Wiki-Driven AI-Ready Knowledge Engine

1k

Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis

1.3k
ggozad/
haiku.rag

Agentic RAG for local and self-hosted document search: hybrid retrieval, reranking and multimodal RAG on embedded LanceDB, with Docling parsing and an MCP server

606
wanshuiyin/
ARIS-Movie-Director

Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generation (intelligence lives in the agent; the diffusion model just renders). Image-based today, video next.

60
OpenBMB/UltraRAGFrameworks & SDKs

A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines

5.7k

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.

49