Sandbox
26 repos for image · Any agent · DocsClear
lansespirit/
image-gen-mcp

An MCP server that integrates with gpt-image-1 & Gemini imagen4 model for text-to-image generation services

69
wuyoscar/
GPT-Image2-Skill

GPT Image 2/2.5 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing

5.3k
shinpr/
mcp-image

MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream.

161

Use your ChatGPT subscription to generate images from the command line — no OPENAI_API_KEY, no gateway, no daemon. Zero-dep Python CLI + AI-agent skill.

350
tonkotsuboy/
github-upload-image-to-pr

AI agent skill(e.g., Claude Code, Codex): Upload local images to a GitHub PR and embed them in the description or comments

38
wanshuiyin/
ARIS-Movie-Director

Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generation (intelligence lives in the agent; the diffusion model just renders). Image-based today, video next.

60
ShunmeiCho/
cc-clip

Paste images into remote Claude Code & Codex CLI over SSH — clipboard bridging for macOS and Windows.

155
I-CAN-hack/
pdf-mcp

PDF MCP server with image rendering capabilities. Useful for automatically searching datasheets, manuals, etc...

77
JKc66/
custom-icons-skill

skill to create custom icons using IDEs or extentions that have image generation support

35
SkyworkAI/
Skywork-Skills

Skywork Agent Skills for AI office suites, including AI PPT, AI Document, AI Excel, AI Image, AI Search/DeepResearch and AI Music. These skills can be used by any skills-compatible agent, including Claude Code, Codex CLI and OpenClaw.

204
AeternaLabsHQ/
pullmd

Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.

480

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.

49

Desktop AI Assistant powered by GPT-6, GPT-5, o1, o3, Gemini, Claude, Ollama, DeepSeek, Perplexity, Grok, Bielik, chat, vision, voice, RAG, image and video generation, agents, tools, MCP, plugins, speech synthesis and recognition, web search, memory, presets, assistants,and more. Linux, Windows, Mac

1.9k
jztan/pdf-mcpConnectors

An MCP server that gives your AI agent agentic RAG over your PDFs, one file or a whole folder: hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts. The agent decides when to search; pdf-mcp does the retrieval.

133
OpenSenseNova/
SenseNova-Skills

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

5.5k

An Agent Skill and Dify plugin to transform Markdown to files of DOCX, PPTX, XLSX, PNG, PDF, HTML, MD, CSV, JSON, XML.

268
jinzcdev/
markmap-mcp-server

An MCP server for converting Markdown to interactive mind maps with export support (PNG/JPG/SVG).

285

Windows 11 shell extension (Rust): Explorer thumbnails for 331 file types Windows can't show, including camera RAW, PSD, HEIC/AVIF, video, ebooks, comics and CAD. Crash-isolated, source-available rebuild of SageThumbs.

174

Open source implementation and extension of Google Research’s PaperBanana for automated academic figures, diagrams, and research visuals, expanded to new domains like slide generation.

2.3k