An MCP server that integrates with gpt-image-1 & Gemini imagen4 model for text-to-image generation services
GPT Image 2/2.5 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing
Generate GPT images from Codex or Claude Code using a ChatGPT subscription, without the Images API.
A Cli, a webUI, and a MCP server for the Z-Image-Turbo text-to-image generation model (Tongyi-MAI/Z-Image-Turbo base model as well as quantized models)
Local MCP server for ChatGPT image generation.
Turn slide screenshots and generated images into editable PowerPoint decks with visual-layer splitting, OCR evidence, and QA.
AI agent skill(e.g., Claude Code, Codex): Upload local images to a GitHub PR and embed them in the description or comments

A powerful OCR extension with area selection tool and more
Connect Claude to image generation with Agent Skills. Three levels: a zero-cost code-based design engine, a Three.js 3D renderer, and a real diffusion model on Cloudflare. Plus an AI Storybook pipeline that turns a plain-English story into an illustrated, narrated HTML book.
An advanced in-memory image visualization plugin for GDB and LLDB on Linux, with experimental support for MacOS and Windows. Previously known as gdb-imagewatch. Also available as an extension for VSCode and forks
Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal generation (intelligence lives in the agent; the diffusion model just renders). Image-based today, video next.
MCP server for OpenRouter — chat with 300+ LLMs (Claude, Gemini, GPT), analyze images / audio / video, generate images / speech / music / video (Veo 3.1, Sora, Seedance, Wan) from Claude Desktop, Cursor, Kiro, VS Code.
Agnes AI skill for text, image, and video APIs with persistent auth and OpenAI-style workflows

Official agent skills from Black Forest Labs for FLUX image and video generation — prompting guides and API integration patterns for Claude Code, Codex, and any agentskills.io-compatible agent.
🔎 A MCP server for Unsplash image search.
Use ChatGPT Web (including Pro) as a native model in Codex — with context, tools, streaming and images, without using Codex quota.
Paste images into remote Claude Code & Codex CLI over SSH — clipboard bridging for macOS and Windows.
Agent skill from Y Build for turning AI images and videos into playable game art assets
PDF MCP server with image rendering capabilities. Useful for automatically searching datasheets, manuals, etc...
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
Local MCP server that lets AI assistants automate Origin/OriginPro for data and image processing.
Wonda CLI — AI-powered content creation from your terminal
Agent skills for healthcare and life sciences: genomics, imaging, claims, drug discovery, and more. Works with Amazon Quick, Kiro, Amazon AgentCore, AWS Strands SDK, Claude Code, Codex, and any Agent Skills-compatible platform.