Sandbox
@lownamlee/gpt-image-2-mcp

MCP server for ChatGPT image generation

This server plugs into an MCP client and exposes a small tool surface for generating images. It accepts a prompt, chooses a backend, writes the results to a local output folder, and returns file paths and metadata to the caller.

92 stars2 forksTypeScriptUpdated 2mo ago
Who it's for

Builders who want their agent to generate and save images from an MCP client.

What it delivers

You can ask for images inside your agent workflow and get saved files back instead of juggling a separate app.

What it does

Image generation tool

`generate_image(prompt, backend?, n?, size?, quality?, output_format?, conversation_mode?, timeout_seconds?)` creates images from a prompt.

Backend selection

Supports `api`, `chatgpt-web`, and `auto`, with fallback from API to ChatGPT website mode when needed.

Saved output files

Writes each generation to a prompt-derived directory with numbered image files and a `metadata.json` file.

Session and visibility control

`browser_visibility(action?, start_browser?)` can start the ChatGPT session and control whether the browser window stays visible.

Backend readiness check

`backend_status(backend?)` reports whether a backend is ready and shows the effective output root.

How to get it

  1. 1Default output roots
    Windows: %LOCALAPPDATA%\gpt-image-2-mcp\output\chatgpt-images
    macOS:   ~/Library/Application Support/gpt-image-2-mcp/output/chatgpt-images
    Linux:   ${XDG_DATA_HOME:-~/.local/share}/gpt-image-2-mcp/output/chatgpt-images
  2. 2Run the server in ChatGPT website mode
    $env:GPT_IMAGE_BACKEND = "chatgpt-web"
    node dist/index.js
  3. 3The local ChatGPT sign-in profile is stored under the same per-user app data directory…
    $env:CHATGPT_WEB_PROFILE_DIR = "C:\path\to\profile"
  4. 4Optional settings
    $env:CHATGPT_WEB_LOGIN_TIMEOUT_SECONDS = "900"
    $env:CHATGPT_HIDE_WINDOW = "0"
  5. 5Run the server in direct API mode
    $env:OPENAI_API_KEY = "sk-..."
    $env:GPT_IMAGE_BACKEND = "api"
    node dist/index.js
  6. 6Install and build
    npm install
    npm run build

README

@ramlyburger/gpt-image-2-mcp

npm version npm downloads Node.js 20 or newer Model Context Protocol server MIT license

GPT Image 2 MCP banner showing prompts flowing through an MCP server into generated images

Turn any MCP-compatible AI client into an image generator. Send a normal prompt, choose a backend mode, and get real saved image files back.

Popularity

Pulse MCP popularity ranking for GPT Image 2

PulseMCP: https://www.pulsemcp.com/servers/ramlyburger-gpt-image-2

🖼️ What It Does

  • ✍️ Prompt in: ask for an image from your MCP client.
  • ⚙️ MCP server runs: gpt-image-2-mcp handles the image request.
  • 💾 Files out: every result includes output_dir, image_path, and metadata.
  • 🔐 No ChatGPT API key needed in chatgpt-web mode. You only need a ChatGPT account and a successful sign-in at chatgpt.com.

Beginner-friendly flow from prompt to GPT Image 2 MCP to saved images

🚀 Quick Start

Add the server to your MCP client:

{
  "mcpServers": {
    "gpt-image-2": {
      "command": "npx",
      "args": ["-y", "@ramlyburger/gpt-image-2-mcp"],
      "env": {
        "GPT_IMAGE_BACKEND": "chatgpt-web"
      }
    }
  }
}

That is enough for the ChatGPT website mode. The first run opens ChatGPT so you can sign in or complete verification. After that, the local profile can be reused across restarts.

For direct API generation, set OPENAI_API_KEY and change GPT_IMAGE_BACKEND to api.

🧭 Pick A Mode

ModeWhat you needBest whenNotes
chatgpt-webA ChatGPT account and sign-in at chatgpt.comYou want a simple setup without a ChatGPT API keyGood beginner default
apiOPENAI_API_KEYYou want the direct API pathUses gpt-image-2
autoPreferably an API key; otherwise a usable ChatGPT website sessionYou want API first with fallback behaviorTries API first, then falls back only when the API backend is unavailable

🎬 Demo

Demo

Click the GIF to open the full MP4.

🧰 Tool Surface

  • generate_image(prompt, backend?, n?, size?, quality?, output_format?, conversation_mode?, timeout_seconds?)
  • backend_status(backend?)
  • browser_visibility(action?, start_browser?)

Backend values are api, chatgpt-web, or auto.

Use conversation_mode="new" or conversation_mode="continue" with the ChatGPT website mode.

📄 Technical Reference

The section below is the implementation-oriented view.

Figure 1. System Model

flowchart LR
    A["MCP client"] --> B["gpt-image-2-mcp<br/>stdio server"]
    B --> C["Input validation<br/>Zod schemas"]
    C --> D{"Backend selection"}
    D --> E["OpenAI API mode"]
    D --> F["ChatGPT website mode"]
    E --> G["Saved image files<br/>metadata.json"]
    F --> G

Academic paper-style architecture figure for GPT Image 2 MCP

Abstract

gpt-image-2-mcp is a small TypeScript MCP server that exposes image generation through a narrow tool contract. The server validates MCP tool input, resolves the requested backend, persists generated artifacts to disk, and returns structured metadata plus image content to the caller.

Method

The implementation follows a five-stage pipeline:

  1. parse and validate MCP tool input
  2. resolve the backend from api, chatgpt-web, or auto
  3. execute the selected image-generation path
  4. write generated images and metadata.json to a prompt-derived output directory
  5. return output_dir, image_path, images, and backend metadata

The auto mode attempts the API backend first and falls back to chatgpt-web only when the API backend is unavailable.

Artifact Model

Each generation creates one output directory. Images are written as numbered files such as image-01.png, and metadata is written beside them.

Default output roots:

Windows: %LOCALAPPDATA%\gpt-image-2-mcp\output\chatgpt-images
macOS:   ~/Library/Application Support/gpt-image-2-mcp/output/chatgpt-images
Linux:   ${XDG_DATA_HOME:-~/.local/share}/gpt-image-2-mcp/output/chatgpt-images

Operational notes:

  • backend_status returns the effective output_root
  • generate_image returns output_dir, image_path, and the full images array
  • image filenames are deterministic within one output directory: image-01, image-02, and so on
  • metadata is written as JSON alongside the image files

ChatGPT Website Mode

Run the server in ChatGPT website mode:

$env:GPT_IMAGE_BACKEND = "chatgpt-web"
node dist/index.js

When the server starts, it opens ChatGPT in Chrome or Edge. Sign in or complete verification there. Once the normal composer is visible, the session is ready for tool calls. No ChatGPT API key is required for this mode.

The local ChatGPT sign-in profile is stored under the same per-user app data directory by default. Override it with:

$env:CHATGPT_WEB_PROFILE_DIR = "C:\path\to\profile"

Optional settings:

$env:CHATGPT_WEB_LOGIN_TIMEOUT_SECONDS = "900"
$env:CHATGPT_HIDE_WINDOW = "0"

CHATGPT_HIDE_WINDOW defaults to enabled. The ChatGPT window stays visible for login or verification, then hides after chatgpt.com is ready. Use 0 if you want the window to remain visible after sign-in.

API Mode

Run the server in direct API mode:

$env:OPENAI_API_KEY = "sk-..."
$env:GPT_IMAGE_BACKEND = "api"
node dist/index.js

This mode uses the configured OpenAI image model directly. By default the model is gpt-image-2, and the selected output format can be png, jpeg, or webp.

Tool Contract

generate_image returns a structured result with these important fields:

  • status
  • requested_backend
  • backend
  • fallback_from
  • prompt
  • output_dir
  • image_path
  • images
  • metadata

backend_status returns readiness and configuration information for the selected backend or for both backends when auto is requested.

browser_visibility controls the visibility of the ChatGPT window and can also start the ChatGPT session when requested.

Local Development

The TypeScript MCP server is the only supported entry point.

Install and build:

npm install
npm run build

Useful local commands:

npm run typecheck
npm run build
npm run start

Repository Notes

  • src/index.ts registers the MCP tools
  • src/config.ts resolves environment-driven configuration
  • src/backends/ contains backend implementations and selection logic
  • src/output.ts is responsible for output-directory naming and file writes

The public MCP surface stays intentionally small while backend-specific behavior remains isolated in the backend layer.

Files in the repo

Repository payload9 top-level entries
  • assets
  • src
  • .gitignore
  • LICENSE
  • mcp_config.example.json
  • package-lock.json
  • package.json
  • README.md
  • tsconfig.json

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

86k

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k