Sandbox
@xyTom/coding-tools-mcp

MCP server for coding tools and workspace access

Coding Tools MCP is a model-neutral coding runtime served over MCP. It gives an agent file search, structured patches, shell execution, interactive PTYs, and git operations while keeping actions inside one workspace and under permission modes. It works with many MCP clients, and the same server can also be exposed over HTTP or wrapped for local, Docker, and remote sandbox use.

1,038 stars173 forksPythonUpdated 13d ago
How to use Coding Tools MCP with ChatGPT, turn your ChatGPT into Codex
Coding-Tools-MCP809 views • 2 months ago
Who it's for

Builders who want their agent to read, edit, test, and review code in a real repository without hand-rolling tool access.

What it delivers

You can let an agent work on code with safer file, shell, and git access instead of stitching those tools together yourself.

What it does

Workspace-bound file tools

Read, list, search, patch, and apply multi-file changes inside one allowed workspace.

Command execution with interactive sessions

Run shell commands, keep REPLs alive across turns, stream output, and stop processes when needed.

Git tools for review and context

Inspect status, diffs, logs, blame, and file history from inside the agent loop.

Permission modes and safety checks

Gate risky commands, block path escapes, and keep execution within the workspace boundary.

MCP transport support

Serve the same tools over stdio or Streamable HTTP for different clients and setups.

Docker and remote sandbox paths

Run the server in a disposable container or through the bundled tunnel and cloud sandbox setup.

How to get it

  1. 1Run it with whichever toolchain you already have (the server is Python ≥ 3.11 from PyPI;…
    uvx coding-tools-mcp --stdio --workspace /path/to/repo   # Python toolchain
    npx coding-tools-mcp --stdio --workspace /path/to/repo   # Node toolchain

README

Coding Tools MCP

English | 简体中文

Give any AI chat or agent a safe pair of hands on your codebase.

PyPI npm Python compliance release License

Coding Tools MCP is a model-neutral coding runtime served over the Model Context Protocol: file reading and search, structured multi-file patches, command execution, interactive sessions, and git — one server that any MCP client can drive. Claude Desktop, Claude Code, Codex, Cursor, Cline, VS Code, Windsurf, Gemini CLI, or an agent you build yourself gets the default catalog of 18 battle-tested tools, confined to one workspace and gated by permission modes.

Watch the demo

Why people use it

  • It turns a chat app into a coding agent. Claude Desktop — or any MCP chat client — gets real repo access with the subscription you already have. No extra product required.
  • Safety is the product, not an afterthought. One workspace root per server. Absolute paths, .. traversal, and symlink escapes are rejected. Permission modes gate network access, shell expansion, inline scripts, and destructive commands. On Linux, Landlock adds kernel-level filesystem confinement.
  • It is model- and vendor-neutral. A truthfully annotated, mode-aware catalog — no profile switching, no annotation games. Swap models or clients freely; the runtime contract stays put.
  • It is engineered for context windows. Results are summarized, paginated, and capped by design; serialized tool-result bytes dropped 37% release-over-release on the deterministic dogfood workload with unchanged task completion.

Quickstart

Run it with whichever toolchain you already have (the server is Python ≥ 3.11 from PyPI; the npm package is a thin launcher that starts it via uv or pipx):

uvx coding-tools-mcp --stdio --workspace /path/to/repo   # Python toolchain
npx coding-tools-mcp --stdio --workspace /path/to/repo   # Node toolchain

Wire it into Claude Desktop, Claude Code, Codex, Cursor, VS Code, Windsurf, Gemini CLI, or Cline — the JSON is the same everywhere (swap uvx for npx if you prefer Node):

{
  "mcpServers": {
    "coding-tools": {
      "command": "uvx",
      "args": ["coding-tools-mcp", "--stdio", "--workspace", "/path/to/repo"]
    }
  }
}

Then ask your client: "run the test suite and fix the first failure."

Prefer HTTP? Drop --stdio and the server speaks Streamable HTTP on http://127.0.0.1:8765/mcp. Both protocol eras are served on either transport: MCP 2026-07-28 in full, with tools as the only advertised capability, and the handshake era 2025-11-25 with 2025-06-18 compatibility. Neither has sessions. A one-line installer, per-client walkthroughs, and troubleshooting live in docs/quickstart.md and docs/mcp-client-config.md.

Seven things to try

1. Make Claude Desktop your coding agent. The config above is all it takes — the chat window you already pay for can now read, patch, test, and commit-review a real repository.

2. Code on your own machine from anywhere.

CODING_TOOLS_MCP_AUTH_MODE=bearer ./integrations/tunnels/tunnel.sh cloudflared /path/to/repo

Loopback bind + authenticated HTTPS tunnel (cloudflared, ngrok, or Microsoft Dev Tunnel). Point claude.ai on your phone at https://<tunnel-host>/mcp and drive your home workstation from anywhere. ChatGPT and Grok connect through their connector settings the same way. Bearer tokens and OAuth 2.1 + PKCE (with RFC 7591 dynamic registration) are built in. → docs/remote-mcp.md

3. Let an agent loose on untrusted code — inside a disposable sandbox.

docker build -t coding-tools-mcp-sandbox:local .
docker run --rm --init -it -p 8765:8765 -v "$PWD:/workspace" coding-tools-mcp-sandbox:local

A containerized server with toolchains and caches preconfigured, safe to point at a sketchy PR and destroy afterwards. → docs/docker.md

4. Spin up a cloud sandbox with one MCP call. The bundled Cloudflare Worker control plane exposes start_coding_tools_sandbox as an MCP tool: one call dispatches a GitHub Actions runner that boots the Docker sandbox and publishes it behind an authenticated Cloudflare Tunnel. Ephemeral compute, no server of your own.

5. Drive it from a GUI.

python -m pip install "coding-tools-mcp[desktop]"
coding-tools-mcp-desktop

Per-workspace profiles, server and tunnel start/stop, credential setup with clipboard helpers, live health checks. English and 简体中文.

6. Keep an interactive command alive. exec_command starts a REPL or debugger under a real PTY; write_stdin feeds it across turns; read_output pages long output; kill_command cleans up. Long-running processes are first-class, with deadline watchdogs and bounded buffers.

7. Give your own agent production-grade hands. Building an agent loop with the Anthropic SDK or anything else? Don't hand-roll file and exec tools — speak MCP to this server and inherit the whole safety boundary. → docs/embedding.md

The tool catalog

The registry contains 19 truthfully annotated tools. The default safe and trusted modes advertise 18; dangerous also advertises request_permissions, the only mode in which that tool can grant anything. apply_patch and apply_changes are the file-mutation primitives: both are staged, baseline-checked, atomic across files, and support rollback.

GroupTools
Files & searchread_file · list_dir · list_files · search_text · apply_patch · apply_changes · view_image
Executionexec_command · write_stdin · read_output · kill_command · request_permissions (dangerous only)
Gitgit_status · git_diff · git_log · git_show · git_blame
Runtimeserver_info · check_exec_environment

Root AGENTS.md/CLAUDE.md files load automatically and come back in the instructions of initialize, or of server/discover for a client that never handshakes. Tool content is concise agent-facing text; structuredContent carries the complete machine result. Schemas and result envelopes: docs/tools-and-schemas.md · docs/runtime-contract-v0.3.md.

Safety Boundary

ModeMeant forWhat it allows
safe (default)day-to-day agent workfile tools and vetted commands; network-looking commands, shell expansion, inline scripts, and destructive commands all require explicit permission
trustedlocal developmentopens network, shell expansion, and inline scripts; keeps secret filtering and destructive-command checks
dangerousisolated containers/VMs onlydisables exec_command permission gates; workspace path boundaries still apply

Recursive listing and search exclude .git, node_modules, build outputs, virtualenvs, and caches. Commands run with workspace-bound cwd, scrubbed environment, timeouts, and output caps. Linux hosts with Landlock get kernel-enforced filesystem confinement; other platforms get an explicit warning — this is still not a complete OS sandbox, so use the Docker image or a VM for genuinely untrusted work. Details: SECURITY.md · docs/security-boundary.md · docs/permission-modes.md

Telemetry

The server sends anonymous usage telemetry (per-tool success/latency counters and version/platform dimensions — never paths, arguments, commands, or file contents) to help prioritize fixes. Disable it with CODING_TOOLS_MCP_TELEMETRY=off or DO_NOT_TRACK=1; it is automatically off in CI. CODING_TOOLS_MCP_TELEMETRY=debug prints every event to stderr instead of sending. The full event list and guarantees are in docs/telemetry.md.

Evidence, Dogfood and SWE-bench

Every release ships through a tag-triggered pipeline in which the compliance suite, real-workload benchmark, and SWE-bench harness run from the same commit that publishes to PyPI and npm — both via trusted publishing, npm with provenance. Dogfood efficiency metrics are reproducible (make dogfood-smoke) and checked in under reports/. This repository does not claim a model-generated SWE-bench leaderboard result — see docs/swe-bench.md for exactly what is and is not measured. More: COMPLIANCE.md · BENCHMARK.md · docs/dogfood.md

Documentation

Documentation mapBrowse docs by topic
Getting startedQuickstart · Client configuration · Troubleshooting
Remote & sandboxedRemote MCP · Docker sandbox · Cloud sandbox worker
Tools & contractTools and schemas · Runtime contract · Migrating to 0.3 · Permission modes
ExecutionExec recipes · Exec troubleshooting
IntegrationEmbedding · npm launcher
Security & qualitySecurity policy · Security boundary · CI and tests · Limitations · Competitive analysis

Development

python -m pip install -e ".[dev]"
make ci        # lint, typecheck, tests, protocol/integration suites, gates

The full gate matrix is in docs/ci-and-tests.md.

License

This project is licensed under the Apache License 2.0.

If you use code, documentation, substantial implementation details, or derivative work from this project, preserve the copyright notice, license notice, and NOTICE file, and clearly attribute the original project.

Project: Coding Tools MCP
Author: Coding Tools MCP Contributors
Source: https://github.com/xyTom/coding-tools-mcp

Citation metadata is available in CITATION.cff.

Files in the repo

Repository payload31 top-level entries
  • .devcontainer
  • .github
  • apps
  • benchmarks
  • coding_tools_mcp
  • docs
  • infra
  • integrations
  • media
  • packages
  • reports
  • scripts
  • tests
  • .dockerignore
  • .gitignore
  • AGENTS.md
  • BENCHMARK.md
  • CHANGELOG.md
  • CITATION.cff
  • COMPLIANCE.md
  • docker-compose.yml
  • Dockerfile
  • LICENSE
  • Makefile
  • NOTICE
  • pyproject.toml
  • README.md
  • README.zh-CN.md
  • SECURITY.md
  • SPEC.md
  • uv.lock

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k
t8y2/dbxConnectors

20 MB lightweight cross-platform database client for 90+ databases, including MySQL, PostgreSQL, SQLite, Redis, MongoDB, DuckDB, SQL Server, and Dameng. Built-in AI, MCP Server, CLI, desktop and Docker. | 轻量级跨平台数据库管理工具,支持 MySQL、PostgreSQL、SQLite、Redis、MongoDB、达梦等 90+ 数据库,提供桌面端、Docker、CLI、内置 AI 助手和 MCP Server。

19k