Sandbox
@jztan/pdf-mcp

MCP server for PDF search and reading

pdf-mcp gives an agent a set of PDF tools instead of a single dump of text. The server handles hybrid search, selective page reads, OCR, tables, images, charts, and corpus-wide search across a folder of PDFs.

133 stars12 forksPythonUpdated 10d ago
Who it's for

Builders who want their agent to search, read, and compare PDFs without stuffing whole documents into context.

What it delivers

You can answer questions from large or scanned PDFs by retrieving only the relevant pages and excerpts.

What it does

Hybrid PDF search

Searches by both semantic meaning and keyword matching, then returns excerpts with source coordinates.

Selective page reading

Reads specific pages or ranges on demand, instead of loading an entire PDF at once.

Corpus tools for folders

Warms, triages, and searches a whole folder of PDFs with provenance back to each document and page.

OCR and layout handling

Extracts text from scanned PDFs and handles multi-column, vertical Japanese, and other complex layouts.

Tables, images, and charts

Pulls out structured tables, embedded images, and chart data from plot geometry.

Cache and remote transport

Keeps a SQLite cache across restarts and can serve the same tools over STDIO or HTTP.

How to get it

  1. 1Run
    pip install pdf-mcp
  2. 2OCR on scanned PDFs additionally needs system Tesseract
    brew install tesseract        # macOS
    apt install tesseract-ocr     # Ubuntu/Debian
    winget install Tesseract-OCR  # Windows
  3. 3Run
    claude mcp add pdf-mcp -- pdf-mcp

README

pdf-mcp

PyPI version Python 3.10+ License: MIT GitHub Issues CI codecov Downloads

Agentic RAG over your PDFs, one file or a whole folder, as a single MCP tool.

The agent decides when to search; pdf-mcp does the retrieval and hands back excerpts. It is an MCP server that lets Claude Code and other AI agents search one PDF or a whole folder by meaning or keyword, read only the pages that matter, and cleanly pull out tables, images, and scanned text, even from multi-column and Japanese layouts, with optional CUDA acceleration for warming large corpora.

mcp-name: io.github.jztan/pdf-mcp

Try it in your browser

See what your AI agent sees →

Drop in any PDF, or a whole folder of them, and watch an agent triage the corpus, search across every document at once, and read only the pages that matter, using a fraction of the tokens. 100% client-side, no install required.

pdf-mcp browser demo: an AI agent warms a 6-PDF corpus, triages it, searches across all six documents, and reads only the matching page, with 97.3% of the corpus never entering the context window

Why pdf-mcp?

Without pdf-mcpWith pdf-mcp
Large PDFsContext overflowRead only the pages you need
Finding contentLoad everythingHybrid search: BM25 keyword + semantic
Folders of PDFsOne document at a timeWarm, triage, and search a whole folder
Warming a big folderMinutes of CPU embeddingLength-sorted small-batch CPU encode; optional CUDA embedding, one to two orders of magnitude faster on an NVIDIA card
Tables and chartsLost in raw textStructured rows, and (x, y) data from vector charts
Multi-column and vertical layoutsColumns interleavedCorrect reading order, including Japanese tategaki
Scanned PDFsNo text at allOCR via Tesseract, parallel across pages
Repeated accessRe-parse every timeSQLite cache that survives restarts
Hidden or injected textSilently ingestedFlagged as untrusted, nothing stripped

Installation

pip install pdf-mcp

That is the whole install: hybrid search, corpus tools, multi-column and CJK reading order all work out of the box.

OCR on scanned PDFs additionally needs system Tesseract:

brew install tesseract        # macOS
apt install tesseract-ocr     # Ubuntu/Debian
winget install Tesseract-OCR  # Windows

GPU embedding is optional and off by default. On an NVIDIA card it makes the embedding pass one to two orders of magnitude faster; set PDF_MCP_CUDA=1 after installing the CUDA build of onnxruntime. Setup per CUDA series is in docs/configuration.md.

Quick Start

claude mcp add pdf-mcp -- pdf-mcp

Then ask Claude to read a PDF. For Claude Desktop, VS Code, Codex CLI, Kiro, or any other MCP client, see docs/clients.md.

pdf-mcp's tools are also plain Python functions, so you can import them and hand a PDF to the Anthropic SDK without running a server. Two runnable scripts, for a question and for a whole document: examples/.

Why this exists, and what broke along the way: Claude's 100-page PDF limit and how I got around it

Tools

13 specialized tools rather than one monolithic one. Typical pattern: pdf_info to plan, pdf_search to locate (its paragraph excerpts often answer the question outright), pdf_read_pages when you need more. For a folder, pdf_corpus_overview to triage, then pdf_corpus_search.

ToolWhat it does
pdf_infoPage count, metadata, TOC summary, scanned-page detection. Call first.
pdf_searchHybrid search (keyword + semantic), page or section granularity, paragraph or context-window excerpts with source coordinates
pdf_read_pagesRead specific pages or ranges, with OCR on demand, tables, and embedded images
pdf_read_allRead a whole document in one call, byte-capped
pdf_get_tocFull table of contents for documents with many bookmarks
pdf_render_pagesRender pages as PNG for vision models: diagrams, handwriting, scans
pdf_extract_chartChart data as exact (x, y) tables, read from plot geometry
pdf_corpus_warmWarm a folder of PDFs into the cache within a time budget
pdf_corpus_overviewPer-document triage cards for a folder
pdf_corpus_searchSearch across a folder, with document and page provenance; excerpt_style="auto" picks the excerpt unit per query
pdf_cache_statsPer-document cache breakdown and total size
pdf_cache_clearClear expired or all cache entries
server_infoWhich optional features and config are active

Text returned by any of these is untrusted content extracted from a PDF. pdf_info(content_trust=True) reports hidden text a human reader cannot see, and the read tools flag it per page.

Example prompts:

"Read the PDF at /path/to/document.pdf"
"Which pages discuss supply chain risks?"
"Find sections about the training process"
"Show me what page 5 looks like"
"OCR pages 3-5 of the scanned PDF"

Full reference, every parameter and response shape: docs/tool-reference.md. Embedding model selection: docs/embedding-models.md.

Example Workflow

For a large document (e.g., a 200-page annual report):

User: "Summarize the risk factors in this annual report"

Agent workflow:
1. pdf_info("report.pdf")
   → 200 pages, TOC shows "Risk Factors" on page 89

2. pdf_search("report.pdf", "risk factors")
   → Matches with structural paragraph excerpts: each excerpt
     is the bullet, paragraph, or heading that matched, not a
     fixed-width window. Often enough to answer directly.

3. If excerpts are sufficient → synthesize answer

4. If more context needed:
   pdf_read_pages("report.pdf", "89-95")
   → Full page text for deeper reading

Remote / HTTP transport

STDIO is the default and is what every example above uses. pdf-mcp-http serves the same tools over HTTP, for clients that cannot spawn a process (the Anthropic API MCP connector, claude.ai custom connectors) and for a warm corpus shared by several clients.

export PDF_MCP_AUTH_TOKEN="$(openssl rand -hex 32)"
pdf-mcp-http

Paths resolve on the server, so an HTTP agent reads what is already there: files under an allow-listed root, or a URL the server fetches. It cannot hand over a file from its own machine. It is single-tenant and fails closed: with no auth token and no [paths] allow list, the process exits rather than serving an open endpoint.

Docker images are published to GHCR for amd64 and arm64, with everything baked in, so every tool works on the first request:

./deploy.sh              # token, image, start, health-check
cp your.pdf documents/   # this folder is the server's /data/pdfs

Read docs/remote-access.md for the trust boundary and threat model before deploying, and docs/configuration.md for setup, client config, and token rotation.

Configuration

pdf-mcp works out of the box. To restrict which paths and URL hosts the server may touch, tune cache and worker settings, or add your own content-trust phrases, see docs/configuration.md.

Roadmap

See ROADMAP.md for planned features and release history.

Contributing

Contributions are welcome. See docs/contributing.md for setup, checks, the coherence eval harness, and quality-loop guidelines.

Contributors

Thank you to everyone who has helped improve this project through code, reviews, testing, and feature requests:

@Summer907 · @ebbsanchez · @VooDisss · @DerDennisOP · @deepdmk · @TheSOV

Contributors

Per-release contributor credits are listed in the Changelog.

Security

Found a vulnerability? See SECURITY.md for the threat model, reporting channel, and expected response timeline. Please do not open a public GitHub issue for unpatched security reports.

License

MIT. See LICENSE.

Links

Blog posts

The story behind the releases. Building pdf-mcp keeps surprising me: benchmarks that go the wrong way, formats that break everything, features I had to remove. I write about that thinking in The Dispatch. Come along if that's your kind of thing.

Background, benchmarks, and design notes from building pdf-mcp:

Getting started

Corpus & multi-document search

Search & retrieval

Engineering & security

Files in the repo

Repository payload27 top-level entries
  • .github
  • benchmark_data
  • deploy
  • docs
  • examples
  • infra
  • pages
  • scripts
  • src
  • tests
  • .dockerignore
  • .env.example
  • .flake8
  • .gitignore
  • .pre-commit-config.yaml
  • CHANGELOG.md
  • CODE_OF_CONDUCT.md
  • codecov.yml
  • deploy.sh
  • docker-compose.yml
  • Dockerfile
  • LICENSE
  • pyproject.toml
  • README.md
  • SECURITY.md
  • server.json
  • uv.lock

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

86k

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k