Sandbox
@runesleo/x-reader

MCP server for reading web content from URLs

x-reader turns a URL into unified content by detecting the platform, fetching text or transcripts, and returning structured output. It can run as a CLI, a Python library, Claude Code skills, or an MCP server with tools like `read_url`, `read_batch`, `list_inbox`, and `detect_platform`.

961 stars93 forksPythonUpdated 18d ago
Who it's for

Builders who use Claude Code or another MCP client and want one reader for posts, pages, videos, and podcasts.

What it delivers

You can hand your agent a link and get back readable, structured content instead of scraping each site by hand.

What it does

Universal URL reading

Detects the platform and fetches content from articles, tweets, video pages, RSS feeds, Telegram, and generic web pages.

MCP tools

Exposes `read_url`, `read_batch`, `list_inbox`, and `detect_platform` through `mcp_server.py`.

Video and podcast transcription

Uses subtitles or Whisper-based transcription through the optional `skills/video` layer.

Structured analysis skill

Includes `skills/analyzer` for turning content into a structured report and action items.

Browser and login fallback

Supports saved browser sessions and Playwright fallback for pages that block simple fetches.

How to get it

  1. 1Or clone and install locally
    git clone https://github.com/runesleo/x-reader.git
    cd x-reader
    pip install -e ".[all]"
    playwright install chromium
  2. 2Run
    # macOS
    brew install yt-dlp ffmpeg
    
    # Linux
    pip install yt-dlp
    apt install ffmpeg
  3. 3For Whisper transcription, get a free API key from Groq and set
    export GROQ_API_KEY=your_key_here

README

x-reader

Python 3.10+ License: MIT

Universal content reader — fetch, transcribe, and digest content from any platform.

Give it a URL (article, video, podcast, tweet), get back structured content. Works as CLI, Python library, MCP server, or Claude Code skills.

简体中文: README.zh.md / README.zh-CN.md

What It Does

Any URL → Platform Detection → Fetch Content → Unified Output
              ↓                      ↓
         auto-detect           text: Jina Reader
         7+ platforms          video: yt-dlp subtitles
                               audio: Whisper transcription
                               API: Bilibili / RSS / Telegram

The Python layer handles text fetching and YouTube subtitle extraction. The Claude Code skills (optional) add full Whisper transcription for video/podcast and AI-powered content analysis.

Three Layers

x-reader is composable. Use the layers you need:

LayerWhatFormatInstall
Python CLI/LibraryBasic content fetching + unified schemaSee InstallRequired
Claude Code SkillsVideo transcription + AI analysisCopy skills/ to your Claude Code skills directoryOptional
MCP ServerExpose reading as MCP toolspython mcp_server.pyOptional

Layer 1: Python CLI

# Fetch any URL
x-reader https://mp.weixin.qq.com/s/abc123

# Fetch a tweet
x-reader https://x.com/elonmusk/status/123456

# Fetch multiple URLs
x-reader https://url1.com https://url2.com

# Login to a platform (one-time, for browser fallback)
x-reader login xhs

# View inbox
x-reader list

Layer 2: Claude Code Skills

Requires cloning the repo (not included in pip install).

For video/podcast transcription and content analysis:

skills/
├── video/       # YouTube/Bilibili/podcast → full transcript via Whisper
└── analyzer/    # Any content → structured analysis report

Install:

export CLAUDE_SKILLS_DIR="/path/to/claude-code-skills"
mkdir -p "$CLAUDE_SKILLS_DIR"
cp -r skills/video "$CLAUDE_SKILLS_DIR/video"
cp -r skills/analyzer "$CLAUDE_SKILLS_DIR/analyzer"

Then in Claude Code, just send a YouTube/Bilibili/podcast link — the video skill auto-triggers and produces a full transcript + summary.

Layer 3: MCP Server

Requires cloning the repo (mcp_server.py is not included in pip install).

git clone https://github.com/runesleo/x-reader.git
cd x-reader
pip install -e ".[mcp]"
python mcp_server.py

The MCP server currently targets FastMCP 1.x. The mcp and all extras pin mcp<2; moving to MCP 2.x requires a server migration rather than removing the version cap.

Tools exposed:

  • read_url(url) — fetch any URL
  • read_batch(urls) — fetch multiple URLs concurrently
  • list_inbox() — view previously fetched content
  • detect_platform(url) — identify platform from URL

Claude Code config (~/.claude/claude_desktop_config.json):

{
    "mcpServers": {
        "x-reader": {
            "command": "python",
            "args": ["/path/to/x-reader/mcp_server.py"]
        }
    }
}

Supported Platforms

PlatformText FetchVideo/Audio Transcript
YouTube✅ Jina✅ yt-dlp subtitles → Groq Whisper fallback
Bilibili (B站)✅ API✅ via Claude Code skill
X / Twitter✅ oEmbed → FxTwitter → Article/Jina → Playwright
WeChat (微信公众号)✅ Jina → Playwright
Xiaohongshu (小红书)✅ Jina → Playwright*
Telegram✅ Telethon
RSS✅ feedparser
小宇宙 (Xiaoyuzhou)✅ via Claude Code skill
Apple Podcasts✅ via Claude Code skill
Any web page✅ Jina fallback

*XHS requires a one-time login: x-reader login xhs (saves session for Playwright fallback)

X Articles and login-required X pages can use a saved local browser session: x-reader login twitter

YouTube Whisper transcription requires GROQ_API_KEY — get a free key from Groq

X / Twitter Reading Path

x-reader uses a lightweight public-first chain for X:

  1. X oEmbed for fast public tweet text.
  2. FxTwitter for structured public tweet fallback.
  3. Jina Reader for public Articles and long-form pages.
  4. Generic Jina Reader for profiles and non-status X pages.
  5. Playwright with saved session for login-required content.

For Articles or gated pages, run:

x-reader login twitter
x-reader "https://x.com/user/status/123"

By default, local X cookies stay local. If you explicitly want to let Jina use your saved X session for gated Articles, set:

export X_READER_ALLOW_EXTERNAL_SESSION_COOKIES=1

Install

# From GitHub (recommended)
pip install git+https://github.com/runesleo/x-reader.git

# With Telegram support
pip install "x-reader[telegram] @ git+https://github.com/runesleo/x-reader.git"

# With browser fallback (Playwright — for XHS/WeChat anti-scraping)
pip install "x-reader[browser] @ git+https://github.com/runesleo/x-reader.git"
playwright install chromium

# With all optional dependencies
pip install "x-reader[all] @ git+https://github.com/runesleo/x-reader.git"
playwright install chromium

Or clone and install locally:

git clone https://github.com/runesleo/x-reader.git
cd x-reader
pip install -e ".[all]"
playwright install chromium

Dependencies for video/audio (optional)

# macOS
brew install yt-dlp ffmpeg

# Linux
pip install yt-dlp
apt install ffmpeg

For Whisper transcription, get a free API key from Groq and set:

export GROQ_API_KEY=your_key_here

Use as Library

import asyncio
from x_reader.reader import UniversalReader

async def main():
    reader = UniversalReader()
    content = await reader.read("https://mp.weixin.qq.com/s/abc123")
    print(content.title)
    print(content.content[:200])

asyncio.run(main())

Configuration

Copy .env.example to .env:

cp .env.example .env
VariableRequiredDescription
TG_API_IDTelegram onlyFrom https://my.telegram.org
TG_API_HASHTelegram onlyFrom https://my.telegram.org
GROQ_API_KEYWhisper onlyFrom https://console.groq.com/keys (free)
INBOX_FILENoPath to inbox JSON (default: ./unified_inbox.json)
OUTPUT_DIRNoDirectory for Markdown output (default: disabled)
OBSIDIAN_VAULTNoPath to Obsidian vault (writes to 01-收集箱/x-reader-inbox.md)

Architecture

x-reader/
├── x_reader/              # Python package
│   ├── cli.py             # CLI entry point
│   ├── reader.py          # URL dispatcher (UniversalReader)
│   ├── schema.py          # Unified data model (UnifiedContent + Inbox)
│   ├── login.py           # Browser login manager (saves sessions)
│   ├── fetchers/
│   │   ├── jina.py        # Jina Reader (universal fallback)
│   │   ├── browser.py     # Playwright headless (anti-scraping fallback)
│   │   ├── bilibili.py    # Bilibili API
│   │   ├── youtube.py     # yt-dlp subtitle extraction
│   │   ├── rss.py         # feedparser
│   │   ├── telegram.py    # Telethon
│   │   ├── twitter.py     # oEmbed → FxTwitter → Article/Jina → Playwright
│   │   ├── wechat.py      # Jina → Playwright fallback
│   │   └── xhs.py         # Jina → Playwright + session fallback
│   └── utils/
│       └── storage.py     # JSON + Markdown dual output
├── skills/                # Claude Code skills
│   ├── video/             # Video/podcast → transcript + summary
│   └── analyzer/          # Content → structured analysis
├── mcp_server.py          # MCP server entry point
└── pyproject.toml

How the Layers Work Together

User sends URL
    │
    ├─ Text content (article, tweet, WeChat)
    │   └─ Python fetcher → UnifiedContent → inbox
    │
    ├─ Video (YouTube, Bilibili, X video)
    │   ├─ Python fetcher → metadata (title, description)
    │   └─ Video skill → full transcript via subtitles/Whisper
    │
    ├─ Podcast (小宇宙, Apple Podcasts)
    │   └─ Video skill → full transcript via Whisper
    │
    └─ Analysis requested
        └─ Analyzer skill → structured report + action items

Star History

Star History Chart

Author

Leo (@runes_leo) — AI × Crypto independent builder. Trading on Polymarket, building data and trading systems with Claude Code and Codex.

leolabs.me — writing · community · open-source tools · indie projects · all platforms.

X Subscription — paid content weekly, or just buy me a coffee 😁

Learn in public, Build in public.

License

MIT

Files in the repo

Repository payload21 top-level entries
  • .github
  • docs
  • examples
  • skills
  • tests
  • x_reader
  • .editorconfig
  • .env.example
  • .gitignore
  • .python-version
  • CHANGELOG.md
  • CODE_OF_CONDUCT.md
  • LICENSE
  • mcp_server.py
  • pyproject.toml
  • README.md
  • README.zh-CN.md
  • README.zh.md
  • REVIEW-codex-pass.md
  • SECURITY.md
  • test_wechat_dom.py

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface

86k

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k

x64dbg-MCP Server is a native MCP (Model Context Protocol) plugin for x64dbg that exposes the debugger's full functionality over HTTP. Connect any MCP-compatible AI assistant and control x64dbg programmatically: set breakpoints, step through code, read memory, dump registers, and more. Built with Zig — zero dependencies, single-binary output, cros

1.9k