Sandbox
30 repos for web-scrapping · Any agentClear
D4Vinci/ScraplingFrameworks & SDKs

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

80k

Adaptive Python web scraping toolkit + MCP server for AI agents. Self-healing selectors that survive site changes, TLS-fingerprint stealth to bypass anti-bot filters, CSS/XPath parsing, and 24 built-in scrapers, clean, structured, LLM-ready data from any URL.

200

🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.

7.4k
just-every/
mcp-read-website-fast

Quickly reads webpages and converts to markdown for fast, token efficient web scraping

161
antibrow/
anti-detect-browser-skills

Launch and manage anti-detect browsers with unique real-device fingerprints for multi-account operations, web scraping, ad verification, and AI agent automation.

218
Sharan-Kumar-R/
Custom-MCP-Server

MCP server for scraping LinkedIn, Facebook, Instagram profiles and Google search.

95
tinyfish-io/
agentql-mcp

Model Context Protocol server that integrates AgentQL's data extraction capabilities.

179
pinkpixel-dev/
web-scout-mcp

An MCP server providing web search and content extraction capabilities. Integrates DuckDuckGo search functionality and URL content extraction into your MCP environment, enabling AI assistants to search the web and extract webpage content.

133
runesleo/
x-reader

Universal content reader MCP Server for 10+ platforms

961
sami-mag07/
scraping-agent-skeleton

Research and scraping agent skeleton: tool loop, loadable skills, fallback chains. No data included, configured via .env.

36
konippi/
servo-fetch

A self-contained browser engine that fetches, renders, and extracts web content as Markdown, JSON, or screenshots — no Chromium, no API key, no setup.

144
Sriram-PR/
doc-scraper

Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).

99
web-agent-master/
google-search

A Playwright-based Node.js tool that bypasses search engine anti-scraping mechanisms to execute Google searches. Local alternative to SERP APIs with MCP server integration.

622
Bin-Huang/
camoufox-cli

Anti-detect browser automation CLI & Skills for AI agents — Camoufox-powered fingerprint spoofing, no bot-detectable Playwright leaks

340
EndymionLee/
PilotBrowseMCP

A browser runtime that lets AI agents control your real Chrome browser via MCP. Agents can explore websites, generate operation manuals, and reuse them to save tokens. AI操控你的真实浏览器。

99
flack0x/
trendspyg

Free, maintained Python library + CLI for Google Trends: trending now, plus keyword interest over time, related queries & interest by region. A modern pytrends alternative.

49

A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

2.6k

Extract public Spotify data — tracks, albums, artists, playlists, podcasts & lyrics — without the official API. Sync + async, typed models, one dependency.

301
AeternaLabsHQ/
pullmd

Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Claude Code skill, Reddit-aware, refreshable share links.

480

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.

49
brettdavies/
crawl4ai-skill

Scrape JavaScript-heavy sites and extract structured data via reusable CSS schemas. Portable agent skill wrapping the Crawl4AI CLI and Python SDK.

46