Sandbox
22 repos for web-scraping · DataClear
yfe404/
web-scraper

Intelligent web scraping Claude Code skill with automatic strategy selection and TypeScript-first Apify Actor development

90
D4Vinci/ScraplingFrameworks & SDKs

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

80k
us/crwConnectors

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

970

Adaptive Python web scraping toolkit + MCP server for AI agents. Self-healing selectors that survive site changes, TLS-fingerprint stealth to bypass anti-bot filters, CSS/XPath parsing, and 24 built-in scrapers, clean, structured, LLM-ready data from any URL.

200
Sharan-Kumar-R/
Custom-MCP-Server

MCP server for scraping LinkedIn, Facebook, Instagram profiles and Google search.

95

scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place

5.8k

Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.

5.9k
sami-mag07/
scraping-agent-skeleton

Research and scraping agent skeleton: tool loop, loadable skills, fallback chains. No data included, configured via .env.

36
konippi/
servo-fetch

A self-contained browser engine that fetches, renders, and extracts web content as Markdown, JSON, or screenshots — no Chromium, no API key, no setup.

144

CLI, MCP server, and npm library that turns any website into an API — no docs, no SDK, no browser.

125
tinyfish-io/
agentql-mcp

Model Context Protocol server that integrates AgentQL's data extraction capabilities.

179
Sriram-PR/
doc-scraper

Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).

99
flack0x/
trendspyg

Free, maintained Python library + CLI for Google Trends: trending now, plus keyword interest over time, related queries & interest by region. A modern pytrends alternative.

49

A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

2.6k

Extract public Spotify data — tracks, albums, artists, playlists, podcasts & lyrics — without the official API. Sync + async, typed models, one dependency.

301

Your agent seeks what search can't find. A self-hosted perception MCP server that transcribes speech, reads behind logins, sees images and video frames, crosses languages, and remembers.

49
brettdavies/
crawl4ai-skill

Scrape JavaScript-heavy sites and extract structured data via reusable CSS schemas. Portable agent skill wrapping the Crawl4AI CLI and Python SDK.

46

Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.

2.3k
sderosiaux/
chrome-agent

Web tasks that compile. Other browser tools tell your AI agent the click worked. chrome-agent reads the page back after every action and answers in JSON what actually happened. One 3 MB Rust binary, no Node, no cloud.

88

This GitHub repo is a powerhouse collection of APIs you can start using immediately to build everything from simple automations to full-scale applications. One of the most valuable API lists on GitHub—period. 💪

7.6k

MCP server for AI agent payments: pay-per-call APIs (search, scraping, data, compute, media, research) with no vendor keys and a spend ceiling on every call. Works in Cursor, Claude Code, Claude Desktop, and Codex.

43