Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
Sriram-PR/
doc-scraper
99
0xMassi/webclawTools
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
2.3k
adityaarsharma/
librecrawl-technical-seo-audit-mcp
The AI-native technical SEO crawler. Open-source MCP server for Claude / Cursor / Codex — 37 tools, 50+ checks, unlimited pages, WAF detection, ephemeral by design. Built on LibreCrawl. MIT.
39
us/crwConnectors
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
970