AI Crawlers
Crawling tools that use semantic models to understand pages, generate or repair extraction rules, or convert web pages into LLM-usable formats.
Scoring method (open_source_ai_tools_v1): 40% GitHub activity (time-decay since last push) + 40% stars (log-normalized within category) + 20% license permissiveness, minus 5 points if archived. Last generated: 2026-09-13T07:03:13.193747+00:00
#1
9.69
Crawl4AI
Local, self-hosted, LLM-friendly web crawler
+ 82790 GitHub stars (top of category)
+ recently active
+ permissive license (Apache-2.0)
#2
9.33
ScrapeGraphAI
Natural-language-driven graph-based data extraction framework
+ 30876 GitHub stars (top of category)
+ recently active
+ permissive license (MIT)
#3
8.8
Firecrawl
Full Web Context API and agent web-data platform
+ 179650 GitHub stars (top of category)
+ recently active
- restrictive/copyleft license (AGPL-3.0)
#4
7.36
WebClaw
Rust-built CLI, MCP, and self-hosted web extraction tool
+ recently active
- restrictive/copyleft license (AGPL-3.0)
#5
4.91
ai-web-research
Deterministic, LLM-independent robots/sitemap-aware crawler producing LLM-ready Markdown with a resumable frontier
+ recently active
No tools match your filter.