Universal Dynamic Curated DirectoryUDCD MVP
← Back to home Archive

AI Crawlers

Crawling tools that use semantic models to understand pages, generate or repair extraction rules, or convert web pages into LLM-usable formats.

Scoring method (open_source_ai_tools_v1): 40% GitHub activity (time-decay since last push) + 40% stars (log-normalized within category) + 20% license permissiveness, minus 5 points if archived. Last generated: 2026-09-13T07:03:13.193747+00:00

#1 9.69

Crawl4AI

Local, self-hosted, LLM-friendly web crawler

+ 82790 GitHub stars (top of category)
+ recently active
+ permissive license (Apache-2.0)
#2 9.33

ScrapeGraphAI

Natural-language-driven graph-based data extraction framework

+ 30876 GitHub stars (top of category)
+ recently active
+ permissive license (MIT)
#3 8.8

Firecrawl

Full Web Context API and agent web-data platform

+ 179650 GitHub stars (top of category)
+ recently active
- restrictive/copyleft license (AGPL-3.0)
#4 7.36

WebClaw

Rust-built CLI, MCP, and self-hosted web extraction tool

+ recently active
- restrictive/copyleft license (AGPL-3.0)
#5 4.91

ai-web-research

Deterministic, LLM-independent robots/sitemap-aware crawler producing LLM-ready Markdown with a resumable frontier

+ recently active

No tools match your filter.