Common Crawl
Free, open repository of web crawl data usable by anyone.
Common Crawl is a 501(c)(3) non-profit founded in 2007 that crawls the web and freely provides its archives and datasets to the public. It maintains an open repository of web crawl data collected since 2008, totaling more than 10 petabytes. Crawls are published approximately once a month, each typically containing more than two billion web pages, hosted on Amazon Web Services through its Open Data Sponsorship Program and downloadable at no cost.
The organization's crawler, CCBot, is based on Apache Nutch and identifies itself via the User-Agent string CCBot, obeying robots.txt directives and rate-limiting its requests to individual servers. Crawl data is provided in WARC (raw HTTP responses), WAT (metadata and link graphs), and WET (plaintext extracted from HTML) formats. Monthly index files are published in the CDXJ Index and URL Index, the latter queryable with Amazon Athena. Common Crawl also publishes host-level and domain-level Web Graphs with computed Harmonic Centrality and PageRank values, and experimental data products on Hugging Face.
Common Crawl is intended for researchers and others who need wholesale extraction, transformation, and analysis of open web data without maintaining their own crawling infrastructure. The dataset is freely available at no cost.
12 alternatives to Common Crawl
Ranked by how well each tool replaces Common Crawl: shared features, audience, price and popularity.
Distributed data warehouse for massive-scale analytics.
Covers 3 of 15 key features and is open source.
Open source56 out of 100 match—A platform for analyzing large data sets with a high-level language and infrastructure for
Covers 3 of 15 key features and is open source.
Open source55 out of 100 match—Highly extensible, scalable, production-ready web crawler
Covers 4 of 15 key features and is open source.
Free planOpen source50 out of 100 matchFree- 50 out of 100 match$99/mo
Unified engine for large-scale data analytics.
Covers 1 of 15 key features and is open source.
Open source48 out of 100 match—- 48 out of 100 match—
- 48 out of 100 matchUsage-based
A platform purpose-built for high-speed data engineering.
Covers 3 of 15 key features and is open source.
Free planOpen source47 out of 100 matchFreeOpen-source self-hosted web archiving
Covers 5 of 15 key features and is open source.
Open source47 out of 100 match—- 47 out of 100 matchContact sales
The web data API to search, scrape, and interact with the web at scale, turning websites L
Covers 1 of 15 key features and is open source.
Free planOpen source47 out of 100 match$16/mo- 47 out of 100 matchContact sales